1. Introduction
Emilia sonchifolia (L.) DC is a traditional Chinese herbal medicine commonly used in southern China. It is composed of the dry whole grass of
E. sonchifolia. It is named for its red spots in the green clumps and is known by several aliases including Yexiahong, Yangshicao, Yemuercai, Hongbeiye, Honghuacao, and Qishierzhihua. It is used for clearing away heat and detoxification, dispersion of blood stasis and detumescence, and it is a common component of anti-inflammatory and detoxifying Chinese medicines in China [
1,
2]. The whole plant of
E. sonchifolia can be used as medicine, and it can be harvested throughout the year. It is a Chinese herbal medicine with extremely high medicinal value, and it is also the main raw material of Huahong tablets and capsules, which are commonly used in gynecological clinics [
3].
E. sonchifolia is rich in polysaccharides, alkaloids, terpenoids, flavonoids, and other organic substances with potential medicinal value. Among them, flavonoids are the main active components, possessing anti-inflammatory, anti-oxidative, anti-viral, and anti-tumor properties, and the plant is widely used in anti-cancer therapies [
4]. Flavonoids can improve hematopoiesis and prevent DNA gene mutation [
5]. Previous studies have shown that there are significant differences in the total flavonoid content in the different vegetative organs of
E. sonchifolia, with the highest amount in the leaves, followed by stems and roots [
6]. Shen et al. [
7] isolated and identified several flavonoids in the aerial parts of
E. sonchifolia, including rhamnetin, isorhamnetin, quercetin, luteolin and tricin-7-O-β-D-glucopyranoside. Other studies on the flavonoids of
E. sonchifolia have focused on content determination [
8], the extraction process [
9] and activity analysis [
10], while the biosynthetic genes in the different parts of the plants have not been elucidated.
The combined analysis of integrated metabolome and transcriptome has become the mainstream paradigm for analyzing the regulation of secondary metabolism in medicinal plants, which effectively breaks through the technical limitations of single transcriptome or metabolome that can only unilaterally describe molecular or metabolic characteristics and cannot analyze causal regulatory relationships [
11]. As an important secondary metabolite of plants, the biosynthetic pathway of flavonoids has been relatively clear, starting from the phenylpropanoid pathway, and then catalyzed by chalcone synthase (CHS) and chalcone isomerase (CHI) to form the core skeleton of flavonoids [
12]. In the downstream branch pathway, flavonol synthase (FLS), isoflavone synthase (IFS), and anthocyanin synthesis-related enzymes (ANS/UGT) further regulate structural diversity [
13]. Recent studies have revealed the regulation of the MYB-bHLH-WD40 transcription complex and epigenetic mechanisms on the pathway [
14]. Metabolic engineering was used to achieve efficient synthesis of flavonoids in crops and microorganisms [
15]. Some researchers have identified key enzymes such as CmCHS and CmFNS and MYB and bHLH to regulate the differential accumulation of flavonoids by WGCNA-associated Chrysanthemum morifolium petal development stage genes and flavonoids and chlorogenic acid metabolites, and MYB and bHLH regulate the differential accumulation of flavonoids [
16]. Studies have combined transcriptome and extensive targeted metabolomics to analyze the differential accumulation of flavonoids in multiple tissues of
Panax japonicus, identify nine key synthetic genes and regulate WRKY transcription factors, and elucidate the regulatory mechanism of flavonoid spatial synthesis [
17]. However, the current research system for
E. sonchifolia still stays at the phenotypic and macro-functional levels, and there are still obvious knowledge shortcomings, which is difficult to support the accurate development of its active ingredients and molecular breeding optimization.
In this study, we analyzed the differences in the composition and content of flavonoids in different tissues of
E. sonchifolia, focusing on the key enzyme of flavonoid biosynthesis by using transcriptomics. The internal reference genes stably expressed in
E. sonchifolia were screened by quantitative real-time-PCR (qRT-PCR) [
18], and we determined the tissue-specific expression of the differentially expressed genes (DEGs) of these enzymes. This study provides a scientific basis for further studies on the synthesis and metabolism of flavonoids in
E. sonchifolia, in order to improve the production of its bioactive metabolites. It can serve as a guide to increase the efficiency of harvesting
E. sonchifolia and improve the production of this widely used resource.
3. Discussion
There are about 50 species in the genus Emilia, but there are few studies on the flavonoid spectrum of other representative species in the genus. In this study, the accumulation patterns and molecular regulatory networks of flavonoids in roots, stems, leaves, and flowers were systematically compared for the first time by integrating metabolomics and transcriptomics analysis. A total of 73 flavonoid metabolites were identified by metabolomics, of which 24 key differential metabolites (such as quercitrin, luteolin, naringenin chalcone, etc.) were significantly enriched in the flavonoid biosynthesis pathway (ko00941), flavonoid and flavonol biosynthesis pathway (ko00944), and isoflavone biosynthesis pathway (ko00943).
E. sonchifolia has been used as a whole herb with rich pharmacological activities. It has antibacterial, anti-inflammatory, analgesic, anti-oxidation, immune enhancement, hypoglycemic, anti-tumor, and other effects [
19,
20,
21]. Among them, flavonoids are its main active ingredients [
22]. As early as 2012, Shen Shoumao et al. [
23] studied the chemical constituents of the aerial parts of
E. sonchifolia and isolated quercetin, luteolin, and other compounds. Quercetin is a flavonoid monomer compound widely found in plants, with anti-tumor, anti-inflammatory, anti-oxidation, and hypoglycemic effects [
24]. In this study, the quantitative analysis of metabolomics showed that most of the flavonoid metabolites were higher in flowers, including quercitrin, luteolin, kaempferol, naringenin chalcone, etc. At present, in addition to quercetin, rutin, afzelin, etc., 48 metabolites such as daidzin and phloretin are reported for the first time to separate potential new components in
E. sonchifolia, which can be used as candidate targets for subsequent LC-MS targeted screening and column chromatography separation. The principal component analysis showed that the content of flavonoids in flower organs was significantly higher than that in other tissues. In particular, the average content of quercitrin in flowers was 974.55 nmol/g, accounting for more than 44% of the total flavonoids. These preliminary results indicated that the highest distribution of flavonoids was in flowers. It is speculated that flowers are the main medicinal parts of
E. sonchifolia, which provides a direction for guiding the harvest of medicinal materials and giving full play to the efficacy of
E. sonchifolia.
It has been reported that the vegetative organs and floral organs of
Tagetes patula generally contain basic flavonoids such as quercetin [
25], and enrichment of Bidens alba flavonoids in aerial tissues [
26]. The results of this study showed that the flavonoid components of
Emilia sonchifolia had significant tissue-specific distribution characteristics, and the flavonoid accumulation level in floral organs was significantly higher than that in vegetative organs such as stems and leaves. The results of this study showed that the flavonoids of
Emilia sonchifolia had significant tissue-specific distribution characteristics, and the accumulation of flavonoids in flower organs was significantly higher than that in stems and leaves. This unique metabolic distribution pattern shows potential chemotaxonomic characteristics, which further confirms the uniqueness of
Emilia sonchifolia in the allocation strategy of flavonoid metabolic organs, and provides an important basis for the study of chemical classification, phylogenetic evolution, and secondary metabolic diversity of
Emilia sonchifolia. Previous studies have shown that rutin, quercetin, and kaempferol are the main flavonol components of parsley plants [
19], and the results of this study are consistent with them. The floral organ is a sensitive reproductive tissue, and long-term natural selection promotes
Emilia sonchifolia to preferentially allocate metabolic resources to the flower and enrich flavonols to protect the reproductive structure. The stems and leaves rely on morphology and basic metabolism to resist stress, and the accumulation level of flavonoids is lower, forming the distribution characteristics of flavonoid-specific enrichment in flower organs. This differential metabolic allocation is an evolutionary strategy for the adaptation of
Emilia sonchifolia to the environment and also provides theoretical support for the screening of medicinal parts and the development of flavonoid resources. The accumulation of flavonoids in floral organs is an adaptive metabolic feature formed by the long-term evolution of
Emilia sonchifolia. A high content of flavonols can effectively absorb ultraviolet light and reduce the oxidative damage of germ cell DNA, which may be the key to ensuring the pollination and fruiting of
Emilia sonchifolia [
27]. The content of flavonols in flowers may also be able to regulate the coloration of petals and assist in attracting pollinators, thereby improving the success rate of plant reproduction. Studies have shown that flavonoids have significant inhibitory activity against fungi and bacteria, thereby reducing the risk of infection of floral diseases during flowering [
28]. Therefore, it is possible to promote the synthesis of flavonoids in petals by regulating light, water, and fertilizer conditions during the cultivation of
Emilia sonchifolia, which may not only reduce flowering diseases and reduce the cost of plant protection drugs but also improve the pollination and fruit setting rate, improve plant breeding yield and seedling cultivation efficiency.
The distribution of flavonoids in different tissues of plants is significantly different. The important reason for this tissue difference is that the enzyme content and activity related to the metabolism of flavonoids have tissue differences [
29]. Metabonomics and transcriptomics were integrated to analyze the pathway relationship between differential metabolites and differential genes in the flavonoid biosynthetic pathway of
P. multiflorum. Six key enzyme candidate genes (
CHS1,
CHI4,
F3′H5,
F3H,
FLS4,
GT) involved in the flavonoid metabolic pathway were preliminarily screened. The expression trends of different parts of
E. sonchifolia were basically consistent with the expression trends of metabolites. Through the analysis of transcription factor annotation, the core transcription factor families MYB, bHLH, and WRKY involved in the regulation of flavonoid biosynthesis in
Emilia sonchifolia have high abundance. The MYB family (including MYB-related) [
30] is the core transcription regulator of plant flavonoid synthesis, and its family members can directly activate or inhibit the expression of downstream genes by identifying the cis-acting elements on the promoters of key enzyme genes for flavonoid synthesis, such as PAL, CHS, CHI, and FLS. At the same time, MYB protein often interacts with bHLH family members to form an MBW (MYB-bHLH-WD40) ternary complex, which synergistically enhances the transcriptional regulation ability of the flavonoid biosynthesis pathway [
31]. The high abundance distribution of the bHLH family in this study further confirms the potential importance of this complex in the regulation of flavonoid biosynthesis.
CHS and CHI are key enzymes in the flavonoid synthesis pathway and play an important regulatory role. Their gene expression levels are related to metabolite production [
32]. F3′H5 is one of the key enzymes regulating the hydroxylation mode of the flavonoid B ring, which mainly plays a regulatory role in flower color modification and resistance to stress [
33]. In addition, the expression of the F3H gene is inhibited or overexpressed, which will directly lead to the down-regulation or up-regulation of flavonoid metabolism and synthesis, and can participate in the regulation of flavonoid resistance to abiotic stress [
34]. Studies have shown that dihydroflavonol promoted by F3H plays an important role in the response of Dendrobium officinale plants to stress [
35,
36]. In this study, CHI, as the rate-limiting enzyme of flavonoid skeleton synthesis, was the most active in leaves, and CHS was the most active in flowers, which was consistent with the accumulation trend of downstream metabolites. F3′H and F3H significantly affected the conversion of dihydroflavonol to flavonol by regulating the hydroxylation mode of the B ring, and their high expression in flowers was synchronized with the accumulation of kaempferol and eriodictyol. FLS is also one of the key enzymes in the synthesis of flavonols. It is highly expressed in leaves, but its catalytic product kaempferol is more abundant in flowers, suggesting that there may be a cross-tissue transport or enzyme activity regulation mechanism [
37,
38,
39]. The expression level of GT in leaves is higher than that in flowers, but the accumulation of its downstream product rutin is higher in flowers than in stems [
40]. The anti-inflammatory activity of rutin is better than that of quercetin, and the presence of glucosyl and rhamnosyl has anti-inflammatory activity.
qRT-PCR showed that the expression trends of the key enzyme candidate genes in the roots, stems, leaves, and flowers were basically consistent with the sequencing results, which suggested that the sequencing results obtained were reliable. However, there is a slight inconsistency between transcriptome data and qRT-PCR results in gene expression patterns such as CHS and CHI, which may be due to the multiple complexities of technical methods and biological backgrounds. On the one hand, the high-throughput characteristics of RNA-Seq may be limited by batch effects, gene isomer annotation bias, or low-abundance transcript detection sensitivity. Although the targeting of qRT-PCR can accurately quantify specific transcripts, it is more sensitive to primer specificity, internal reference gene stability, and sample processing timeliness. On the other hand, post-transcriptional regulatory mechanisms such as mRNA stability, translation efficiency, or tissue-specific subcellular localization differences may lead to a decoupling of transcript abundance and protein functional activity. In order to reconcile this contradiction, it is suggested to expand biological repetition to exclude individual heterogeneity, integrate metabolomics data to explore the dynamic correlation between gene expression and metabolic flux, and re-evaluate the potential interference of experimental conditions such as sampling time points on the two technologies. Future research needs to analyze the spatial and temporal specificity of ‘transcription–translation–function’ cascade regulation under the framework of multi-omics, so as to more comprehensively explain the biological logic of gene expression pattern and phenotype association.
4. Materials and Methods
4.1. Sample Collection
The samples of this study were collected in the summer of August 2023 in the Qingxiu District, Nanning City, Guangxi Zhuang Autonomous Region, at a longitude of about 108.5045° E and a latitude of about 22.7988° N. The soil was loose and humid, the temperature was about 32 °C, and the altitude was about 100 m. The roots, stems, leaves, and flowers of three positive flowering plants with the same growth size were selected as repeated samples. Each tissue had three samples, and all samples were used for RNA sequencing. The numbers allocated to the samples were as follows: CBr3-1 (root), CBr3-2 (root), CBr3-3 (root), CBs3-1 (stem), CBs3-2 (stem), CBs3-3 (stem), CBl3-1 (leaf), CBl3-2 (leaf), CBl3-3 (leaf), CBh3-1 (flower), CBh3-2 (flower), and CBh3-3 (flower). After harvesting, the samples were preserved at a low temperature, and they were returned to the laboratory for cleaning with deionized water. Samples were taken from different parts of roots, stems, leaves, and flowers. The samples were placed in liquid nitrogen for quick freezing for 5–10 min, and then they were stored at −80 °C for later use. They were sent to Wuhan Metware Biotechnology Inc. for flavonoid quantitative metabolomics measurements and transcriptome sequencing. The
Emilia sonchifolia (L.) DC specimen was collected in the Herbarium of Guangxi University of Traditional Chinese Medicine. The collection number was 20235090002, and the collector was Yao Jinli. All samples were identified as the whole plant of
Emilia sonchifolia (L.) DC by Professor Dongping Tu of Guangxi University of Chinese Medicine (
Figure 13).
4.2. Metabolomic Sequencing of E. sonchifolia
The samples of root, stem, leaf, and flower of
E. sonchifolia were dried by vacuum freezing, and then they were ground for 1.5 min to powder at 30 Hz using a grinding instrument. The internal standard of Wuhan Maiwei Metabolic Biotechnology Co., Ltd. (Wuhan, China) was adopted; that is, the company uses the internal standard to correct according to the samples provided by us. The selected internal standard is generally the isotope internal standard or structural analog of the substance to be detected (
Supplementary Table S5). In total, 20 mg of powder was weighed, and 10 μL of internal standard at a working solution of 4 mmol/L and 500 μL of 70% methanol solution were added, and the mixture was sonicated for 30 min. The samples were centrifuged at 4 °C and 12,000 rpm for 5 min, and the supernatant was filtered through a 0.22 μm filter membrane. The samples were stored in the injection bottle and analyzed by ultra-high performance liquid chromatography-mass spectrometry (UPLC-MS). The UPLC analysis conditions were: Waters ACQUITY UPLC HSS T3 C18 column (1.8 μm, 100 mm × 2.1 mm); the mobile phase A was added with 0.05% formic acid, and the mobile phase B was added with 0.05% formic acid. The elution gradient was 0 min A/B 90: 10 (
v/
v), 1 min A/B 80: 20 (
v/
v), 9 min 30: 70 (
v/
v), 12.5 min A/B 5: 95 (
v/
v), 13.5 min 5: 95 (
v/
v), 13.6 min 90: 10 (
v/
v), 15 min 90: 10 (
v/
v). The flow rate was 0.35 mL/min, and the column temperature was maintained at 40 °C. The injection volume was 2 μL.
The mass spectrometry analysis conditions used were electrospray ionization (ES) at a temperature of 550 °C, mass spectrometry at a voltage of 5500 V (positive ion mode), and mass spectrometry at a voltage of 4500 V (negative ion mode), with curtain gas (CUR) at 35 psi. In the triple quadrupole (QQQ), each ion pair was scanned and detected based on the optimized de-clustering potential (DP) and collision energy (CE). Subsequently, the Metware database (MWDB) constructed by Wuhan Metware Biotechnology Co., Ltd. (Wuhan, China) standardized the original readings by the normalization method and corrected the sequencing depth of each sample, and the statistical model was used to calculate the hypothesis test probability (p-value). Finally, multiple hypothesis test correction was performed to obtain the FDR value (false discovery rate), and, finally, the flavonoid components in E. sonchifolia root, stem, leaf, and flower were qualitatively and quantitatively analyzed. After using Analyst 1.6.3 software to process the mass spectrometry data, the chromatographic peaks detected in different samples were corrected by MultiQuant 3.0.3 software. The detected peak area ratio obtained was substituted into the standard curve linear equation for calculation, which ultimately resulted in the content data of the metabolites in the samples.
4.3. Transcriptomic Sequencing of E. sonchifolia
The total RNA of roots, stems, leaves, and flowers was extracted using a polysaccharide polyphenol plant total RNA extraction kit. The integrity of RNA and the presence of DNA contamination were assessed by agarose gel electrophoresis. The RNA concentration was measured using a Qubit 4.0/MD microplate reader (Thermo Fisher Scientific, Waltham, MA, USA), and for the integrity of RNA, a Qsep400 biological analyzer (BiOptic, New Taipei City, Taiwan) was used. Following assessment, a library was constructed, and this was qualified; then, an Illumina platform was used for sequencing. The raw reads were obtained, and the fastp software (0.23.2) was used to control the quality of the data, and some of the reads were removed. After removal of the N content that exceeded 10% of the read base number, those with low quality (Q ≤ 20) that exceeded 50% of the read base number were the high-quality clean reads were obtained. These were then assembled using Trinity software (v2.13.2). DIAMOND BLASTX software (v2.0.9) was used to compare the unigene sequences obtained with the KEGG NR, Swiss-Prot, GO COG/KOG, and TrEMBL databases.
After predicting the amino acid sequences of the unigenes, HMMER software (3.3.2) was used to compare the data obtained with the PFAM database to obtain their annotation information. The expression levels of transcripts were calculated by using fragments per kilobase of transcript per million fragments (FPKM). A differential analysis was performed using DESeq2 software (1.22.2), and the false discovery rate (FDR) was obtained. The screening criteria for DEGs were set as |log2Fold Change| ≥ 1 and FDR < 0.05. After that, the selected DEGs were compared with the KEGG database for functional annotation and enrichment, and to analyze and predict their functions and the metabolic pathways they were involved with.
4.4. Combined Analysis of Metabolomics and Transcriptomics
Based on the screening results of differential genes and differential metabolites in flavonoid metabolic pathways, through multi-dimensional bioinformatics analysis, including FPKM standardization of gene expression profiles and LC-MS/MS mass spectrometry peak area normalization calculation of metabolites, the transcriptome sequencing data and metabolomics characteristic peak clustering analysis results were co-located in the KEGG database (ko00940, ko00941, ko00943, ko00944). Cytoscape platform (3.10.4) was used to construct the metabolic network map of flavonoids. Combined with weighted gene co-expression network analysis (WGCNA) and metabolite ion flow fingerprint (EIC) technology, the regulatory relationship between gene expression and metabolite accumulation was revealed. The spatial distribution characteristics of tissue-specific metabolomics and transcriptome sequencing data were further integrated, and the differential accumulation of flavonoids in different tissues and organs (roots, stems, leaves, and flowers) and its molecular regulation mechanism were elucidated by system biology methods. It provides an important theoretical basis for further understanding the biosynthetic pathway of flavonoids and their tissue-specific accumulation.
4.5. qRT-PCR Verification of the Key Enzyme Genes
According to the literature research and the data of the
E. sonchifolia transcriptome, the DESeq2 algorithm (|log2FC| ≥ 1.5,
padj < 0.01) was used to screen the core regulatory gene clusters of the flavonoid biosynthesis pathway. Combined with the metabolomics data, MapMan 3.6.0 was used to carry out hierarchical clustering analysis of metabolic pathways, and a pathway map based on Z-score standardized metabolite content and FPKM gene expression was constructed. Finally, four internal reference genes, Tubulin, Ubiquitin, GAPDH, and UPL, were selected as candidate genes, and six key enzyme genes involved in the synthesis of
E. sonchifolia flavonoids were selected. Primer Premier 5.0 software was used to design primers (
Supplementary Table S1) according to their CDS sequences. The gene sequences were synthesized by Nanning Dennis Biotechnology Co., Ltd. (Nanning, China). The HiScript
® III RT SuperMix for qPCR (+gDNA wiper) kit was used for the measurements. The total reaction system was 10 μL and consisted of a 0.5 μL cDNA template, 0.195 μL each of the forward and reverse primers, 4.885 μL of 2 × ChamQ SYBR qPCR Master Mix, 0.195 μL of 50 x ROX Reference Dye 1, and 4.03 μL pure water. The qRT-PCR program (CFX96) used was as follows: 95 °C for 3 min; 95 °C for 10 s and 60 °C for 30 s for a total of 39 cycles. This was followed by a melting curve analysis, starting at 65 °C in steps of 5 s until 95 °C was used with a final step at that temperature for 50 s. Each sample had three replicates, and the relative expression was calculated using the 2
−ΔΔCt method.