The transcriptome is the complete set of transcripts (RNA molecules) present in a cell, tissue, or organism at any given time1. It identifies those genes that are being actively transcribed or expressed. Unlike the relatively static genome, the transcriptome is highly dynamic and changes in response to environmental conditions and developmental stages.

The genome can be partly changed by attaching methyl groups to a cytosine (C) or adenine (A). This converts them into N6-methyladenine, 5-methylcytosine, and N4-methylcytosine. Spontaneous deamination of 5-methylcytosine produces thymine and a mismatch between thymine (T) and guanine (G), also known as a T:G mismatch. DNA repair can occur, often converting it back into the original C:G base pair. However, an A may be substituted for a G, which creates a mutation. DNA repair enzymes will not be corrected when double-stranded DNA in the cell replicates during the cell cycle. The strand carrying the T will be complemented by an A in one of the daughter cells, such that the mutation becomes permanent.

When the parts of DNA that are in a gene promoter region, this usually turns off the transcription of the gene. In mammals, DNA methylation is essential for normal development and is important in genomic imprinting, inactivating the X chromosome, repressing the transcription of transposable elements, aging, and carcinogenesis. DNA methylation is almost exclusively found in CpG dinucleotides, with the cytosines on both strands usually being methylated. About 75% of CpG dinucleotides are methylated in somatic cells. By contrast, the genomes of most plants, invertebrates, fungi, and protists show a mosaic of methylation. That is, only specific genomic elements are methylated.

Methyl groups can also be removed from DNA and then re-established between generations in mammals. Almost all of the methyls from the parents are removed, first during gametogenesis, and again in early embryogenesis, with demethylation and remethylation occurring each time.

However, in many types of cancer, gene promoter CpG islands become hypermethylated. This causes transcriptional silencing that can be inherited by daughter cells following cell division. Alterations of DNA methylation are an important component of cancer development in many cases. Hypomethylation, in general, arises earlier and is linked to chromosomal instability and loss of imprinting, whereas hypermethylation is associated with promoters and can arise secondary to the silencing of oncogene suppressors.

In humans and other mammals, DNA methylation levels can be used to accurately estimate the age of tissues and cell types, forming an accurate epigenetic clock. Moreover, high-intensity exercise can cause reduced DNA methylation in skeletal muscle cells. Also, changes in DNA methylation in the brain may be important in memory and learning.

These changes in DNA methylation usually do not cause permanent changes in the underlying base sequence. In contrast, the way that the DNA is transcribed is temporary and dynamic. This enables cells and organisms to adapt quickly to changes in their environment and those caused by the many biological cycles that are essential for life.

There are many types of RNA. For many decades, scientists focused on messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA). These were thought to be part of a machine. Francis Crick proposed a central dogma of molecular biology: information flows from DNA to RNA to proteins, which are the machinery of life. There was even a one gene-one polypeptide hypothesis, in which every single gene codes for a single peptide or protein. That is, a gene was said to be a piece of DNA that codes for a protein or polypeptide (such as insulin).

Now, we know that there are important exceptions to this. Millions of different types of proteins can be made in humans from only about 20,000 genes. That is, as we are exposed to hundreds of thousands or millions of antigens, our immune system makes millions of antibodies. This is done by splicing portions of gene segments, such as for genes that code the variable, diversity and joining regions of antibodies in the human immune system.

We have also learned that there are other important types of RNA. Only about 2% of the nuclear human genome codes for mRNA that codes for proteins (or its precursor, hnRNA, or heterogeneous nuclear RNA). To make human proteins, first the DNA is transcribed into hnRNA, which is converted to mRNA, which can be translated into polypeptides or proteins. Proteins provide structural support, catalyze many different reactions in cells, and assist in intercellular communication. When RNA molecules act as catalysts, they are called RNA enzymes, ribozymes, or catalytic RNA.

Genes can also be silenced by small interfering (siRNAs) in plants and some invertebrates. They are being used in clinical trials for treating a variety of diseases. Some non-coding RNAs are processed by Dicer and Drosha to form Dicer and Drosha-dependent RNAs (DDRNAs), which play an essential role in the DNA-damage response of cells.

Dicer catalyzes the hydrolytic cleavage of long double-stranded DNA into shorter fragments of siRNA containing about 20 nucleotides. They enter the cytoplasm and interact with the RNA-induced silencing complex (RISC). They target mRNA by binding to as little as 6-8 nucleotides in the seed region at the 5’ end of miRNA. Argonaute proteins in the RISC can bind miRNAs, siRNAs, and piwi-interacting RNAs (piRNAs) to prevent the translation of mRNA into proteins.

The piRNAs form complexes with piwi proteins that have been linked to epigenetic and post-transcriptional gene silencing of retrotransposons and other genetic elements in germ line cells, especially during spermatogenesis. The piRNAs bind to piwi proteins to repress transposons and protect the genome from this type of mutation. They trigger the generation of another class of small RNAs, called 22G-RNAs, which silence the transposons.

So, the piRNAs act as part of the immune system that distinguishes mRNA from exogenous and deleterious RNAs. The genes coding for piwi proteins were first described as P-element induced wimpy (piwi) testis in Drosophila. They are now known to encode regulatory proteins that help maintain incomplete differentiation in stem cells and the rate of cell division in germline cells.

Some genes code for siRNAs. They ensure genomic stability by silencing endogenous selfish genetic elements such as retrotransposons and repetitive sequences. The siRNAs contain 21-25 base pairs. They are used in research to knock down specific proteins (temporarily reduce or inhibit the expression of specific genes, usually by degrading their corresponding mRNAs before they can be translated into proteins). There are also some pieces of DNA code for other types of RNA molecules, including some that are catalysts and others that regulate transcription and translation (such as siRNA).

Some of these RNAs can bind to DNA or mRNA, turning them off or on. During the life of a cell, different genes need to be turned on or off at just the right times, in response to environmental stimuli. The RNA and proteins that are needed are made at the right times and in the right locations in the cell. On the other hand, in cancer, transcription of some of the DNA (oncogenes) can be left on when it should have been turned off.

There are also thousands of long noncoding RNAs called lncRNAs that orchestrate and regulate other genes and are regulated by some key proteins2-3. At least one of them, lncRNA-p21, is directly regulated by the oncoprotein p53. It forms a lncRNA-ribonucleoprotein (RNP) by combining with a nuclear factor. It is a transcriptional repressor that facilitates p53-mediated apoptosis.

So, some lncRNAs play important roles in cancer etiology. There are also long intergenic RNAs (lincRNAs) that do not overlap protein-coding genes. They have highly conserved promoter regions that recruit the binding and direct regulation of essential transcription factors. There are also antisense lncRNAs that overlap known genes that code for proteins and intronic lncRNAs that are encoded within introns of protein coding genes. The lncRNAs are a cryptic but critical layer in the genetic regulatory code. About one third of them are linked with chromatin-modifying complexes that modulate key cellular pathways. Moreover, purified chromatin contains twice as much RNA as DNA.

Also, some genes code for small pieces of RNA, such as miRNA, rather than proteins. That is, miRNA is single-stranded and contains about 22 nucleotides. Many stay within cells, but others (extracellular miRNAs) can circulate throughout the extracellular environment. Both types can bind to mRNA and keep it from being translated into a polypeptide or protein. First, though, a longer, primary miRNA (pri-miRNA) is transcribed from DNA. The pri-miRNA has double-stranded hairpin structures that are recognized by a protein in the nucleus of the cell. This complex between the nuclear protein and the pri-mRNA associates with the Drosha protein to form a microprocessor complex.

It hydrolyzes the pri-miRNA to form a preliminary miRNA (pre-miRNA). The pre-miRNA is exported out of the cell nucleus and into the cytosol. Here, it is hydrolyzed further in a reaction catalyzed by the RNase enzyme called Dicer to make miRNA. The miRNAs can then bind to select mRNAs to stop them from being translated into proteins. This is an important part of post-transcriptional regulation of gene expression.

Noncoding RNAs (ncRNA) are also important in regulating the expression of genes involved in circadian rhythms4. For example, miRNAs can regulate gene expression by degrading mRNA or repressing translation into peptides or proteins, while lncRNAs and circular RNAs (circRNAs) can modulate chromatin organization, transcription factor activity, RNA stability, protein interactions, and competing endogenous RNA (ceRNA) networks. That is, miRNAs, lncRNAs, and circRNAs may act as molecular guides, scaffolds, decoys, or sponges that collectively fine-tune the expression of several genes.

Many miRNAs are expressed rhythmically. They regulate core oscillator components. For instance, miR-375 is a key post-transcriptional regulator of circadian rhythm through its direct interaction with the clock gene timeless.

Also, miRNAs play an important role in chromosome segregation, differentiation of developing cells, metabolism, and programmed cell death (apoptosis). Defects in miRNA have been implicated in cancer and diabetes. For example, miRNA-802 is upregulated in type II diabetes. Also, double-stranded miRNA can bind to complementary portions of DNA, turning off its transcription into mRNA. Other types of RNA, called noncoding RNA-activating, can activate transcription by interacting with the Mediator protein complex to stabilize the structure of chromatin.

However, some miRNAs are very beneficial to human health. This includes miRNAs in human milk5-6. Human milk has nutrients (proteins, fats, carbohydrates, minerals, and vitamins) and other bioactive compounds, such as lactoferrin, immunoglobulin, growth factors, oligosaccharides, and polyunsaturated fats. Unlike formula, which has a constant nutrient content, breastmilk adapts its nutrient content to the needs of infant growth at different stages. It leaves a lasting molecular signature that optimizes growth and the development of a healthy neuroendocrine immune system and neonatal intestinal health.

Human breast milk-derived exosomes (HMDEs) are bioactive vesicles that play an important role in maintaining the integrity and dynamic balance of the intestinal barrier function of babies and infants. They contain over 1,400 distinct microRNAs. Predominantly breastfed infants often have reduced DNA methylation of the glucocorticoid receptor gene, resulting in decreased cortisol reactivity to stress.

Of the approximately 2650 mature miRNAs in the human body, only a few of them (35-40) are very abundant in the central nervous system (CNS). This strongly suggests that there was highly selective evolutionary pressure for the CNS to use only these single-stranded ncRNAs for productive miRNA–mRNA interactions and the downregulation of gene expression. They are involved in intercellular signaling and help regulate embryonic development, adult neurogenesis and responses to environmental stimuli.

Moreover, DNA methylation and histone modification regulate miRNA expression. In addition, mitochondria store many miRNAs and affect their activities. At the same time, miRNAs are enriched in dendrites and synaptosomes. Six of them are upregulated in Alzheimer’s disease. They can account for much of the neuropathology of this common, age-related inflammatory neurodegenerative disease.

In addition, miRNAs are involved in Fragile X syndrome, Rett syndrome, autism spectrum disorders, major depression disorders, schizophrenia, and Parkinson’s disease. Six inducible, pro-inflammatory miRNAs that are abundant in the CNS are upregulated in sporadic Alzheimer’s disease. They appear to be key contributors to its sporadic process and can explain much of its neuropathology.

Also, the DNA sequences that miRNAs target are being used together with oncolytic viruses to express genes for a variety of therapies. That is, some viruses can be genetically modified so that they contain genes that can kill cancer cells (so they are called oncolytic). This strategy has been developed enough so that it can make potent oncolytic viruses with much lower toxicities than conventional anti-cancer drugs. They improve the delivery, duration, and therapeutic efficacy of miRNA-based therapeutics that either restore or inhibit the function of dysregulated miRNAs in cancer cells. The result is increasingly potent and safe cancer therapies.

Like hnRNA, precursors of rRNA and tRNA are hydrolyzed and then ligated to make mRNA, rRNA or tRNA. Sometimes, the precursor RNA molecules act as their own catalysts, accelerating the process. So, not just mRNA, but also other forms of RNA are important. In fact, rRNA, which is produced in the nucleolus in eukaryotic cells, accounts for most of the RNA in eukaryotic cells. Ribosomes are assembled in the nucleolus. Also, other forms of RNA can regulate gene expression and the translation of mRNA into many kinds of proteins.

There are also crucial parts of the human genome that code for the microRNAs or miRNAs. They cause some mRNAs to be translated into proteins and prevent others from being translated. Many miRNAs have important roles in cardiovascular health, Alzheimer’s disease, amyotrophic lateral sclerosis (ALS), autoimmune diseases, inflammatory responses, insulin sensitivity, and antidepressant sensitivity.

At the same time, miRNAs regulate gene expression in other mammalian cells and expressed at different levels in different tissues. Researchers are using miRNA to target sequences that control the expression of exogenous genes as well as the tissues that will support viral growth. This is being used to enhance the therapeutic indices of oncolytic viruses that express therapeutic transgenes. Their efforts have produced exciting new therapies and possible cures for many types of cancer. These oncolytic viruses are more potent, can infect many species and have lower toxicities than conventional anti-cancer drugs.

Furthermore, oncolytic viruses have been used to enhance the delivery, duration, and therapeutic efficacy of therapies based on miRNAs that are designed to either restore or inhibit the function of dysregulated miRNAs in cancer cells. Recent efforts focused on combining oncolytic viral therapy and miRNA regulation have generated much interest and investment of resources in developing potent and safe cancer therapies and cures.

The miRNAs are usually encoded within introns. First, they are first transcribed from DNA to make a long RNA, called primary RNA, or pri-RNA, which may contain many mi-RNAs, including short hairpin RNAs (shRNAs). They are partly broken down in the nucleus of the cell to make shorter (about 70 nucleotides) precursor mi-RNAs (pre-mi-RNAs). This is catalyzed by a protein complex called Microprocessor. It contains an RNA hydrolase (RNase III) called Drosha. The pre-mi-RNA is exported out of the nucleus and is hydrolyzed in a reaction catalyzed by another RNase III enzyme called Dicer, producing the mature mi-RNA.

The imperfectly complementary miRNA duplexes contain a passenger and guide strand. This binds to an RNA-induced silencing complex (RISC). The RISC is guided to its mRNA target by the miRNA strand. The passenger strand is removed, and the guide strand binds to its mRNA target. The RISC complex can silence genes by either inhibiting the initiation of translation or transporting the complex to cytoplasmic processing bodies (p-bodies) where the mRNA is deadenylated and destroyed.

One type of miRNA, called mirtrons, is not cleaved by Drosha. Instead, splicing and debranching produces a pre-miRNA substrate for Dicer. Once they are hydrolyzed by Dicer, they function the same as Dicer-dependent miRNA. There are also some non-coding RNAs that come from introns but aren’t hydrolyzed by Dicer. They are called Agotrons. Their mature structure resembles pre-miRNA. They are highly conserved in mammals and are stabilized by Argonaute proteins. They induce repression of the transcription of mRNA targets in the way that miRNA does. This is another way that short introns can regulate the expression of genes.

After many introns are removed from hnRNA they are broken down by hydrolysis. However, when an RNA gene is embedded within an intron, it is expressed to form miRNAs, small nucleolar RNAs (snoRNAs), small-interfering RNAs (siRNAs), and various long non-coding RNAs (lncRNAs). Some lncRNAs can function as oncogenes. It is thought that many genes self-regulate their expression through the ncRNAs within their introns. It should be noted that when introns were first discovered, it was thought that they were only present in eukaryotes and not in prokaryotes.

Also, miRNAs may be useful biomarkers for disease, such as early-stage lung cancer7. The transcriptomes of early-stage lung cancer patients were analyzed. The goal was to discover promising biomarkers and their associated metabolic functions. Statistical analysis found that gene expression occurs in three stages. Two major clusters were identified. They consisted of highly variant genes among the three stages. Also, 7742, 6611, and 643 genes were identified for stages I-III, respectively. Topological analysis of the protein–protein interaction network found seven candidate biomarkers.

Moreover, AI is becoming useful in discovering biomarkers and developing new therapies8. Machine learning (ML) is being used to analyze data from transcriptomes. It’s increasing our understanding of disease mechanisms. It also helps to identify key biomarkers and accelerates new drug development. and suggests therapeutic interventions. AI models have analyzed cancer transcriptomes by automating analyses and facilitating explanation of RNA-sequence data sets to reveal cancer subtypes, biomarkers, and tumor heterogeneity.

AI models divide data sets into training and testing models. Samples containing the training data sets are subject to preprocessing and feature selection to reduce the dimensions of the data. Features that represent key genes are identified while eliminating irrelevant data sets. Supervised classification algorithms are used to identify gene signatures that classify genes based on their expression patterns.

AI utilizes different models for performing RNA-sequence analysis, such as support vector machines, random forest, and logistic regression. Regression models predict target labels based on the type of training data. They analyze the expression levels of various genes that are present in the sample. Several regression models can be used to perform pathway enrichment analysis of the differentially expressed genes based on changes in gene expression. This minimizes data redundancy and noise while identifying important genes involved in pathogenesis.

There is a Cancer Treatment Response gene signature DataBase (CTR-DB)9. It is the first data resource for clinical transcriptomes that contains responses to cancer treatment. It also supports various data analysis functions, providing insights into the molecular determinants of drug resistance.

There is an upgraded version, CTR-DB 2.0. It has about 190 up-to-date source datasets with primary resistance information and 13 acquired-resistant datasets, covering 10,856 patient samples, 39 cancer types and 346 therapeutic regimens. Also, CTR-DB 2.0 added new gene set enrichment, tumor microenvironment (TME) and signature connectivity analysis functions to help identify mechanisms of drug resistance and discover candidate combinational therapies.

Furthermore, biomarker-related functions were greatly extended. CTR-DB 2.0 supported the validation of cell types in the TME as predictive biomarkers of treatment response. It was especially useful in the validation of a combinational biomarker panel and the discovery of the optimal biomarker panel using user-customized CTR-DB patient samples. The analysis of users’ own datasets, application programming interface, and data crowdfunding were also added.

So, we now know that genes code for much more than mRNA, tRNA, and rRNA. Moreover, we realize that information does not flow in one linear direction. It does not always flow from DNA to RNA to proteins. Many types of RNA and proteins can send information back to DNA in our genomes. They can affect and control when genes are transcribed and how the transcription process can be accelerated or decelerated. It is a circular and nonlinear process. The one gene - one polypeptide hypothesis is not supported by epigenetic data.

Genes in our human cells can code for either no protein at all (and just RNA), just one polypeptide or protein or many proteins. However, 99% of the protein-coding genes in our body are not in our human eukaryotic cells. They are in the bacteria in our intestines. The human body is a complex ecosystem with a diverse microbiome. Its genes are not a blueprint for a machine. They are part of a dynamic living system that adjusts rapidly and appropriately to environmental changes and the body’s many rhythms. It does this through epigenetics. So, the human transcriptome is being studied and analyzed with all the tools available to us, including AI.

Notes

1 Casamassimi, Amelia, et al. "Transcriptome profiling in human diseases: new advances and perspectives." International journal of molecular sciences 18.8 (2017): 1652.
2 Rinn JL, Chang HY."Genome regulation by long noncoding RNAs". Ann. Rev. Biochem. (2012): 81, 145-166.
3 Kyung, Jaeryoung, et al. "Multi-layered gene regulation by long non-coding RNAs: from chromatin to genome architecture". BMB reports 59.2 (2026): 112.
4 Abdelmawla, Mai A., et al. "Circadian rhythms and non-coding RNAs: mechanistic insights, clinical impact, and future opportunities for personalized medicine". Cancer Cell International (2026).
5 Markonda, Lakshmi Prasanna. "Human Milk: Nature's Epigenetic Prescription. MCN: The American Journal of Maternal/Child Nursing". 51.1 (2026): 49.
6 Chen, Gen, et al. "Human breast milk-derived exosomes and their positive role on neonatal intestinal health". Pediatric Research 98.1 (2025): 72-79.
7 Thirunavukkarasu, Muthu Kumar, et al. "Transcriptome profiling and metabolic pathway analysis towards reliable biomarker discovery in early-stage lung cancer". Journal of Applied Genetics 66.1 (2025): 115-126.
8 Bakhtiyar, Anam, et al. "AI-driven approaches in therapeutic interventions: Transforming RNA-seq analysis into biomarker discovery and drug development". Drug Discovery Today 30.7 (2025): 104391.
9 Jiang, Jianzhou, et al. "CTR-DB 2.0: an updated cancer clinical transcriptome resource, expanding primary drug resistance and newly adding acquired resistance datasets and enhancing the discovery and validation of predictive biomarkers". Nucleic Acids Research 53.D1 (2025): D1335-D1347.