NSF grant ~$1.69M: predicting the "start switch" of plant genomes with AI for crop improvement — BIO-AI (Michigan State)
An NSF award of ~$1.69 million to Michigan State University. Where a gene begins to be read — the transcription start site (TSS) — is a switch for when a gene turns on and which protein version it produces, yet no complete map exists across plant diversity. The project builds BioLearnTSS, a machine-learning framework predicting TSS location and core-promoter structure from DNA sequence, validated from algae to crops, with a maize focus on how TSS choice alters splicing and translation. Runs 2026-2029.
Grant overview (primary data)
- Award amount$1,689,412
- RecipientMichigan State University (Michigan)
- ProgramTools, Inform & Func Geno
- Period2026-09-01 〜 2029-08-31
- FunderU.S. National Science Foundation (NSF) / NSF
Key points
- Recipient: Michigan State University, ~$1.69M, September 2026 to August 2029 (BIO-AI program)
- The TSS is a switch for gene on/off and which protein version is made, but no cross-plant map exists
- Develops BioLearnTSS, an ML framework predicting TSS location and core-promoter structure from DNA sequence
- Trained on flowering plants, generalization tested on a conifer, green alga, and moss; maize focus on how TSS choice alters splicing and translation
- Genome editing in maize and Arabidopsis validates the mechanism; datasets, tools, and edited resources released for agricultural innovation
What makes this award interesting is that it points AI not at language or images but at how a plant genome begins to be read. As the program name BIO-AI signals, it throws machine learning at an unsolved problem in biology.
1One genome, different uses in every cell
The underlying biology is elegant. Every cell in a plant carries the same blueprint yet uses it differently, and the key to that control is where a gene begins to be read — the transcription start site (TSS). The TSS acts as a switch determining not only when a gene turns on but which version of a protein it ultimately produces. Yet no complete map of these central switches exists across plant diversity, from algae to crops.
2The BioLearnTSS framework
The core of the project is BioLearnTSS, a machine-learning framework that predicts, directly from DNA sequence, TSS locations and core-promoter structure. Training on flowering plants and then testing generalizability on a conifer, a green alga, and a moss — lineages far apart — sits squarely at the confluence of modern genomics and AI: reading function from sequence.
It also examines how alternative TSS usage shapes the proteome by altering splicing and translation efficiency, focused on maize.
3Testing in maize and Arabidopsis
Experiments accompany the modeling. Using genome editing in maize and the model plant Arabidopsis, the team directly tests how specific regulatory DNA elements within the core promoter influence TSS selection, gene expression, and protein production.
The work is organized around three aims: ML models for TSS prediction across lineages; the effect of alternative TSS usage on splicing and translation; and how core-promoter elements shape TSS selection, mRNA accumulation, and splicing.
Why it matters
A case of AI-for-Science-meets-agriculture: applying sequence-to-function AI to plant genomics for crop improvement. A model predicting function from DNA sequence, validated across lineages, is a reference for those tracking genomics, agricultural biotech, and sequence foundation models.
FAQ
What is a transcription start site (TSS)?
Why use machine learning?
How does it help agriculture?
Sources (primary)
Source: NSF Award Search (U.S. National Science Foundation, public domain). Amounts are the obligated amount. For privacy, we do not handle principal investigator names.
- NSF Award (original, official)
- NSF Award ID: 2620375