"Ideas have genomes" — IG-Bench for scientific lineage reasoning and lineage-grounded idea generation
Scientific ideas rarely start from a blank page: they inherit mechanisms, repair limitations, and recombine earlier work, like genomes. Existing benchmarks say little about whether AI can follow this inheritance. IG-Bench represents each paper as typed, evidence-grounded Idea Genome objects, with a GenomeDiff recording inheritance, mutation, loss, import, and insertion, across 10 domains. The strongest of 14 LLM systems reaches only 27.3% exact accuracy on lineage reasoning — a compositional bottleneck.
Paper overview (our summary)
- Field (arXiv category)cs.AI
- AuthorsYifan Zhou, Qihao Yang, Yan Li, et al. (17)
- Submitted2026-07-09
- arXiv ID2607.08758v1
Key points
- Tests whether AI can follow how scientific ideas inherit mechanisms, repair limitations, and recombine work (a genome metaphor)
- Represents each paper as typed, evidence-grounded Idea Genome objects; GenomeDiff records inheritance, mutation, loss, import, insertion
- Spans 10 domains with 1,961 lineage traces, 1,085 Idea Genome objects, 920 GenomeDiff records
- Two evaluations: IG-Exam (closed-form lineage reasoning, 42 task types) and IG-Arena (generation scored by lineage-conditioned PES)
- On 14 LLM scientists the strongest hits only 27.3% exact accuracy — a compositional bottleneck; lineage context reshuffles rankings
This work (IG-Bench) builds a new yardstick for whether AI can trace the lineage of scientific progress and extend it with meaningful ideas.
1Ideas rarely start from a blank page
The starting insight is compelling: scientific ideas rarely start from a blank page. Like biological genomes, they inherit mechanisms, repair known limitations, and recombine pieces of earlier work. Yet current benchmarks say almost nothing about whether AI systems can follow this inheritance structure.
2The Idea Genome framework
The authors present IdeaGene-Bench (IG-Bench). Following the IdeaGene framework, it represents each paper or proposal as a set of minimal, typed, evidence-grounded Idea Genome objects. A GenomeDiff then aligns these objects under six operational evolutionary dynamics, recording inheritance, mutation, loss, external import, and novel insertion.
The benchmark spans 10 scientific domains and contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records.
3IG-Exam and a second evaluation track
It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification.
IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score (PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research.
Why it matters
A read on AI for Science, autonomous hypothesis generation, and research-assistant AI. A framework quantifying the ability to inherit lineage while producing novelty helps readers gauge the limits and potential of research-assistant and idea-generation tools.
FAQ
What is the "genome of an idea" a metaphor for?
What does 27.3% exact accuracy mean?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2607.08758