Source: arXiv — 2026-07-09
Summary
A 16-author team introduces IdeaGene-Bench, a large benchmark — 1,961 golden lineage traces, 1,085 curated "Idea Genome" objects, and 920 pairwise GenomeDiff records across 10 scientific domains — testing whether LLMs can reason about how scientific ideas inherit, repair, and recombine prior work, and generate new ideas grounded in that lineage.
Key Takeaways
- Frames research-agent idea generation as a lineage/genealogy problem rather than open-ended brainstorming.
- Covers 10 scientific domains at large scale, with paired diff records enabling fine-grained evaluation of how well models track idea provenance.
- Relevant to the "AI research agent" thread already present in this digest, alongside coverage of Prompt-to-Paper and the coding-agent PR mutation taxonomy.
- Positions lineage-grounding as a distinct capability from raw idea novelty, which most prior scientific-idea-generation benchmarks measure instead.