Source-linked AI summary
Comment on arXiv:2607.01233: Survivorship Bias in Published-Paper Baselines for Research-Idea Distributions
Fredrik A. Dahl
TL;DR
The comment questions whether published human papers are a valid baseline for one-shot LLM ideas because publication selects which human ideas survive. It uses survivorship-bias logic and proposes stage-matched comparisons, concluding that the observed gap may not establish an LLM-specific research-taste difference.
Problem
Published human papers are survivors rather than random first-pass ideas, making their distribution potentially incomparable with one-shot LLM proposals.
Method
The comment constructs a schematic counterexample and proposes comparing first-pass human ideas with LLM proposals under matched prior-work conditions.
Results
Bridge-like ideas can be common in human brainstorming yet scarce among published papers if their publication probability is low.
Takeaways & Limitations
The human–LLM difference should be interpreted cautiously until pre-publication human ideas or stage-matched first-pass ideas are evaluated.
Takeaways & Limitations
The counterexample is schematic, not a numerical estimate, because survival probabilities are unobserved.
Abstract
from arXiv · showhide
Chen, Zhao, and Cohan introduce a valuable distributional evaluation of LLM-generated research ideas. This comment raises a narrower identification concern: their human baseline consists of published papers, whereas the LLM baseline consists of one-shot proposals. If bridge-like or synthesis-like ideas are relatively easy to generate but relatively unlikely to survive publication, then the published human baseline will understate their prevalence in the unseen human idea pool. The observed human--LLM gap may therefore be partly, or even largely, a consequence of survivorship bias.
1 Survivorship bias, not merely refinement
The published human baseline is a survivor-conditioned distribution rather than a random sample of first-pass human ideas. Because bridge-like ideas may be easy to generate but easy to reject, their scarcity among published papers may reflect selection rather than ideation differences.
- Survivorship bias, not merely refinement: The descriptive finding that LLM ideas concentrate more in bridge-like and synthesis-like opportunities than published human papers does not establish what the human reference distribution represents.The identification concern targets the baseline, not the reported descriptive comparison.
- Survivorship bias, not merely refinement: Published papers are survivors of meetings, abandoned projects, funding decisions, and peer review rather than random draws from human first-pass ideation.The resulting human statistic is conditional on survival into the published literature.
- Survivorship bias, not merely refinement: Bridge-like ideas may be common in human brainstorming yet scarce in published papers because weak versions are easy to formulate and reject as obvious or incremental.Their low publication probability can underrepresent them among published human papers even without any taxonomy change during development.
2 A counterexample distribution
A schematic counterexample shows how a bridge-heavy zero-shot human distribution could become bridge-light after publication selection. This follows standard selection-bias logic, but the figure is illustrative rather than an estimate.
- A counterexample distribution: A hypothesized bridge-heavy zero-shot human distribution can yield a bridge-light published survivor distribution after differential publication selection.The lower row is similar to the observed LLM distribution, while the upper row is similar to the observed human-paper distribution.
- A counterexample distribution: The counterexample is plausible because the paper characterizes bridge-heavy output as generic and narrow, properties that can make ideas less publishable.Under that diagnosis, their scarcity in published human papers is predicted by survivorship bias.
- A counterexample distribution: Low-publication-probability ideas would receive large inverse-publication weights when reconstructing the pre-publication distribution, but the needed survival probabilities are unobserved.Figure 1 therefore demonstrates selection-bias logic rather than providing a numerical correction.
3 What comparison is needed
The needed test compares human first-pass ideas with one-shot LLM proposals at the same stage, supplemented by samples of rejected and abandoned human work. Until then, the observed gap cannot by itself distinguish research taste from survivorship bias.
- What comparison is needed: A stage-matched test should elicit human first-pass ideas from the same reconstructed prior-work packets before development or publication filtering.If those human ideas are also bridge-heavy, survivorship bias would explain much of the reported gap; otherwise, an LLM-specific taste gap would be stronger.
- What comparison is needed: Samples of rejected submissions, abandoned drafts, and unfunded proposals could estimate the human idea distribution before publication survival.These sources are described as an “idea graveyard” for checking the pre-survival distribution.
- What comparison is needed: The evidence currently supports only that one-shot LLM proposals differ from published human papers, not that the difference is an LLM-specific research-taste gap.The commentary itself is presented as a statistics-to-article bridge, but its authorship anecdote does not resolve the identification issue.