Source-linked AI summary

Measuring the Gap Between Human and LLM Research Ideas

Ziyu Chen, Yilun Zhao, Arman Cohan

arXiv:2607.01233v1cs.CLcs.AI

TL;DR

Existing evaluations usually judge LLM research ideas individually, leaving their broader research taste relative to humans unclear. This paper compares literature-grounded human and LLM ideas using a two-axis taxonomy and finds that LLM ideas occupy a narrower region, especially around bridge-like opportunities and synthesis methods.

  • Problem

    Existing evaluations mostly judge individual LLM-generated ideas, leaving the distribution of research problems, gaps, and contributions relative to human research undercharacterized.

  • Method

    The paper compares human ideas from published papers with LLM ideas generated from reconstructed sets of related works, profiling both with a two-axis research-taste taxonomy.

  • Results

    LLM ideas occupy a narrower research-taste region, with bridge-like opportunities comprising 47.1–64.2% versus 12.1% for human ideas and synthesis methods 22.5–38.7% versus 5.1%.

  • Takeaways & Limitations

    LLM ideation should be evaluated as a distributional alignment problem because reasonable individual ideas can still reflect systematically narrower research taste.

  • Takeaways & Limitations

    The STEM-centered corpus, reconstructed local contexts, discrete taxonomy, finite model set, and one-shot settings limit generalization to other domains and ideation conditions.

Abstract

from arXiv · show

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.

1 Introduction

The introduction reframes LLM research-ideation evaluation from judging individual ideas to comparing the distribution of research tastes across human and LLM-generated ideas. Using a constrained, literature-grounded framework and two-dimensional taxonomy, the paper finds that LLM ideas occupy a narrower, connection- and synthesis-centered region than human ideas.

  • Evaluation perspective: Existing evaluations typically judge LLM-generated ideas individually by novelty, feasibility, impact, or preference rather than examining distributions of research taste.The paper defines research taste as the kinds of problems, gaps, and contributions a source produces across comparable literature-grounded ideation contexts.
  • Evaluation perspective: The distributional view captures whether a source produces varied contribution types, even when its individual ideas appear novel, feasible, and coherent.Human research includes contributions such as discovering failure modes, relaxing assumptions, building measurement instruments, offering formal explanations, and constructing systems or artifacts.
  • Evaluation framework: The evaluation uses a constrained task in which human and LLM outputs generate new motivations and methods from the same small sets of closely related papers represented by titles and abstracts.This literature grounding makes comparisons across human and LLM outputs more consistent than open-ended ideation prompts.
  • Research-taste taxonomy: Research taste is represented along two dimensions: how a proposal frames an opportunity and the intellectual contribution style used to develop it.The opportunity dimension spans missing explanations or overlooked failures to structural disconnects or limitations, while contribution styles include analytical and constructive approaches.
  • Main finding: LLM-generated ideas occupy a substantially narrower region of the taxonomy, especially favoring connection-oriented motivations and methods that integrate, reconcile, or unify existing approaches.Connection-oriented ideas frame the opportunity as linking previously separate literatures, methods, or evidence streams.

2 Related Work

Prior work uses LLMs to generate, refine, retrieve, search for, and evaluate research ideas, mainly along dimensions such as novelty, feasibility, and impact. This work instead examines whether the distribution of LLM-generated ideas resembles ideas realized in human-written scientific papers.

  • LLMs for Research Ideation: LLMs for research ideation have been studied for generating, refining, and evaluating hypotheses and research directions.Subsequent approaches include iterative refinement, retrieval-augmented generation, and search-based ideation pipelines.
  • LLMs for Research Ideation: Directly prompted LLMs can produce highly novel ideas, but these ideas are often less feasible or well-grounded than human proposals.This finding motivates methods that improve ideation through refinement, retrieval, and search.
  • LLMs for Research Ideation: Existing benchmarks evaluate generated ideas along dimensions including novelty, feasibility, and impact.The present work differs by studying distributional resemblance between LLM-generated ideas and ideas realized in human-written scientific papers.
  • Gaps between Human and LLM-Generated Content: Research on human–LLM content gaps shows that fluent and useful LLM outputs can still differ systematically from human outputs.Detection studies identify statistical artifacts in token ranks, sampling behavior, and likelihood geometry that distinguish neural generations from human text.

3 Evaluation Framework for Ideation

The evaluation framework constructs paired human and LLM idea corpora from literature-grounded ideation tasks, representing each idea by its motivation and method. It uses a two-axis research-taste taxonomy and automated annotation to compare the opportunities and contribution strategies emphasized by each source.

  • Ideation task: The framework defines a literature-grounded task in which models identify research gaps across related prior works and generate a coherent idea with motivation and method.Each instance provides prior-work titles and abstracts, while the target idea is y_i = (m_i, s_i).
  • Idea corpora: Human ideas come from published papers in major machine-learning conferences and Nature Communications, spanning 71 scientific disciplines.The human endpoint is the research idea originally devised by paper authors for publication.
  • Research-taste taxonomy: The two-axis taxonomy labels why a study is needed through opportunity patterns and how it turns that gap into a contribution through method paradigms.Opportunity examples include contradictions, missing explanations, scope mismatches, evidence gaps, disconnected literature, failure risks, and resource bottlenecks.
  • Idea corpora: LLMs generate new structured ideas whose motivation synthesizes gaps across the provided papers and whose method gives a concrete high-level approach.The evaluated model families include Claude, Gemini, GPT, DeepSeek, and Qwen.
  • Annotation and validation: An automated LLM annotator assigns primary and secondary labels, confidence scores, and diagnostic scores for surface stitching, bottleneck specificity, and boilerplate.Primary labels support distributional comparisons, while diagnostic scores support mechanism analysis; the annotator was audited on a held-out set of 150 papers.

4 Experiments

Experiments find a consistent gap between human and LLM research-idea distributions: model ideas are more concentrated around bridge-like opportunities and synthesis methods. Reasoning further sharpens this template, while mechanism analyses suggest models integrate salient technical concepts whereas humans make narrower local interventions.

  • Experimental design: The evaluation compares human and model taxonomy distributions, diagnostic scores, reasoning effects, and mechanisms using 11,683 human ideas matched to each model’s generations.The experiments condition humans and models on the same local literature context.
  • Distributional gap: Human ideas have normalized entropy above 0.92 on both axes, while model entropy ranges from 0.550 to 0.758 for opportunities and 0.723 to 0.879 for method paradigms.Even the closest model on the opportunity axis has TVD 0.348.
  • Distributional gap: Bridge or fragmentation opportunities comprise 12.1% of human ideas versus 47.1 to 64.2% for main LLMs, while explicit synthesis comprises 5.1% versus 22.5 to 38.7%.Human papers place more mass on explanation, measurement, risk, scope, artifacts, and optimization-style contributions.
  • Reasoning effects: Thinking mode moves both tested model settings farther from the human reference, increasing bridge opportunities from 49.7% to 71.1% and explicit synthesis from 38.7% to 52.2%.For Qwen3-8B, opportunity entropy drops from 0.658 to 0.481, and thinking further reduces generated-idea diversity.
  • Mechanism analysis: Mechanism analyses identify an archetype-level recipe in model ideas: select a salient technical concept cluster, then integrate or unify it with another nearby object.The operation integrate appears 7,994 times in model outputs (34.2%) but 275 times in human ideas (2.35%).
  • Mechanism analysis: Human ideas more often make narrower local interventions, such as replacing brittle components, decoupling confounded mechanisms, or formalizing local structures.Model-enriched clusters are reusable technical motifs, whereas human-enriched clusters are more local.

5 Conclusion and Discussion

The paper introduces a literature-grounded framework for comparing human and LLM research ideas using reconstructed related-work contexts and a two-axis research-taste taxonomy. It finds that LLMs occupy a much narrower region of research taste, overproducing bridge-like opportunities and synthesis-oriented method paradigms.

  • Framework: The framework compares human and LLM research ideas under shared, literature-grounded inputs.Human ideas come from real papers, while LLMs receive reconstructed related-work contexts.
  • Framework: A two-axis research-taste taxonomy labels ideas to characterize their distribution across research taste.
  • Findings: LLMs occupy a much narrower region of research taste than humans, overproducing bridge-like opportunities and synthesis-oriented method paradigms.

Limitations

The evaluation is limited by its STEM-centered corpus, reconstructed local literature contexts, and compression of nuanced ideas into discrete taxonomy labels.

  • The corpus is broad but STEM-centered, so research-taste distributions may differ in social science, humanities, clinical research, or engineering design.
  • The task reconstructs local literature contexts, unlike researchers who draw on tacit expertise, failed attempts, collaborations, reviewer feedback, and long-term research agendas.
  • The human-validated taxonomy and LLM annotation pipeline compress nuanced ideas into discrete labels.

A Dataset Details

The dataset begins with real human papers, their extracted motivations and methods, and reconstructed local literature contexts. The matched evaluation corpus retains papers with valid outputs and annotations for distributional analyses.

  • Data construction: Each data point pairs a real human paper’s extracted idea, represented by its motivation and method, with a reconstructed local literature context.The context contains proximal prior works represented by title and abstract.
  • Data construction: The source papers come from ICLR, ICML, NeurIPS, and Nature Communications.
  • Data filtering: The matched evaluation corpus uses papers for which all evaluated sources have valid outputs and annotations.Rows with missing or invalid model outputs or labels are removed before merging records by paper ID.

B Taxonomy Design

The taxonomy is designed to compare research taste rather than topic, field, or technical substrate. It separates why current knowledge is insufficient from what new contribution addresses that insufficiency, organizing ideas by opportunity patterns and research paradigms.

  • Design principles: The design draws on DARPA, NIH, NSF, and AHRQ research-gap guidance to distinguish identifying insufficiency from specifying a contribution.These sources informed the taxonomy’s guideline primitives.
  • Opportunity patterns: Opportunity patterns include Puzzle / Contradiction, Explanation Gap, Scope Mismatch, Evidence Gap, Bridge Opportunity, and Failure / Risk Gap.They capture paradoxes, missing explanations, unrealistic assumptions, inadequate evidence, disconnected areas, and reliability or risk concerns.
  • Opportunity patterns: The taxonomy also includes Resource Bottleneck, covering cost, compute, time, data, samples, experimentation, deployment, usability, and scalability constraints.Resource Bottleneck is defined as a constraint on the resources or practical conditions needed for research or deployment.
  • Research paradigms: Research paradigms include Synthesis / Unification, Relax / Extend Scope, Robustification, Formal Derivation, Empirical Mapping, Artifact / System, and Optimization / Search.These paradigms respectively cover integration, broader applicability, reduced failures or risks, formalization, systematic measurement, concrete systems, and solution improvement or discovery.

C Experimental Details and Configurations · D Prompts for Idea Generation and Annotation · E Additional Distributional Analyses

The appendices document the study’s public-data and model-compute setup, along with prompt templates for reconstructing prior work, generating ideas, and annotating research taste. The annotation prompt explicitly separates problem-finding patterns from idea-construction paradigms across research domains.

  • C Experimental Details and Configurations: The study uses public scholarly metadata, papers, abstracts, prior-work contexts, generated ideas, taxonomy labels, and annotation outputs solely for research evaluation.The released dataset will preserve source attribution.
  • C Experimental Details and Configurations: The corpus contains no private user text, recruited participant records, or human-subject data, consisting instead of public scholarly text and derived idea summaries.
  • C Experimental Details and Configurations: Each model generates one idea per input context, using specified decoding settings for local open-weight runs, GPT API runs, and thinking-mode variants.Local runs use temperature 0.6, top-p 0.95, top-k 20, and up to 2,048 new tokens; the GPT API uses temperature 1.0 with JSON-schema constraints.
  • C Experimental Details and Configurations: Distributional metrics use no hyperparameter search and apply fixed TF-IDF, MiniBatchKMeans, embedding-clustering, and probe configurations.Archetype clustering uses k = 30, batch size 512, seed 13, and n_init=auto.
  • D Prompts for Idea Generation and Annotation: The prior-work extraction prompt asks an AI research analyst to identify proximal works shaping a paper’s core idea rather than general background literature.It also requests the main innovation, motivating limitation or gap, and non-obvious contribution insight.
  • D Prompts for Idea Generation and Annotation: The idea-generation prompt supplies related papers’ titles and abstracts and asks for research gaps, opportunities, and one coherent novel proposal.Inputs are formatted as title-and-abstract blocks with citation identifiers.
  • D Prompts for Idea Generation and Annotation: The research-taste annotation prompt classifies proposals across broad research domains by separating problem-finding patterns from idea-construction paradigms.The two axes use disjoint labels, and labels from one axis must not be copied into the other.

E.1 Domain-Specific Results

The paper reports domain-specific distributional comparisons for the Machine Learning and Nature Communications corpora. It also provides domain-specific percentages for Bridge Opportunity and Synthesis / Unification.

  • Domain-specific comparisons: The main distributional comparison is presented separately for the Machine Learning and Nature Communications corpora.Results appear in Table 7 and Table 8, respectively.
  • Machine Learning: Table 7 reports Machine Learning distributional distances against the human distribution.Ent. denotes normalized entropy, and header arrows indicate the direction closer to the human distribution.
  • Nature Communications: Table 8 reports Nature Communications distributional distances against the human distribution.Ent. denotes normalized entropy, and header arrows indicate the direction closer to the human distribution.
  • Domain-specific percentages: Table 9 reports domain-specific percentages for Bridge Opportunity and Synthesis / Unification.These labels correspond to the opportunity axis and method-paradigm axis, respectively; header arrows mark the direction closer to the human distribution.

E.2 Full-Paper Context Ablation

A full-paper context ablation on 1,000 inputs replaced abstracts with model-generated summaries containing detailed motivation, methods, and insights. Richer context did not eliminate the bridge-heavy opportunity pattern, which increased for both evaluated models.

  • Ablation setup: The ablation used 1,000 inputs: 500 from the Machine Learning corpus and 500 from the Nature Communications corpus.For each related-work paper, the model read the full text and produced a compact summary covering motivation, method, and insight.
  • Opportunity pattern: For Qwen3-8B, bridge opportunities increased from 456 to 551 cases under full-paper summaries.The comparison used the same subset as the original title-and-abstract condition.
  • Opportunity pattern: For DeepSeek-V4-Flash, bridge opportunities increased from 489 to 521 cases under full-paper summaries.The richer context therefore did not remove the bridge-heavy opportunity pattern.

E.3 Prompt Ablation

The prompt ablation replaces synthesis-oriented wording with more neutral generation language, changing the idea distribution without eliminating its main qualitative pattern. For Qwen3-8B, bridge opportunities decline while failure/risk gaps increase, yet bridge opportunities remain the largest listed category.

  • Prompt design: The ablation relaxes the instruction to analyze gaps and propose a coherent idea, using neutral terms such as “generate” and “describe” instead.The original wording may encourage models to connect prior papers into synthesis-style proposals.
  • Results: The relaxed prompt changes the distribution but does not eliminate the main qualitative pattern.The comparison uses the full matched evaluation set of 11,683 papers and reports taxonomy counts plus divergence and entropy metrics.
  • Results: 5,247 bridge opportunities remain for Qwen3-8B under the relaxed prompt, down from 5,807, while failure / risk gaps rise from 987 to 1,238 cases.Bridge opportunities remain the largest listed opportunity category for Qwen3-8B.
  • Results: 6,368 bridge opportunities are produced by DeepSeek-V4-Flash under the relaxed prompt, up from 6,094.This model therefore shows the opposite bridge-opportunity count change from Qwen3-8B.
Loading 2607.01233v1…