Source-linked AI summary

An Investigation of the NeurIPS and ICML 2025 Position Tracks

Fan Yang, Wenkai Li, Jun Liu

arXiv:2608.16894v1cs.CYcs.CL

TL;DR

The paper asks what kinds of position papers become visible in the newly established NeurIPS and ICML tracks and audits their early composition. It finds a pool dominated by reformist critique and recommends CFP-level signals to solicit direction-setting work with operationalizing artifacts alongside those critiques.

  • Problem

    The early composition of position tracks warrants auditing because agenda-setting positions can become actionable when they provide artifacts for others to build on, test, or contest.

  • Method

    The authors audit 191 publicly accessible, reviewed NeurIPS 2025 and ICML 2025 submissions using a pre-specified rubric and compare them with historical agenda-shifting ML papers.

  • Results

    Three-quarters of audited papers are reformist critiques, while the reference class more often provides operationalizing artifacts; evidentiary depth does not reliably predict reviewer score.

  • Takeaways & Limitations

    The track should solicit direction-setting papers with operationalizing artifacts alongside, rather than instead of, rigorous reformist critiques through modest CFP-level interventions.

  • Takeaways & Limitations

    The audit is cross-sectional, covering only NeurIPS 2025 and ICML 2025 at one point in time.

Abstract

from arXiv · show

ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position Paper Tracks were created for agenda-setting work, making their early composition worth auditing. \textbf{This paper argues that the publicly accessible 2025 reviewed pool is dominated by reformist critique, and that the track should explicitly solicit direction-setting work alongside, not in place of, the reformist critiques it already hosts well.} We audit every accessible submission to the NeurIPS 2025 and ICML 2025 Position Tracks under a pre-specified rubric, and compare the resulting pattern with a reference class of widely recognized agenda-shifting ML papers. Three-quarters of audited submissions critique an existing benchmark, evaluation, or methodology; these papers score highly on our artifact-coupling rubric, but evidentiary depth does not predict reviewer rating. The reference class (AlexNet, the Transformer, Concrete Problems in AI Safety, and others) differs from the accessible reviewed pool in \emph{artifact kind}: agenda-shifting papers typically gave the field something new to build on, test against, or contest, such as a measurement protocol, benchmark proposal, toy implementation, dataset card, audit template, or falsifiable experimental program. We close with four CFP-level interventions aimed at broadening the submission mix without displacing the critiques the track already hosts well.

1 Introduction

The paper audits the publicly accessible reviewed pools of the NeurIPS 2025 and ICML 2025 Position Paper Tracks to examine which forms of position work become visible and reviewable. It argues that the pool favors reformist critique and recommends explicitly soliciting direction-setting work with operationalizing artifacts alongside those critiques.

  • Motivation: Agenda-shifting claims become consequential when they give subsequent researchers something to build on, test against, or contest.Examples include measurement protocols, benchmark proposals, toy implementations, dataset cards, audit templates, and falsifiable experimental programs.
  • Venue context: The NeurIPS and ICML Position Paper Tracks were created to advance new perspectives, ideas, or research directions and shape future scholarly dialogue.As formative venue-design experiments, their reviewed papers teach authors what position papers and reviewable evidence are expected to look like.
  • Audit design: N = 191 publicly accessible reviewed papers from the NeurIPS 2025 and ICML 2025 Position Paper Tracks were classified using a pre-specified rubric.The audit covers claim type, artifact relation, falsifiability, and evidentiary structure, but does not estimate the full submission distribution because withdrawals and desk rejections are not public.
  • Core finding: Roughly three-quarters of audited papers are reformist critiques that adjust existing benchmarks, evaluations, datasets, or methodologies.Many are rigorous and valuable; the concern is that reformist critique may become the most legible and reviewable form of position work by default.
  • Implication: The track should solicit direction-setting papers, especially those pairing a position with an operationalizing artifact, alongside rather than instead of reformist critiques.This recommendation is also motivated by the rubric’s finding that evidentiary depth does not reliably predict reviewer score.

2 An Audit of NeurIPS 2025 and ICML 2025 Position Papers

The audited 2025 Position Track pool contains 191 accessible, reviewed papers and is dominated by reformist claims. Papers are generally artifact-coupled, but artifact coupling does not predict reviewer scores, while explicit falsifiability is nearly absent.

  • Corpus and method: 191 papers comprised the accessible corpus: 95 NeurIPS submissions and 96 ICML submissions with retrievable PDFs and at least one complete review.The corpus included 40 accepted and 55 rejected NeurIPS papers, plus 73 accepted and 23 rejected ICML papers.
  • Claim-type composition: 74.9% of submissions were reformist, compared with 17.3% normative, 5.8% descriptive, and 2.1% predictive.The reformist plurality held in both venues: 70.5% at NeurIPS and 79.2% at ICML.
  • Artifact coupling: The mean artifact-coupling score was 3.08/5, with 43% of papers at the modal score of 4 and only 12.6% lacking an identifiable artifact.Artifacts included measurement studies in 46% of papers, existing-benchmark critiques in 23%, and concrete experiments in 17%.
  • Artifact coupling: Artifact coupling did not predict reviewer score in the normalized pooled sample, with Spearman ρ = +0.042, p = 0.57.Within-venue associations were also near zero: ρ = −0.028 for NeurIPS and ρ = +0.067 for ICML.
  • Falsifiability: Only 1 of 191 papers (0.5%) stated an explicit falsifiability condition, although every paper was judged to contain an inferable falsifiability claim.The contrast indicates that testable implications are often inferable but almost never explicitly stated.

3 Historical Anchor: A Narrative Contrast

The historical reference class is qualitative rather than representative, but it highlights a recurring pattern: agenda-shifting papers coupled claims to new, load-bearing artifacts. This historical mode is scarce in the accessible 2025 pool, while reviewer ratings and acceptance do not vary systematically by artifact type.

  • Historical reference class: The reference class is a qualitative contrast, not a statistically representative sample, used to identify artifacts that anchored new research directions.The rubric interprets these papers at the text level and separates stated claim form from eventual agenda-setting influence.
  • Composition versus evaluation: Neither reviewer rating nor acceptance varies systematically across artifact types in the pooled sample.Table 2 reports pooled N = 191, with ICML ratings normalized to a 1–10 scale; the released_dataset row concerns submission composition, not reviewer preference.
  • Artifact coupling: Agenda-shifting papers typically bundled claims with a new load-bearing artifact rather than only critiquing an existing benchmark or evaluation.Examples include measurement protocols, benchmark proposals, toy implementations, dataset cards, audit templates, and falsifiable experimental programs.
  • Composition versus evaluation: The historical mode of coupling is scarce in the accessible pool, making the gap compositional rather than evaluative.The paper reports no evidence that reviewer ratings track artifact coupling when it appears.

4 Mechanisms: Why the 2025 Reviewed Pool Is Dominated by Reformist Critique

Five interacting mechanisms help explain why reformist critique dominates the publicly accessible 2025 reviewed pool: it is easier to falsify, empirically support, and evaluate, while career incentives and track criteria offer weaker support for normative work. Together, these local pressures produce a field-level signal about what position-paper work looks like, consistent with rigorous submissions not predicting reviewer ratings and differing in artifact kind from the historical reference class.

  • Local incentives and reviewability: Reformist claims are locally falsifiable, whereas normative claims require subjective judgments about research’s long-run trajectory, encouraging reviewers under workload pressure to prefer reformist work.A benchmark-overfitting claim can be checked experimentally; a claim about what the field should prioritize is harder to defend in a written review.
  • Local incentives and reviewability: ML’s evidentiary culture rewards benchmark numbers, ablations, and error bars, giving reformist critiques crisper empirical support than proposals for directions that do not yet exist.The field’s experimental self-image makes empirical rigor especially legible for critiques of existing practices.
  • Career incentives: Career incentives favor discrete, citable reformist contributions, while genuinely normative proposals risk being viewed as speculation and expose junior researchers to asymmetric risk.The current track design does little to offset this asymmetry.
  • Track criteria: Review criteria ask for novelty, rigor, and significance without operationalizing agenda-setting quality, so reviewers default to the dimensions of rigor they can readily observe.Rigor is legible for testing whether a critique holds but more opaque for judging whether a proposed direction is important.
  • Aggregate signal: These local pressures aggregate into a field-level signal: visible pools dominated by one contribution form teach junior researchers what a “Position Paper” looks like.The evidence is compositional rather than longitudinal, and the mechanisms jointly fit rigorous submissions that do not predict ratings and differ in artifact kind from the historical reference class.

5 Implications: Four Interventions for 2026 and 2027

The paper proposes four CFP-level interventions to broaden the Position Paper Track toward direction-setting work without discouraging rigorous reformist critiques. These interventions make claim types, falsifiability, review criteria, and editorial solicitation more explicit.

  • Intervention 1: Claim-type field: Require authors to self-report whether their primary claim is normative, predictive, reformist, or descriptive, while keeping the field visible to reviewers but outside acceptance filtering.The intervention is cheap and reversible, and creates a dataset for checking whether authors’ positioning matches the LLM-audited distribution.
  • Intervention 2: Falsifiability statement: Require a one-sentence falsifiability statement specifying what observation or experiment would undermine the paper’s central claim.Only 0.5% of submissions contained an explicit falsifiability sentence, despite the classifier inferring one in 100% of papers; vacuous statements would prompt revision, not desk rejection.
  • Intervention 3: Claim-type review criteria: Split review criteria by claim type: assess critique validity and actionable alternatives for reformist papers, and specification, novelty, and evidentiary conditions for normative or predictive papers.The proposal responds to the concern that rigor currently means different things across claim types and that bundled criteria favor the most legible form.
  • Intervention 4: Explicit solicitation: Explicitly solicit direction-setting submissions, including normative or predictive claims and positions paired with new operationalizing artifacts rather than critiques alone.This makes deliberate the editorial preference already communicated by the visible reviewed pool’s default composition, while preserving reformist submissions’ role.

6 Limitations

The study’s main limitations concern classifier noise, a post-hoc normative-versus-reformist analysis, venue and outcome-measure constraints, cross-sectional evidence, and claim type as an imperfect proxy for direction-setting. The paper therefore treats F1, F3, F5, and the historical reference class—not F4—as its argumentative foundation.

  • Empirical limitations: 0.50 cross-model agreement between Gemini and Claude indicates classifier noise, but F3 remains null overall and within each venue.The pooled correlation is ρ = +0.042, with ρ = −0.028 for NeurIPS and ρ = +0.067 for ICML; Gemini’s 3-run self-consistency is 0.85.
  • Empirical limitations: F4 is post-hoc and its raw-pooled p = 0.034 effect disappears under ICML×2 normalization, leaving only a modest within-venue trend.The paper states that treating F4 as unsupported is legitimate and that its core argument does not depend on F4.
  • Empirical limitations: Auditing only NeurIPS 2025 and ICML 2025 leaves open whether reformist dominance is venue-specific, while reviewer ratings are an imperfect outcome measure.Within-corpus acceptance leaves F3 null at p = 0.72; acceptance rates are also obscured by withdrawn rejections and cross-venue scale differences.
  • Empirical limitations: The audit is cross-sectional rather than longitudinal and supports only the descriptive claim that the publicly accessible 2025 reviewed pool is dominated by reformist critique.The substantive prescription remains unchanged if the track’s charter names direction-setting as its purpose.
  • Empirical limitations: Claim type is only a weak proxy for direction-setting because the substantive target is artifact mode, which Intervention 4 addresses directly.The paper uses claim type in Interventions 1 and 3 because it is easier to self-report and apply than artifact mode.

7 Alternative Views

The section addresses three objections: reformism may precede direction-setting, the distinction may be artificial, and engineering culture may favor reformist work. It accepts parts of each objection but argues that explicit solicitation remains necessary because the artifact-mode gap persists.

  • Alternative 1: Reformism may be a prerequisite for direction-setting, but the 74.9% reformist plurality suggests the prerequisite is abundant while downstream direction-setting remains undersolicited.The authors accept that historical agenda-setting examples often build on earlier critiques, but reject the assumption that direction-setting follows automatically.
  • Alternative 2: Many papers combine critique and proposal, but the authors distinguish claim type from artifact type when assessing whether a position is operationalized.A paper may be reformist and normative in prose while still revealing whether its position centers something new or critiques an existing object.
  • Alternative 2: Of 191 submissions, the pool rarely centers new operationalizing artifacts, so ambiguity in claim classification does not eliminate the artifact-mode gap.The authors grant that primary-orientation coding compresses papers that bundle critique with proposal, while maintaining that artifact-centered contributions remain scarce.
  • Alternative 3: ML’s engineering culture supports iterative refinement, but the authors reject the claim that it rules out direction-setting as a legitimate mode.They argue that the historical reference class reflects both iterative refinement and occasional direction-setting, rather than only atypical exceptions.

8 Conclusion · A Classification Rubric, Pipeline, and Validation Protocol · B Additional Analyses

The publicly accessible 2025 Position Track pool is compositionally dominated by reformist critiques, while agenda-shifting work typically contributes operational artifacts for future research. The paper therefore recommends modest CFP-level nudges to broaden direction-setting submissions and documents a reproducible, caveated audit pipeline plus post-hoc supporting analyses.

  • 8 Conclusion: Reformist critiques outnumber all other claim types combined by nearly 3:1 in the publicly accessible 2025 reviewed pool.The track was created to advance new perspectives and research directions, but its visible equilibrium is narrower.
  • 8 Conclusion: The stronger evidence is compositional: reviewed papers differ in what artifacts they ask the community to build on, test against, or contest, not in demonstrated reviewer rejection of ambition.Review scores show no reliable sensitivity to rubric-captured evidentiary depth, and the ratings advantage for ambition is too fragile to support the argument.
  • 8 Conclusion: The proposed CFP interventions are modest nudges: normalize lightweight operational artifacts, ask what subsequent work a position enables, encourage genuine contestation, and welcome directions alongside critiques.They are not presented as remedies for deeper structural forces affecting production, resourcing, and review.
  • A Classification Rubric, Pipeline, and Validation Protocol: The rubric classifies primary claim orientation, artifact coupling, artifact type, falsifiability, target audience, and benchmark-data citation using fixed schemas and mutually exclusive artifact categories.Normative claims concern where the field should go, whereas reformist claims concern how an existing practice should change.
  • A Classification Rubric, Pipeline, and Validation Protocol: 80%+ of papers score ≥2 on artifact coupling, but the 0.50 cross-model agreement for that field cautions against interpreting fine score differences.Level 2 denotes re-analysis or critique of an existing benchmark, while level 4 executes a load-bearing measurement study.
  • A Classification Rubric, Pipeline, and Validation Protocol: The audit retained 191 submissions with retrievable PDFs and completed reviews, extracted their text, classified each paper three times, and used Claude Sonnet 4.6 for cross-model validation.The pipeline records per-field self-agreement and applies statistical tests to consolidated labels; code is slated for release with the final paper.
  • A Classification Rubric, Pipeline, and Validation Protocol: The reported 0.85 claim-type agreement is an upper bound because Claude saw Gemini’s output, whereas Gemini self-consistency scores provide the cleaner reliability signal.Gemini’s three-run self-consistency was 0.85 for artifact coupling, 0.90 for claim type, 0.84 for artifact type, and ≥0.93 for other fields.
  • B Additional Analyses: The additional analyses are post-hoc rather than pre-specified in H-A1–H-A4 and are reported transparently as supporting, not primary, evidence.The validation covers the full corpus, N = 191, with exact-match agreement computed for each field, including the 0–5 coupling score.

B.1 Pre-specified hypothesis tests in full

The four pre-specified tests are largely null or unusable: pooled associations between artifact coupling and reviewer outcomes are null, cross-venue sign consistency fails, and falsifiability testing is degenerate.

  • Pre-specified hypothesis tests: Both H-A1 and H-A2 are null when pooled across venues.H-A1 uses the ICML×2-normalized rating; H-A2 compares artifact-coupling scores by acceptance status.
  • Pre-specified hypothesis tests: H-A4 fails because H-A1 and H-A2 do not show sign consistency across venues.The test evaluates whether both hypotheses have consistent signs across the two venues.
  • Pre-specified hypothesis tests: H-A3 is degenerate: falsifiability_inferred is True for 100% of papers, leaving no variance for Fisher’s exact test.The test examines falsifiability_inferred × acceptance.

B.2 Artifact-type composition by venue

NeurIPS and ICML show the same artifact-type pattern: measurement studies dominate, existing-benchmark critiques are substantial, and genuinely released datasets are rare. This venue-level pattern is consistent with, but does not definitively establish, the paper’s interpretation of reviewer behavior.

  • Venue-level composition: 2/95 NeurIPS and 1/96 ICML submissions were genuine released-dataset papers, making new primary artifacts rare at both venues.The apparent released-model case did not release model weights.
  • Venue-level composition: Existing_benchmark_critique was the second-most common artifact mode at both venues: 19/95 NeurIPS and 25/96 ICML.The compositional signal is therefore not an artifact of aggregating venues.
  • Classification and robustness: Table 4 retained three released_dataset papers after manual reassignment of one ICML released_dataset to measurement_study and one ICML released_model to proposed_experiment.The reassignment followed manual re-reading, and the venue columns sum to 95 for NeurIPS and 96 for ICML.
  • Classification and robustness: Because artifact_type had 0.60 cross-model agreement and the pattern was post-hoc, its lack of systematic variation supports F3 only as consistent evidence, not a definitive reviewer-behavior claim.Artifact_type had the second-lowest cross-model agreement in validation.

B.3 Acceptance rate by claim type … B.6 Full claim-type × rating pairwise comparisons

Across the audited pool, acceptance and reviewer ratings show no reliable claim-type differences, while the historical reference class differs mainly in the artifacts used to operationalize claims. The reference-class comparison also highlights a compositional contrast between new models or measurements and the 2025 audit’s concentration in measurement studies and existing-benchmark critiques.

  • B.3 Acceptance rate by claim type: 63.6% vs. 59.4%: normative claims had a slightly higher within-pool acceptance rate than reformist claims, but no table column survived significance testing.Acceptance is noisier than ratings, and the predictive and descriptive cells are too small for their ordering to be interpreted as an effect.
  • B.3 Acceptance rate by claim type: The pooled acceptance-rate table covers 191 papers and reports claim-type differences transparently rather than as statistically supported effects.The table is explicitly described as within-pool acceptance by claim type.
  • B.4 Exploratory ordinal Spearman on claim type: ρ = +0.010, p = 0.89, N = 191: normalized reviewer ratings were essentially uncorrelated with ordinal claim type.The raw-pooled ρ = −0.12, p = 0.10 result is attributed to mixing the NeurIPS and ICML rating scales.
  • B.5 Rubric classification of the historical reference class: Eight canonical reference-class papers were reclassified with the same Gemini 3 Pro rubric pipeline used for the 191-paper audit.The set included AlexNet, the Transformer, two adversarial-example papers, Concrete Problems, the Bitter Lesson, and two RLHF papers.
  • B.5 Rubric classification of the historical reference class: Artifact type drives the compositional contrast: reference-class papers often operationalized claims through new models, measurements, or proposed experiments, unlike the audit’s dominant measurement studies and existing-benchmark critiques.Only 3/191 audited papers released datasets, and none released genuine model weights.
  • B.5 Rubric classification of the historical reference class: Seven of eight historical papers were classified as descriptive or reformist, reflecting text-level primary orientation rather than their downstream normative effects.Their normative weight was delivered through capability or measurement.
  • B.6 Full claim-type × rating pairwise comparisons: None of the six normalized claim-type rating comparisons was significant, even before correction, so the reported normative-versus-reformist test was not selectively chosen.The full matrix was included to expose the complete comparison set.
  • B.6 Full claim-type × rating pairwise comparisons: H = 0.19, p = 0.98: the normalized reviewer-rating differences across claim types were not significant in the omnibus test.The table reports six pairwise tests and applies Bonferroni correction by multiplying each raw p by 6.

B.7 Per-field classifier reliability · B.8 Full list of audited papers, by venue and decision

Reliability is strongest for claim-type classification, while artifact-coupling classification is less reliable across models and is therefore used cautiously. The audit covers all 191 submissions, whose venue- and decision-grouped citations are listed without implying venue-level acceptance rates.

  • B.7 Per-field classifier reliability: Table 8 reports Gemini’s three-run self-consistency and Claude Sonnet 4.6 exact-match agreement across all 191 papers.The cross-model figures should be interpreted as upper bounds under the anchoring caveat in Appendix A.4.
  • B.7 Per-field classifier reliability: Claim type reaches 0.90 Gemini self-consistency and 0.85 cross-model agreement, the two fields supporting findings F1 and F4.The cross-model figure is an upper bound because Claude saw Gemini’s labels during reclassification.
  • B.7 Per-field classifier reliability: Artifact coupling reaches 0.85 self-consistency but only 0.50 cross-model agreement, limiting its use to F3’s near-zero correlation rather than fine score comparisons.The near-zero magnitude is described as robust to noise.
  • B.7 Per-field classifier reliability: A blinded single-expert re-audit assigns claim-type labels with 0.81 raw agreement against Gemini’s majority-vote labels.Because only one annotator participated, the result cannot separate rubric ambiguity from individual annotator idiosyncrasy.
  • B.8 Full list of audited papers, by venue and decision: The within-pool “Rejected” column is not a venue-level acceptance-rate statistic because withdrawn rejections are not public on OpenReview.This limitation is stated in the discussion of Sections 2 and 6.
  • B.8 Full list of audited papers, by venue and decision: Table 9 cites all 191 audited submissions grouped by venue and within-pool decision, with compressed citation ranges backed by auto-generated BibTeX entries.Entries use keys of the form or_<paper_id>.
Loading 2608.16894v1…