Source-linked AI summary

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou

arXiv:2608.08975v1cs.CLcs.AI

TL;DR

The paper asks which rhetorical choices change AI-review scores when underlying scientific content is preserved, addressing limited controlled evidence separating presentation from scientific merit. The authors construct 4,200 full-paper manuscripts from 120 anonymized ICLR 2026 submissions, applying opposing rewrites across six rhetorical dimensions. Evidence framing and novelty stance produce the largest score contrasts, while rhetorical effects vary with initial reviewer scores, rewriting configurations, reviewers, rewriters, and protocols.

  • Problem

    The paper asks which rhetorical choices change AI-review scores when underlying scientific content is preserved, addressing limited controlled evidence separating presentation from scientific merit.

  • Method

    The authors construct 4,200 full-paper manuscripts from 120 anonymized ICLR 2026 submissions, applying opposing rewrites across six rhetorical dimensions.

  • Results

    Evidence framing and novelty stance produce the largest score contrasts, while rhetorical effects vary with initial reviewer scores, rewriting configurations, reviewers, rewriters, and protocols.

  • Takeaways & Limitations

    AI scientific review should be evaluated for rhetorical robustness across multiple models and conditions rather than relying solely on aggregate agreement or average scores.

  • Takeaways & Limitations

    The benchmark is restricted to ICLR 2026 submissions with recoverable full-paper sources and public review metadata, limiting generalization to other venues and scientific fields.

Abstract

from arXiv · show

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.

1 Introduction

The study tests whether content-preserving rhetorical rewrites can reward-hack AI scientific reviewers by changing their judgments. A controlled corpus and multi-reviewer evaluation show that rhetorical influence is selective, reviewer-dependent, and configuration-dependent rather than universal.

  • Introduction: Because AI systems may both revise manuscripts and evaluate them, rhetorical choices can become reviewer-dependent merit signals, but their influence is selective, configuration-dependent, and sometimes counterproductive.The study addresses this concern amid broader evidence that LLM judgments can vary with presentation, evaluation design, and repeated-judgment inconsistency.
  • Introduction: The framework separates rhetorical presentation from scientific merit using 4,200 full-paper manuscripts, including 120 originals and 4,080 content-preserving rewrites across six rhetorical dimensions.It combines controlled rewriting, paired analysis, multi-model AI review, and standard or strict protocols.
  • Introduction: Evidence framing and novelty stance produce the largest rewrite effects, followed by scope framing, with positive evidence framing raising OA by up to 0.93 and negative novelty stance lowering it by up to 0.73.Evidence framing also changes weak-accept probability by 13 percentage points on average, and the hierarchy remains stable across human-assessed quality levels.
  • Introduction: More elaborate rewriting produces configuration-dependent and diminishing gains: joint rewriting helps substantially with Opus 4.8 but nearly not at all with GPT-5.5, while reviewer guidance is inconsistent.Gains diminish after the second pass.
  • Introduction: Rewriters mainly determine separation between positive and negative variants, whereas reviewers determine the magnitude and sign of OA effects and whether changes appear in contribution or soundness.Strict review shifts the absolute OA scale without consistently strengthening or weakening rewrite effects.

2 Related Work

Prior work finds that LLM-based evaluation is sensitive to prompts, presentation, sampling, and evaluator identity, while scientific-review systems often struggle to identify substantive weaknesses and distinguish paper quality.

  • 2 Related Work: LLM evaluations remain sensitive to prompts, presentation, sampling, and evaluator identity, and scientific-review systems can struggle to identify substantive weaknesses and distinguish paper quality.Rubric-based methods may align with human preferences, but manuscript-side studies also examine hidden instructions and other factors affecting evaluation.

3 A Controlled Analysis Framework for Rhetorical Rewriting in AI Review

The framework tests content-preserving rhetorical variation through verified manuscript baselines, bidirectional interventions across six dimensions, and blinded AI-review comparisons. It covers single-dimension, joint, recursive, and reviewer-guided rewriting under standard and strict evaluation protocols.

  • Corpus construction: The corpus samples 120 ICLR 2026 papers evenly across six human overall-assessment ranges to cover a broad quality spectrum.The seed set contains 20 papers from each range, spanning [1, 3) through [7, 8.5].
  • Rhetorical intervention space: Six rhetorical dimensions are rewritten independently in positive and negative directions from the same anonymized baseline, separating directional sensitivity from rewriting-induced score movement.The dimensions cover novelty stance, scope, evidence framing, contribution structure, technical register, and lexical or syntactic complexity.
  • Framework-level effects: Across the intervention space, evidence framing and novelty stance produce the largest overall-assessment changes, followed by scope framing, with effects varying by rewriter, reviewer, and protocol.Table 3 reports paper-level mean changes for positive and negative interventions relative to prompt-matched originals.
  • Single-dimension rewriting: The single-dimension design produces 12 variants per paper and rewriter, totaling 2,880 rewritten manuscripts without stacking interventions.Each paper receives both directions for all six dimensions from the same anonymized baseline.
  • Extended rewriting workflows: The framework extends single-pass rewriting with joint, three-round recursive, and reviewer-guided workflows that add coordinated objectives or review feedback.Joint rewriting combines all six positive directions; recursive rewriting repeats that procedure for three rounds, while reviewer-guided rewriting performs a second pass using a standard review from the rewriting backbone.
  • AI-review evaluation: Blinded evaluation uses five AI reviewers under standard and strict prompts, with the strict protocol requiring concrete evidence for high ratings and discouraging presentation-based rewards.Reviewers are not told the rewrite condition or rewrite model, and judgments are compared within the same paper against the prompt-matched original.

4 Which Rhetorical Dimensions Matter Most?

Rhetorical sensitivity concentrates in evidence framing and novelty stance, with scope framing forming a weaker second tier. Effects persist across human-assessed quality levels, while the AI reviewer’s starting score chiefly determines whether scores rise or fall.

  • 4 Which Rhetorical Dimensions Matter Most?: Evidence framing and novelty stance produce the largest and most consistent overall-assessment contrasts, followed by weaker effects from scope framing.Positive evidence and novelty variants are favored across displayed configurations, while scope follows the same direction.
  • 4 Which Rhetorical Dimensions Matter Most?: 13.0, 12.0, and 9.0 percentage points are the weak-accept contrasts for evidence framing, novelty stance, and scope framing, respectively.Novelty is slightly stronger under standard evaluation, whereas evidence is stronger under strict evaluation; the pattern reflects selective sensitivity to merit framing rather than a general reward for stronger rhetoric.
  • 4 Which Rhetorical Dimensions Matter Most?: The dimensional hierarchy persists across human-assessed quality levels, without a consistent increase or decrease in rewrite effects as human rating rises.Evidence framing, novelty stance, and scope framing retain the clearest contrasts, and each shows the intended ordering for more than 100 of 120 manuscripts.
  • 4 Which Rhetorical Dimensions Matter Most?: +1.05 OA for original scores in [1, 3] versus −0.19 for scores in [8, 10] shows that low starting scores tend to rise while high starting scores tend to fall.This movement may partly reflect scale bounds and regression to the mean, can outweigh the intended rewrite direction at extremes, and varies across reviewers.

5 Do More Complex Rewrites Yield Reliable Gains?

More elaborate rewriting does not reliably improve AI-review scores: gains depend on the rewriter–reviewer configuration, are usually concentrated by the second round, and diminish thereafter. Joint rewriting favors Opus 4.8, while reviewer guidance and additional rounds provide no consistent model-agnostic advantage.

  • Joint rewriting: Joint rewrites do not consistently outperform individual favorable objectives: Opus 4.8 gains +0.160 OA over the six-rewrite mean, whereas GPT-5.5 falls −0.074 below it.This cross-batch comparison is descriptive rather than a synergy estimate.
  • Score dependence: Joint-rewrite effects track starting AI-review scores more closely than human-assessed quality, declining from +1.42 for initial scores in [1, 3] to −0.88 for scores in [8, 10].This descriptive pattern may partly reflect scale bounds and regression to the mean.
  • Joint rewriting: Opus 4.8 joint rewriting improves reviewer-averaged OA from +0.289 to +0.396 under standard evaluation and from +0.463 to +0.557 under strict evaluation, unlike GPT-5.5.GPT-5.5 changes only from +0.021 to +0.015 under standard evaluation and from +0.045 to +0.032 under strict evaluation; the overall increase from +0.204 to +0.250 is driven by Opus 4.8.
  • Guided rewriting: Reviewer guidance performs worse than an unguided second pass across all four rewriter-prompt averages, with guided-minus-unguided contrasts ranging from −0.008 to −0.067 for Opus 4.8 and −0.035 to −0.042 for GPT-5.5.Because the workflows use independently generated first-pass rewrites, this comparison is descriptive rather than a causal estimate of feedback.
  • Recursive rewriting: Recursive rewriting yields most additional improvement by round two, with Opus 4.8 reaching +0.410 under standard and +0.624 under strict evaluation, while the third round adds little.GPT-5.5 remains much less responsive, and reviewer-specific trajectories may flatten or reverse earlier gains.

6 How Do Rewriters, Reviewers, and Protocols Shape Rhetorical Effects?

Rewriters primarily determine the separation between rhetorical variants, while reviewers determine score magnitude, direction, and rubric associations. Strict review lowers absolute OA scores but does not consistently alter rewrite sensitivity.

  • Review protocols: Mean OA decreases by 1.36 points under strict review, from 6.296 to 4.934, while paper rankings remain broadly stable.95.0% of papers receive lower average scores under strict evaluation, and the mean standard–strict Spearman correlation is .862.
  • Review protocols: Strict review changes mean rewrite effects only from +.003 to +.020 and does not consistently strengthen or weaken rhetorical sensitivity.The mean cross-protocol Spearman correlation for paper-level rewrite effects is only .081, with direction varying across reviewers.
  • Rewriters and reviewers: Rewriters shape rhetorical separation, whereas reviewers determine the magnitude, sign, and rubric dimensions associated with score changes.Reviewers generally favor positive novelty, evidence, and scope variants but disagree on how strongly those changes affect scores.
  • Reviewer heterogeneity: Reviewers map similar OA changes onto different merit dimensions: contribution and soundness usually matter more than presentation, with reviewer-specific associations.Qwen 3.5 F aligns OA changes mainly with contribution, while strict Gemini 3.5 FL emphasizes contribution and soundness; Sonnet 5 is weaker and less stable.

7 Conclusion

AI scientific review is systematically sensitive to rhetorical presentation even when reported scientific content is preserved. This sensitivity varies with rhetorical dimension, rewriting process, model identities, and review protocol rather than reflecting a single stable average effect.

  • AI scientific review responds systematically to rhetorical presentation even when the reported scientific content is preserved.
  • Rhetorical sensitivity varies across dimensions, rewriting processes, rewriter and reviewer models, and review protocols.
  • The observed variation cannot be reduced to one average effect or treated as a stable property of a single model configuration.

Limitations

The study’s conclusions may not generalize beyond ICLR 2026 or fields with comparable data, while expensive evaluation incurred non-random missingness that may reflect model safeguards.

  • The benchmark is limited to ICLR 2026 submissions with recoverable full-paper sources and public review metadata.Results may therefore not generalize to other venues or scientific fields.
  • Full-paper rewriting and multi-model evaluation are computationally expensive, constraining the practical scalability of the study.
  • Some planned reviews lacked valid structured records, and this missingness may be non-random because failures appear related to topics triggering model safeguards.

Ethical Considerations · Appendix · A Detailed Related Work

The study uses public sources, removes identity-bearing information before model evaluation, and reports only aggregate results. Its framework is dual use: it can diagnose rhetorical sensitivity but could also tailor manuscripts to AI-reviewer preferences.

  • Ethical Considerations: The study anonymizes manuscripts before API submission, reports aggregate results, and acknowledges that its controlled rewriting framework could be misused to optimize manuscripts for AI reviewers.Removed information includes author names, affiliations, acknowledgments, submission-status statements, and identity-bearing links; individual authors are not assessed.

A.1 LLM-Based Evaluation … C.6 Strict-Prompt Human-OA Range Analysis

Across related work, execution checks, and detailed experiments, the paper shows that AI-review judgments vary with presentation and evaluation conditions, while controlled rhetorical rewriting produces structured, reviewer- and rewriter-dependent effects. The analyses also document review missingness, human-score discrepancies, secondary-score spillovers, weak-accept changes, and heterogeneity across paper quality ranges.

  • A.1 LLM-Based Evaluation; A.2 AI-Assisted Scientific Evaluation: LLMs support reference-free evaluation and scientific reviewing through rubric-conditioned scoring, pairwise comparison, customized evaluators, manuscript prediction, evidence-linked comments, and fluent feedback, but judgment quality depends on the task and evaluation conditions.Prior work reports substantial alignment with human preferences and useful GPT-4 feedback while also motivating scrutiny of reviewing competence and condition sensitivity.
  • A.1 LLM-Based Evaluation; A.3 Rhetorical Rewriting and Presentation Effects: Rhetorical rewriting is motivated by prior evidence that LLM evaluations shift with candidate order, length, prompts, anchors, familiarity, sampling, and model-family identity, while this study tests opposing edits across six dimensions with methods and reported results preserved.The design retains lower-scoring rewrites rather than selecting only favorable candidates, enabling dimension-level comparisons of rating sensitivity.
  • B.1 AI-Review Execution; B.2 Review Missingness and Potential Safeguard Triggers: The evaluation pipeline submitted PDFs through OpenRouter without OCR, browsing, retrieval, or reviewer metadata, and produced 42,396 valid structured reviews from 42,480 planned evaluations.The 84 missing evaluations were retained as missing, with analyses using available matched pairs; Sonnet 5 repeatedly failed on 26 JULI evaluation cells, plausibly reflecting safety mechanisms.
  • B.4 Comparison with Human Overall Assessments: AI and human overall assessments differ by prompt: standard-prompt AI ratings exceed human OA by 0.418–2.934 points, while strict-prompt differences range from below human OA to 1.676 points above it.Under the strict prompt, GPT-5 mini, GPT-5.5, and Sonnet 5 score below human OA, Qwen 3.5 F is nearly aligned, and Gemini 3.5 FL remains 1.676 points higher; paper-level MAE ranges from 1.029 to 2.934 points.
  • C Detailed Results for Individual Rhetorical Dimensions; C.1 Overall-Rating Effects; C.6 Strict-Prompt Human-OA Range Analysis: The appendix decomposes overall-rating changes by fixed reviewer–rewriter pair and human-OA range, retaining original-to-rewrite means, paired changes, confidence intervals, and strict-prompt range-specific effects.Pooled results average reviewer evaluations within paper, while fixed-model and strict-prompt tables preserve configuration-level heterogeneity across the six rhetorical dimensions.
  • C.1 Overall-Rating Effects; C.2 Fixed-Model Directional Heterogeneity; C.4 Paper-Level Directional Consistency: Novelty, evidence, and scope produce the most consistent directional effects, whereas technical register, contribution structure, and lexical complexity are smaller or less stable across fixed reviewer–rewriter pairings.The intended ordering holds for more than 100 of 120 papers for novelty, evidence, and scope; lexical complexity is nearly balanced at 62 wins, 7 ties, and 51 losses.
  • C.3 Secondary AI-Review Score Spillovers; C.5.1 Panel and Configuration Summaries; C.5.2 Full Fixed-Model Rates and Transition Counts: Rhetorical changes spill over into presentation, contribution, and soundness scores, which are treated as reviewer-response changes because methods and reported values remain fixed.The secondary-score analysis pools both rewrite models and all five reviewer models, while weak-accept analyses report configuration-level and panel-aligned probability changes plus upward and downward cutoff crossings.

C.7 AI Original-Score Range Analysis … C.7.5 Detailed Secondary-Score Profiles Across the Rating Scale

The analysis stratifies overall-rating changes by reviewer-matched AI original-score ranges, then provides pooled, boundary-focused, reviewer-specific, and secondary-score profiles. These descriptive decompositions preserve evaluation-condition detail while cautioning that extreme-range movement may reflect scale bounds, regression to the mean, and sparse cells rather than rhetorical steering.

  • C.7 AI Original-Score Range Analysis: Strict-prompt range tables retain fixed rewriter–reviewer rows across four AI original-score ranges, with only within-row averages across the six rhetorical dimensions.The ranges are [1, 3], [4, 5], [6, 7], and [8, 10], assigned using prompt- and reviewer-matched original overall ratings.
  • C.7 AI Original-Score Range Analysis: AI original-score ranges organize descriptive rewrite responses, but extreme-range movement may reflect bounded-scale effects, regression to the mean, and sparse fixed-condition cells.Ranges are assigned separately for each prompt–reviewer pairing, so the same paper can appear in different columns; these profiles are not unbiased estimates of quality moderation.
  • C.7.1 Panel-Aligned Values for the Main-Text Displays: Panel-aligned displays preserve the two rewrite models separately while averaging reviewer-model changes within each paper, protocol, rhetorical condition, and original-score range.Table 23 reports rating changes and Table 24 reports weak-accept probability changes in percentage points, averaging all five reviewers within paper.
  • C.7.2 Pooled Overall-Rating Inference: Pooled inference summarizes original means, rewrite means, signed changes, absolute movement, and paper-bootstrap intervals across all score ranges, dimensions, directions, and protocols.The two rewrite models and five reviewer models are averaged within paper.
  • C.7.3 Pooled Rating-6 Transitions Near the Decision Boundary: Boundary analysis counts rating-6 transitions only in [4,5] and [6,7], where transitions can occur upward or downward, while omitting sparse extreme-range decompositions.Extreme-range panel values remain available, but their complete transition-count breakdown is not reported.
  • C.7.4 Full Reviewer-Model Decomposition: The full reviewer-model decomposition holds both rewrite and reviewer models fixed, reporting original and rewritten values with their changes so reviewer-specific baselines remain visible.This decomposition follows the pooled and boundary summaries without replacing the primary score-range results.
  • C.7.5 Detailed Secondary-Score Profiles Across the Rating Scale: Secondary diagnostics combine presentation, contribution, and soundness responses across overall-rating ranges, but the ranges are defined by original overall rating rather than each secondary score.Table 29 pools both rewrite models and all five reviewer models; these profiles are therefore complementary diagnostics.

D Detailed Results for Multi-Stage Rhetorical Rewriting … E.1 Five-Reviewer Comparison

The appendix details multi-stage rhetorical rewriting across joint, reviewer-guided, and recursive workflows, then compares five reviewers’ response distributions, score changes, and cross-reviewer consistency.

  • D Detailed Results for Multi-Stage Rhetorical Rewriting: The appendix distinguishes joint rewriting, recursive reapplication, and reviewer-guided comparisons, including an independently generated no-review endpoint that is not a matched causal treatment.The no-review reference is a separate two-pass endpoint without intermediate review.
  • D.1 Joint Rewrite by Reviewer Model: Joint-rewrite analyses fix the rewriter, review prompt, and reviewer, while treating comparisons with six single-dimension rewrites as descriptive cross-batch differences.Table 30 reports matched-endpoint effects and a paper-matched mean of the six positive single-dimension rewrites.
  • D.2 Final Reviewer-Guided Endpoints by Reviewer Model: Reviewer-guided endpoint analyses report original and Final levels, paired changes, threshold shifts, and descriptive contrasts with an independently generated no-review two-pass endpoint.The appendix covers all 20 reviewer, prompt, and rewriter configurations under standard and strict review.
  • D.3 Complete Recursive Joint-Rewrite Trajectories: Recursive joint-rewrite analyses track reviewer-pooled overall-rating and weak-accept-probability trajectories across successive rounds under standard and strict review.Reviewer-specific decompositions report cumulative changes from the matched original for each round.
  • E Detailed Reviewer and Secondary-Score Results: The reviewer analyses retain the complete prompt-level calibration comparison while focusing additional results on reviewer differences and cross-reviewer consistency.These analyses extend the reviewer and secondary-score results from Section 6.
  • E.1 Five-Reviewer Comparison: Across five reviewers, Figure 4 presents full response distributions, while Table 37 reports exact mean rewriting-induced OA changes and the range of ten pairwise rank correlations.The compared reviewers are Gemini 3.5 FL, Qwen 3.5 F, GPT-5 mini, GPT-5.5, and Sonnet 5.

E.2 Secondary-Score Associations and Cross-Reviewer Consistency

Secondary-score changes showed reviewer- and prompt-dependent associations with ΔOA, while reviewers weakly agreed on which papers were most affected by rewriting. Record-level within-paper correlations preserved stronger co-movement for soundness and contribution than presentation.

  • Secondary-score associations: Secondary-score associations with ΔOA varied by reviewer and prompt, including ρΔOA=.635 for contribution under GPT-5.5 and ρΔOA=.587 for soundness under Strict Gemini 3.5 FL.The paper-first estimates average the 12 single-dimension rewrites within each paper and use paper-bootstrap intervals.
  • Cross-reviewer consistency: Cross-reviewer agreement was weak for OA and all three secondary scores, even when reviewers produced similarly signed aggregate changes.Table 39 measures agreement using pairwise Spearman correlations between reviewers’ paper-level mean rewrite responses.
  • Record-level analysis: After removing each paper’s mean response, correlations ranged from .210 to .691 for soundness, .020 to .239 for presentation, and .171 to .841 for contribution.These ranges cover 20 rewriter, reviewer, and prompt combinations.

F Stability Under Repeated Review Sampling · G Complete Rewrite and Review Prompt Templates

The appendix shows that single-review estimates closely match three-review means, supporting the primary protocol, and specifies content-preserving prompts for directional, full-paper, joint, and reviewer-guided rewrites. These templates systematically alter rhetorical presentation while preserving the manuscript’s scientific content, evidence, and protected technical structure.

  • F Stability Under Repeated Review Sampling: Single-review estimates closely matched three-review means, with Pearson r = 0.995 and Spearman ρ = 0.993 across 12 baseline-paired effects, and mean absolute difference 0.040 points.The auxiliary audit found Pearson r = 0.996 and Spearman ρ = 0.943 across six positive-versus-negative contrasts; the main analyses therefore used one designated review per cell.
  • G Complete Rewrite and Review Prompt Templates: The review prompts preserve the complete evaluation and structured-output instructions supplied with each PDF, while advanced templates reuse the combined six-dimension prompt across rounds.Only the reviewer-guided final rewrite differs in the advanced workflow templates.
  • G.1 Direction-Specific Rewrite Prompts: The direction-specific prompts change one rhetorical dimension at a time while retaining the paper’s underlying content and scientific accuracy.Dimensions include claim and novelty stance, scope and generalization, quantitative evidence framing, contribution salience, technical register, and lexical and syntactic complexity.
  • G.1 Direction-Specific Rewrite Prompts: Positive and negative instructions reverse rhetorical framing—for example, assertive versus cautious claims and broader versus bounded scope—without changing datasets, tasks, methods, or reported findings.Quantitative-evidence prompts also require preserving underlying numerical facts and comparisons, while technical-register and complexity prompts preserve technical content and meaning.
  • G.2 Shared Full-Paper Rewrite Requirements: Shared requirements mandate paragraph-by-paragraph rewriting across all major sections, including captions and result prose, while prohibiting new experiments, data, baselines, numbers, or unsupported conclusions.Protected LaTeX commands, displayed mathematics, code, algorithms, bibliography files, graphics paths, URLs, labels, and filenames must remain valid.
  • G.3 Joint and Reviewer-Guided Rewrite Prompts: The joint rewrite applies one coordinated six-dimension editorial profile to the original paper, and recursive rewriting reapplies the unchanged instruction to the immediately preceding source.The joint profile combines stronger supported claims, broader justified scope, clearer evidence, more salient contributions, greater technical formality, and sophisticated expression.
  • G.3 Joint and Reviewer-Guided Rewrite Prompts: All rewrite variants require confidence, breadth, and empirical emphasis to remain grounded in existing evidence, without introducing unsupported certainty, domains, capabilities, or generalization claims.The prompts explicitly preserve qualifications that materially define claims and retain the same methods, datasets, tasks, assumptions, evaluation settings, and demonstrated capabilities.
Loading 2608.08975v1…