Source-linked AI summary
Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing
Hyeonchu Park, Gahye Jeong, Bugeun Kim
TL;DR
AI text detectors may respond to polished academic style as well as authorship, but prior comparisons of writer populations are confounded by topic, domain, and writing style. Using paired original and professionally edited manuscripts, this study finds detector-specific score changes associated with editing intensity, raising fairness and reliability concerns.
Problem
Prior writer-population comparisons confound detector responses with topic, domain, writing purpose, and individual writing characteristics.
Method
The study compares human-authored manuscripts with their professionally edited versions and evaluates publicly available detectors using baseline and paired false-positive analyses.
Results
Detector responses varied: some assigned lower AI scores after editing while others assigned higher scores, with score shifts correlating with editing intensity.
Takeaways & Limitations
AI detector outputs should not be treated as definitive evidence of AI use in academic settings and should be considered alongside contextual information.
Takeaways & Limitations
The paired design excludes AI-generated texts receiving comparable editing, so the findings do not establish whether the effects generalize to AI-generated text or differ by text origin.
Abstract
from arXiv · showhide
AI text detectors are increasingly employed in academic settings, but it remains unclear whether their outputs reflect AI authorship itself or broader linguistic features associated with polished academic English. Previous studies have reported high false-positive rates (FPRs) for non-native English writing, but population-level comparisons confound authorship with differences in topic, domain, and writing style. Professional editing provides a useful setting for examining this issue because it changes the linguistic form of manuscripts while preserving authorship and content. We examined 135,389 document pairs from a professional English editing service (2018-2025), comprising non-native manuscripts and their native-edited versions, to assess how editing affects detector responses controlling for content and authorship. For the 13 AI text detectors, FPRs for human-written texts varied widely, from 0.0% to 100.0%. Responses varied across detectors: the same edits increased AI scores in some detectors but decreased them in others. Notably, score changes correlated with the extent of editing. The findings identify professional editing style as a key confounding variable in AI detector outputs, rather than establishing a full separation of text origin from linguistic style, raising concerns about fairness and reliability in academic settings.
1 Introduction
AI detectors can falsely label human academic writing as AI-generated, especially when linguistic style differs across writer populations. This study uses professionally edited document pairs to examine detector sensitivity to style while holding content and authorship constant.
- Non-native English writing is more likely than native English writing to be labeled AI-generated.Prior studies link this disparity to linguistic properties such as fluency and lexical diversity.
- Writer-population comparisons are difficult to interpret because topic, domain, purpose, and author characteristics vary simultaneously.
- 135,389 original–edited manuscript pairs provide a natural experiment holding content, topic, and authorship constant.The study evaluates 13 AI text detectors and examines baseline false-positive rates, editing effects, editing intensity, and temporal differences.
- Professional editing produces detector-specific changes: some detectors assign lower AI scores, while others assign higher scores to the same edited text.Score shifts also correlate with editing intensity, and practical significance appears in only 7 of 13 detectors.
2 Related Work
Prior research documents higher false-positive rates for non-native writing but cannot cleanly distinguish linguistic style from text origin because writer groups differ in content and context. This study addresses that gap with professionally edited document pairs and compares broad detector paradigms.
- Non-native English writers often face higher false-positive rates, associated with lower lexical diversity, simpler syntax, and increased repetition.
- Population comparisons confound linguistic proficiency with subject matter, disciplinary conventions, and individual writing habits.
- Comparing original and professionally edited versions isolates linguistic refinement while keeping content and authorship constant.
- AI text detectors comprise token-statistics, zero-shot, and classifier-based approaches.These families use token distributions, language-model probabilities, or supervised decision boundaries.
- The study asks whether detector outputs primarily reflect generation origin or broader linguistic properties correlated with AI-generated text.
3 Method
The study analyzes professionally edited non-native academic manuscripts as paired observations, measuring detector responses, editing intensity, and temporal differences. Statistical comparisons quantify how linguistic revision relates to false positives and AI scores.
- Editors preserve meaning while revising grammar, clarity, naturalness, word choice, and sentence structure under standardized guidelines.
- 135,389 original–edited document pairs were collected from a professional academic English editing service between 2018 and 2025.The dataset spans more than 40 academic disciplines.
- In a 6,000-pair subsample, median document length changed by 6.5 words, while editing intensity had a median of 196 and mean of 463 edit units per document.These edits substantially affect wording while largely preserving overall content and structure.
- The evaluation covers 13 publicly available detectors spanning token-statistics, zero-shot, and classifier-based paradigms.Detectors were tested off the shelf without dataset-specific fine-tuning.
- Pre-ChatGPT documents serve as human-written baselines for false-positive analysis, with FPR, accuracy, F1-score, and paired ΔFP reported.
- Paired t-tests, Cohen’s d, bootstrap confidence intervals, correlations, regression, quintile analyses, and corrected temporal comparisons quantify detector responses.Score shifts are examined separately for pre-ChatGPT and post-ChatGPT periods.
4 Results
Across detectors, false-positive behavior varied substantially, and professional editing shifted scores in opposite directions. Greater editing generally amplified each detector’s existing response pattern, although linguistic-feature associations were correlational.
- 4.1 False-Positive Results Vary by Detector: FPRs ranged from less than 1% to nearly 100% on the same human-written academic texts.Token-statistics detectors often misclassified most documents, whereas zero-shot likelihood-based detectors remained near zero.
- 4.2 Effects of Professional Editing: Professional editing significantly shifted scores across nearly all detectors, but the direction and magnitude differed by system.MAGE decreased by Δscore = −0.108 and ΔFP = −10.7 pp, whereas Fast-DetectGPT and LastDE+ increased false-positive rates by +4.6 and +5.8 pp.
- 4.2 Effects of Professional Editing: The same revisions produced contradictory responses across detector families, with classifier-based systems tending toward human-written classifications and many token-statistics systems toward AI-generated classifications.These patterns were largely consistent across pre- and post-ChatGPT subsets.
- 4.3 Responses Scale with Editing Intensity: Editing intensity amplified existing detector patterns: MAGE’s mean score shift ranged from −0.032 in Q1 to −0.214 in Q5, while Entropy showed ρ = +0.343.The explanatory power of editing intensity was modest, with R2 ≤ 0.041.
- 4.4 Linguistic Drivers of Detector Score Shifts: Edit-token ratio correlated positively with score shifts for statistical detectors (ρ = +0.286) but negatively for supervised detectors (ρ = −0.089).Changes in function-word ratio correlated positively across detector families, whereas type–token-ratio changes correlated negatively; these associations were stronger than those involving word or sentence length.
- 4.4 Linguistic Drivers of Detector Score Shifts: The linguistic-feature interpretations are correlational and do not establish causal effects for any individual feature.Score shifts may reflect combined effects of lexical diversity, syntax, discourse organization, predictability, and other stylistic properties.
5 Conclusion
Using paired original and edited manuscripts, the study finds that detectors disagree widely and respond differently to the same linguistic revisions. It concludes that detector outputs combine generation-origin and linguistic-quality signals, so scores should not be treated as definitive evidence of AI use.
- 5 Conclusion: The study analyzed 135,389 original–edited document pairs to isolate linguistic changes while controlling for topic, content, and authorship.The paired design compared pre- and post-editing versions from an academic editing platform.
- 5 Conclusion: False-positive rates and editing effects varied widely across detectors, with some scores decreasing and others increasing for the same edits.Detector responses also correlated with editing intensity, linking score changes to linguistic revision rather than random variation.
- 5 Conclusion: The findings suggest detectors combine generation-origin and linguistic-quality signals rather than measuring AI-generated content uniformly.Token-statistics detectors had high false-positive rates, while likelihood-ratio-based zero-shot detectors showed lower rates and greater robustness to editing.
- 5 Conclusion: AI detector outputs should not be treated as definitive evidence of AI use, especially in academic settings where editing is common.The study recommends considering detector scores alongside contextual information and evaluating systems under real-world writing conditions.
6 Limitations
The study’s limitations constrain mechanistic interpretation, generalizability, separation of text origin from style, absolute FPR interpretation, and coverage of evolving detectors.
- Mechanistic interpretation: The linguistic analyses identify revisions associated with detector-score shifts but do not fully explain the mechanisms underlying detector behavior.Score changes may reflect combined effects of lexical diversity, syntactic complexity, discourse organization, predictability, and other stylistic properties; no single feature is shown to causally determine outputs.
- Scope: The dataset predominantly contains academic manuscripts by non-native English-speaking researchers, so findings may not generalize to other writing domains.Detector behavior may differ for student essays, journalism, creative writing, social media, and other genres.
- Design boundary: The paired design examines human-authored manuscripts and edited versions but does not fully separate text origin from writing style.Because AI-generated controls with comparable editing are absent, the study cannot determine whether effects generalize to AI-generated text or differ by origin.
- Measurement: Absolute false-positive rates depend on detector-specific thresholding and normalization procedures, although relative editing-condition patterns are more robust.Different threshold choices may produce different absolute FPR estimates.
- Detector coverage: The evaluation covers 13 publicly available detectors assessed off-the-shelf, not future systems or proprietary commercial detectors.The findings should therefore be interpreted as evidence of a broader confounding phenomenon rather than definitive assessments of individual detectors.
- Implication: Despite these limitations, the study provides large-scale evidence that professional editing and stylistic refinement should be considered when evaluating detector fairness and reliability.The evidence comes from more than 135,000 paired documents in a real-world academic editing environment.
The Use of Large Language Models
The manuscript reports using AI-assistance tools during writing to polish language and improve clarity, alongside standard computational tools and hardware for the experiments.
- AI assistance: Claude-Sonnet 4.6 and Grammarly were used to polish language and improve clarity of expression.
- Software: All experiments were conducted with Python 3.10, statsmodels 0.14, and scipy 1.11.
- Hardware: The experimental platform used an AMD Ryzen Threadripper 3960X processor and four NVIDIA RTX A6000 GPUs.The GPUs supported detector models requiring GPU acceleration.
B Dataset
The dataset contains paired original and edited documents, records extensive metadata and editing information, spans multiple periods and domains, and evaluates diverse detector paradigms and edit categories.
- Metadata and preprocessing: Each pair records text versions, service type, submission year, English variant, subject domain, and editing ratio.The editing ratio is the proportion of modified tokens relative to tokens in the original; 417 cases above 5 were excluded, and 5,916 unmodified orders were handled separately.
- Temporal composition: 104,752 pairs (77.4%) belong to the pre-ChatGPT period, compared with 30,637 (22.6%) in the post-ChatGPT period.
- Service and domain composition: Academic Editing contributes 71,595 documents (52.9%), followed by Admission Editing with 50,971 (37.6%).Application Essays, Medicine, and Engineering & Technology are the three largest subject domains and together comprise approximately 36.9%.
- Edit composition: Mixed edits dominate the representative subsample at 66.1%, followed by grammar edits at 15.8% and lexical edits at 14.0%.Purely syntactic edits are rare at 0.1%; score shifts differ significantly by edit category across all 13 detectors, and RoBERTa shows dz = +0.225 for syntactic versus dz = −0.023 for grammar edits.
- Detector coverage: The study evaluates 13 publicly available detectors spanning token-statistics-based, zero-shot, and classifier-based paradigms.The methodological diversity tests whether editing sensitivity is associated with detector paradigm rather than an individual detector.
C.0.1 Thresholding and Normalization Protocol
The protocol converts heterogeneous detector outputs to a common AI-likelihood scale, applies one fixed threshold, and distinguishes threshold-sensitive absolute FPRs from more robust paired changes.
- Normalization: Raw detector outputs are converted to a common [0, 1] AI-likelihood scale using each detector’s original or commonly used normalization.
- Thresholding: A fixed decision threshold of 0.5 is applied across all detectors and both document versions.Per-detector threshold optimization is intentionally omitted so detector behavior is compared under one consistent operating point.
- Change measures: ∆score denotes the edited-minus-original normalized AI-likelihood score, while ∆FP denotes the corresponding change in binary decisions under the 0.5 threshold.
- Interpretation: Paired ∆score and ∆FP estimates are invariant to the specific threshold, whereas absolute FPR values are not.This distinction motivates interpreting absolute FPRs in light of the adopted thresholding protocol and relative editing-condition patterns as more robust evidence.
D.1 Detector Sensitivity Weakens in the post-ChatGPT Era
Professional editing had weaker effects on many detector outputs after 2023, while baseline false-positive scores for some human-written texts increased. These patterns indicate that detector behavior changes alongside writing practices over time.
- Temporal changes in editing effects: 76%: MAGE’s average editing-induced score shift decreased from -0.130 before ChatGPT to -0.031 after ChatGPT.The comparison covers the pre-ChatGPT period (2018–2022) and post-ChatGPT period (2023–2025).
- Temporal changes in editing effects: Editing effects also weakened for RADAR, Fast-DetectGPT, and LastDE+ after 2023.Detectors with previously positive editing responses showed smaller score increases in the post-ChatGPT period.
- Temporal changes in baseline false positives: RoBERTa’s human-text FPR increased from 92.7% to 96.4% between the pre-ChatGPT and post-ChatGPT periods.DetectLLM-LRR also showed a significant increase in baseline scores.
- Interpretation: Detector behavior evolves with broader writing practices, reducing the incremental effect of professional editing over time.The distinction between professionally edited human writing and detectors’ internal AIGT representations may become less pronounced as polished language becomes more common.
D.2 Results on the Strong Editing Subset
Analyses of heavily edited documents amplified the editing-response patterns observed in the full dataset. Some detectors increasingly treated edited texts as human-written, whereas others treated them as more AI-generated, with saturation limiting changes for detectors already producing very high false-positive rates.
- Subset definition and analysis: 57,207 document pairs with edit ratio redit > 0.20 formed the heavy-editing subset.For each detector, the analysis compared pre- and post-editing FPRs, average score shifts, and Cohen’s d with the full dataset.
- Stronger negative editing effects: BiScope’s Cohen’s d changed from −0.488 in the full dataset to −0.833 under heavy editing.MAGE and RADAR showed similar increases in effect magnitude.
- Stronger negative editing effects: MAGE’s FPR decreased from 25.3% to 9.4% under heavy editing.This pattern indicates that MAGE increasingly classified heavily edited texts as human-written.
- Stronger positive editing effects: Entropy’s effect size increased from d = 0.444 to d = 0.707, while GLTR’s increased from d = 0.335 to d = 0.576.DiVeye and DetectLLM-LRR also showed larger positive editing effects under heavy editing.
- Interpretation: Extensive linguistic revisions can make human-written texts appear more AI-generated to certain detectors.The direction of change remained consistent with the full-dataset patterns for most detectors exhibiting editing effects.
- Saturation effects: Log-Likelihood and Log-Rank changed little because their pre-editing FPRs already exceeded 99%.Their high baseline left limited room for additional linguistic modifications to influence detector outputs.