Source-linked AI summary
Tortured phrases: A dubious writing style emerging in science. Evidence of critical issues affecting established journals
Guillaume Cabanac, Cyril Labbé, Alexander Magazinov
TL;DR
Synthetic scientific writing can evade conventional scrutiny as advanced language models generate increasingly human-like text, raising concerns for research-literature integrity. The paper introduces tortured phrases, searches for them across the literature, examines one journal with concentrated occurrences, and combines detector results with analyses of editorial irregularities and dubious articles. It reports likely questionable operational changes, synthetic texts, and an estimate of around 500 questionable articles, while calling for better-characterized detection and analysis of peer-review failures.
Problem
Advanced language models can generate human-like scientific texts, creating a research-integrity concern that conventional detection and review may not adequately address.
Method
The study searches literature for tortured phrases, examines Microprocessors and Microsystems, applies synthetic-text detection to abstracts and controls, and analyzes timelines and individual articles.
Results
The journal showed likely questionable operational changes, acceptance of evidently synthetic texts, and an estimated 500 questionable articles, while other venues also contained tortured phrases and high GPT detector scores.
Takeaways & Limitations
The findings support broader investigation and awareness of questionable AI-generated or rewritten scientific texts that pass peer review.
Takeaways & Limitations
The authors state that detection methods should be characterized for false positives and false negatives and provide a rationale for their decisions.
Abstract
from arXiv · showhide
Probabilistic text generators have been used to produce fake scientific papers for more than a decade. Such nonsensical papers are easily detected by both human and machine. Now more complex AI-powered generation techniques produce texts indistinguishable from that of humans and the generation of scientific texts from a few keywords has been documented. Our study introduces the concept of tortured phrases: unexpected weird phrases in lieu of established ones, such as 'counterfeit consciousness' instead of 'artificial intelligence.' We combed the literature for tortured phrases and study one reputable journal where these concentrated en masse. Hypothesising the use of advanced language models we ran a detector on the abstracts of recent articles of this journal and on several control sets. The pairwise comparisons reveal a concentration of abstracts flagged as 'synthetic' in the journal. We also highlight irregularities in its operation, such as abrupt changes in editorial timelines. We substantiate our call for investigation by analysing several individual dubious articles, stressing questionable features: tortured writing style, citation of non-existent literature, and unacknowledged image reuse. Surprisingly, some websites offer to rewrite texts for free, generating gobbledegook full of tortured phrases. We believe some authors used rewritten texts to pad their manuscripts. We wish to raise the awareness on publications containing such questionable AI-generated or rewritten texts that passed (poor) peer review. Deception with synthetic texts threatens the integrity of the scientific literature.
1 Introduction
The paper situates synthetic scientific writing within a history of computer-generated and nonsensical publishing stings, while emphasizing that modern language models may threaten literature integrity. It investigates a reputable journal using tortured phrases, possible AI-generated abstracts, questionable texts and images, and editorial irregularities.
- Scholarly publishing stings have exposed dysfunctional peer review through human-written and computer-generated nonsensical papers.
- Modern neural-network language models can generate scientific writing and may threaten the integrity of the scientific literature.
- The study reports tortured phrases, possible AI-generated abstracts, questionable texts and images, and shortened reception-to-acceptance intervals in one reputable journal.
- The investigation focuses on Microprocessors and Microsystems, then examines editorial timelines, synthetic-text detection, and factual evidence of inappropriate practices.
2 Tortured phrases found in published academic articles
The paper defines tortured phrases as unconventional substitutions for established scientific terminology and searches the literature for them. It finds that Microprocessors and Microsystems ranked first among venues with matching articles and therefore became the study’s focus.
- Tortured phrases incorrectly replace well-established scientific terms with unconventional phrases, often through word-by-word synonym substitution.
- The authors identified phrases first by chance, expanded them through snowballing, and retro-engineered the wording readers would expect.
- On May 25, 2021, the authors queried Dimensions’ full-text index for papers containing known tortured phrases.
- Microprocessors and Microsystems ranked first among venues listed by Dimensions by decreasing number of matching articles and was selected for further investigation.
- Figure 1 summarizes articles retrieved using 30 tortured phrases listed as of May 25, 2021.
3 The Microprocessors and Microsystems journal
The paper characterizes Microprocessors and Microsystems as an Elsevier computer-science journal and examines its publication record from February 2018 to June 2021. It documents a radical increase in articles per volume beginning in 2020 and constructs a dataset from Crossref and Elsevier metadata.
- Microprocessors and Microsystems is an Elsevier journal classified by Scopus across four Computer Science subject areas.
- The journal is ranked Q3 by Scimago across four subject areas and appears in Clarivate’s Science Citation Index Expanded under three categories.
- The Journal Citation Reports covered 2017–2019, when the journal published 378 articles from its top contributing countries and organizations.
- The journal’s Journal Impact Factor increased from 0.471 to 1.161 over 2015–2019, a 146% increase over four years.
- The analysis covers February 2018 to June 2021 and finds a radical change in articles published per volume beginning in 2020.
- The authors collected DOIs for volumes 56–83 through Crossref and extracted publication metadata from full-text XML obtained through the Elsevier API.
- After filtering out non-full-length publications and removing two articles with missing acceptance dates, the final dataset contained 1,078 articles.
4 Irregularities of the editorial assessment in Microprocessors and Microsystems
The journal’s editorial-assessment timelines changed sharply in early 2021, with shorter processing times, clustered dates, and an over-representation of China-affiliated papers among rapid acceptances. These irregularities motivate investigation, although the study does not provide a satisfactory normal-editorial explanation.
- Editorial assessment covers the interval from manuscript submission to acceptance, including screening, reviewer invitation, peer-review rounds, and the final decision.
- Shortening duration of editorial assessment: 5-fold lower average and 6-fold lower median processing times were observed in early 2021 than in volumes from 2018–2020.Processing times below 40 days became prevalent from volume 80 onward.
- Quicker editorial assessment and over-representation of some author countries: 97.5% of the 404 papers accepted within 30 days had authors affiliated with mainland China, compared with 9.5% among papers processed for more than 40 days.The authors describe this tenfold imbalance as suggesting differentiated processing characterized by shorter peer-review duration.
- Quicker editorial assessment and over-representation of some author countries: 186% more papers were accepted in early 2021 than in early 2020, while median editorial assessment fell from 108 to 25 days.The comparison covers 499 papers in volumes 80–83 and 174 papers in volumes 74–77.
- Blocks of similar editorial timelines: The study identified 111 overlapping blocks containing at least 10 papers and 40 blocks containing at least 20 papers with similar submission, revision, and acceptance dates.A block allows each date to vary by at most one day from a defining triple.
- Blocks of similar editorial timelines: One 30-paper date block combined papers from three special issues and four regular papers, while another 24-paper block combined a special issue with regular papers.The four special issues represented in these blocks were not listed online and had no identifiable prefaces.
- Discussion: The authors could not propose a satisfactory explanation for the date-block phenomenon within the bounds of a normal editorial process.They note that shortened submission-to-acceptance times may reflect poor or deficient editorial assessment.
5 Abstracts with high Generative Pre-Training (GPT) detector score
The study compares GPT detector scores for abstracts from Microprocessors and Microsystems with several control sets, finding an unusually high concentration of flagged abstracts in the experimental journal.
- Datasets for evaluation: The controls represented least-concerning journal articles, high-quality SIAM articles, translated abstracts, Chinese abstracts translated into English, and 139,236 randomly selected Elsevier abstracts.These sets were designed to compare prevalence and assess explanations other than GPT, including translation or cross-language plagiarism.
- Evaluation method: The analysis constructed empirical distributions and Dvoretzky–Kiefer–Wolfowitz confidence bands for the experimental and five control sets.The method uses the GPT detector scores for each abstract and accounts for multiple comparisons.
- Results and analysis: At cut-offs from 0.3 through 0.9, experimental-set scores exceeded every control set significantly more often.At cut-offs 0.1 and 0.2, the distinction held for every control except the Chinese translated set D, for which the comparison was undecided.
- Results and analysis: 72.1% of Microprocessors and Microsystems abstracts had GPT detector scores of at least 70%, compared with a maximum of 13.6% in the other tabulated journals.The authors caution that a high score does not necessarily indicate flaws in an individual paper, but concentrated high scores warrant further assessment.
- Results and analysis: Visual examination of flagged publications found tortured writing, plagiarism, and image theft alongside the detector results.The authors therefore treated automatic screening as a basis for further investigation rather than definitive proof.
6 Critical flaws found in questionable and problematic publications: individual cases
The case analyses identify multiple concrete irregularities in questionable publications, including tortured language, reused images, nonexistent references, incoherent technical descriptions, and inconsistent internal content.
- Case 1: Case 1 combines 99.98% and 88.22% GPT detector scores with difficult-to-understand prose and a water-leak-detector description whose figure and paragraph are largely irrelevant to each other.The source text and image were heavily modified, making the resulting technical explanation hard to understand.
- Case 1: Case 1’s source passage received a 3.53% GPT detector score, contrasting with the heavily altered version’s 88.22% score.The comparison accompanies the authors’ discussion of rewritten technical text.
- Image reuse: Case 1 reused images without acknowledgement, including a figure identical to a 2015 article’s figure and another image taken from a water-leak-detector webpage.The paper’s figure and caption were described as mostly irrelevant to each other.
- Cases 2.1 and 2.2: Case 2.2 used tortured phrases such as “irregular timberland” for random forest and “innocent Bayes” for naïve Bayes.The cases also contained questionable technical prose and shared an image with an unexpected Spanish annotation.
- Case 3: Case 3 reused a heart-rate-monitor image while citing an unrelated breast-cancer-segmentation reference.The figure contained a source indication and was probably reused from another webpage.
- Cases 4 and 5: Other cases contained tortured phrases, nonexistent or unidentifiable references, inconsistent author attribution, nonexistent theorems, unexplained variables, and unacknowledged image reuse.One paper’s stated research interests in Chinese calligraphy and fine arts also appeared inconsistent with its IoT and VR title.
7 Potential sources of problematic papers
The authors discuss paper mills and Spinbot-like software as suspected sources of problematic papers, based on recurring manuscript and production features and the availability of automatic synonym substitution.
- Suspected sources: The authors suspect paper mills produced part of the problematic papers and identify recurring features suggesting that several may come from a single source.This is presented as a suspicion rather than definitive attribution.
- Recurring paper features: Many questionable papers shared a five-section composition that was uncommon before volume 80 or among papers with longer editorial assessment.The recurring structure included Introduction, Related work, Materials and methods, Results and discussion, and Conclusion.
- Recurring paper features: Most inspected questionable papers used a similar light-blue, orange, grey, yellow, and blue diagram palette, suggesting common preparation software.The authors qualify that this pattern was not common to every questionable paper.
- Recurring paper features: Block diagrams and low variability in image and table preparation were common, leading the authors to expect non-standard images may indicate unacknowledged reuse.This expectation is framed as an observation-based suspicion.
- Editorial operation: The journal’s increased output and other operational changes resembled features reported for hijacked journals.The authors describe this as a resemblance, not proof that the journal was hijacked.
- Spinbot-like software: Spinbot offered free rewriting of up to 10,000 characters, replacing words with synonyms such as “big data” with “enormous data” or “huge data.”The paper reports that the service provided no information about its technology or computer code.
8 Conclusion and Call for Action
The study identifies concentrated signs of questionable publication activity, including tortured phrases, synthetic texts, irregular editorial patterns, and potentially hundreds of questionable articles. It calls for monitoring, case-by-case analysis, and impartial investigation while stressing that the evidence is not definitive proof of misconduct.
- Findings: Tortured phrases were found mainly in Computer Science, and multiple venues also published papers with tortured phrases and abstracts with high GPT detector scores.Preliminary probes suggested that several thousands of papers with tortured phrases were indexed in major databases.
- Findings: Microprocessors and Microsystems showed abrupt editorial changes, including shorter assessment times, more accepted articles, synthetic texts, and unusual author backgrounds.Several blocks of articles shared identical submission and acceptance dates, departing from the journal’s typical output before 2021.
- Findings: The authors examined 7 cases covering 8 papers, while none of the journal’s 1,078 papers had received PubPeer comments as of 17 June 2021.The authors interpreted this absence as suggesting that the issues they found had gone unnoticed.
- Findings: Around 500 questionable articles were estimated for Microprocessors and Microsystems, including 389 papers with short editorial assessments and 225 additional articles queued in press.No systematic screening of papers containing tortured phrases had been performed at the time of the estimate.
- Call for Action: The authors recommend monitoring publication activity, followed by individual-paper analysis because hints alone do not establish misconduct.They encourage broader case-by-case studies and ask relevant organizations to initiate an impartial, efficient, transparent, and wide investigation if concerns are grounded.
- Call for Action: They argue that detection methods should be evaluated for false positives, false negatives, and explainability, while screening may provoke an arms race.The paper also attributes the underlying problems to the publish-or-perish atmosphere affecting authors and publishers.
- Call for Action: Elsevier issued Expressions of Concern for six special issues of Microprocessors and Microsystems, although the status of regular papers remained unclear.The update was reported on 12 July 2021.
Appendix: List of supplementary materials
The supplementary materials were released on Zenodo to support transparency and reproducibility, alongside acknowledgements of contributors and data access.
- Supplementary Materials: Supplementary materials were released on Zenodo under DOI 10.5281/zenodo.5031935 for transparency and reproducibility.
- Acknowledgements: The acknowledgements thank PubPeer, Digital Science, and individuals who supported reconnaissance of the subject.