Source-linked AI summary
AI Hallucinations: A Misnomer Worth Clarifying
Negar Maleki, Balaji Padmanabhan, Kaushik Dutta
TL;DR
AI hallucination lacks consistent usage across the expanding domains where large language models are applied, creating a need to clarify the term and its alternatives. The paper systematically reviews definitions across 14 databases, categorizes them by application, and finds inconsistent and sometimes contradictory characteristics while identifying alternative terminology. It concludes that more systematic, consistent, and semantically nuanced terms are needed.
Problem
AI hallucination lacks a formal, consistent definition, with varying characteristics across applications and concerns about the term’s appropriateness in domains including medicine.
Method
The paper systematically reviewed literature across 14 databases, manually identifying definitions of AI hallucination and grouping them by application.
Results
The review found inconsistent and sometimes contradictory uses of AI hallucination and identified alternative terms proposed in the literature.
Takeaways & Limitations
The findings support using more systematic, consistent, and semantically nuanced terminology for phenomena currently labeled AI hallucinations.
Abstract
from arXiv · showhide
As large language models continue to advance in Artificial Intelligence (AI), text generation systems have been shown to suffer from a problematic phenomenon termed often as "hallucination." However, with AI's increasing presence across various domains including medicine, concerns have arisen regarding the use of the term itself. In this study, we conducted a systematic review to identify papers defining "AI hallucination" across fourteen databases. We present and analyze definitions obtained across all databases, categorize them based on their applications, and extract key points within each category. Our results highlight a lack of consistency in how the term is used, but also help identify several alternative terms in the literature. We discuss implications of these and call for a more unified effort to bring consistency to an important contemporary AI issue that can affect multiple domains significantly.
I. INTRODUCTION
The term “hallucination” shifted from a constructive concept in computer vision to a label for erroneous AI outputs, while lacking a precise, universally accepted definition. The paper motivates a systematic review because the term’s uses vary across domains and may be inappropriate or harmful, especially in medicine.
- In early computer vision, “hallucination” described constructive processes such as super-resolution, image inpainting, and image synthesis.Low-resolution images could be made more useful by generating additional pixels.
- Recent computer-vision research uses “hallucination” for errors including nonexistent-object detection and incorrect localization.
- Language-model hallucination includes fluent outputs unrelated to inputs or content that is incorrect and unsupported by information.
- No precise, universally accepted definition exists, and interpretations can be diverse or contradictory across AI applications.
- Medical perspectives criticize the term because AI lacks sensory perception and the metaphor may stigmatize people with mental illness.
- The review examines definitions across 14 databases and domains beyond healthcare and computer science, including ethical, legal, physical-science, and sports settings.
II. METHODOLOGY
The review searched major databases using adapted queries, manually screened records, and retained English publications defining AI hallucination in large language models. From an initially broad retrieval, the authors identified 333 definition-bearing records while acknowledging that manual filtering may have missed newer meanings.
- The database search covered computer science and health sources, including PubMed, Scopus, Web of Science, ACM, IEEE Xplore, ScienceDirect, Google Scholar, and arXiv.
- Search strategies were adapted to each database’s query behavior and result volume, with screening focused on titles, abstracts, introductions, full text, or body sections as appropriate.
- 333 records provided an AI-hallucination definition independently or through an cited reference after searches covering January 1, 2013, to October 1, 2023.
- The authors manually reviewed papers to identify definitions in context, but this process may have missed definitions with newer connotations.
- Google Scholar searches were narrowed from 17,000 records to 89 using exact phrases, while arXiv retrieved 40 relevant papers through an advanced keyword search.
- Eligibility included published research and preprints containing specified AI-hallucination search terms, while non-English records and non-LLM hallucination types were excluded.
III. RESULT
The review finds that AI hallucination lacks a formal, consistent definition, with characteristics varying—even contradictorily—across applications. It also identifies alternative terms and organizes definitions across diverse application categories.
- The review found no formal, consistent definition of AI hallucination and little agreement on its characteristics across applications.The observed characteristics sometimes directly contradict one another.
- In text translation, hallucination was variously described as fluent but irrelevant, fluent but inadequate, or abnormal and unrelated.
- In text summarization, hallucination denotes content inconsistent with the source document and includes intrinsic and extrinsic subtypes.
- Definitions became more relevant to ChatGPT after its November 30, 2022 launch, while retaining different characteristics across applications.
- Researchers proposed alternative terms because hallucination may give AI unintended human characteristics, but the alternatives also reveal inconsistent terminology.
- The review grouped papers across chatbots, generative AI, health, legal and ethical settings, science, translation, question answering, summarization, and other applications.Definitions were extracted from the grouped papers, with key points extracted using ChatGPT 3.5.
IV. DISCUSSION
The review highlights growing public and research attention to AI hallucination, while finding inconsistent terminology and calling for more systematic alternatives and taxonomy.
- Public and research attention: Preventing hallucinations in chatbot outputs remains a key research goal, although few solutions have emerged so far.The popular press discusses future prevention efforts by major technology companies.
- Application-specific definitions: Table III organizes hallucination definitions by application and shows that they share characteristics while differing in degrees of inaccuracy.The listed applications span chatbots, dialogue, generative AI, academia, health, legal and ethical settings, science, technology, translation, question answering, and summarization.
- Terminological alternatives: The review identifies alternative terms intended to replace hallucination with more systematic, consistent, and semantically nuanced language.The authors present these alternatives as an initial step toward more specific definitions and characteristics.
- Future work: More work is needed to develop a systematic taxonomy that can be widely adopted across AI applications.
APPENDIX REVIEW OF THE "AI HALLUCINATION" DEFINITIONS
The appendix compiles definitions of AI hallucination across applications, revealing recurring themes of unsupported, inaccurate, fabricated, or source-inconsistent generated content.
- Scope of definitions: The appendix catalogs AI hallucination definitions across applications including text summarization, translation, dialogue, natural language generation, health, and large language models.Table V presents definitions alongside their associated studies and application areas.
- Text summarization: In text summarization, hallucination commonly involves content that is unsupported by, inconsistent with, or not directly inferable from the source document.Several definitions distinguish intrinsic hallucinations that contradict or misrepresent the source from extrinsic hallucinations that add unverifiable information.
- Text translation: In text translation, hallucination is variously described as fluent but inadequate output or generated content that is unrelated to the source sentences.
- Recurring characteristics: The reviewed definitions also describe hallucination as fictional content, false statements, irrelevant text, nonexistent entities, or information absent from a source.
- Dialogue and language-model settings: Across dialogue, health, and language-model settings, definitions emphasize factually incorrect, nonsensical, confident, or unfaithful responses.