Source-linked AI summary
FaithDial: A Faithful Benchmark for Information-Seeking Dialogue
Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar Zaiane, Mo Yu, Edoardo M. Ponti, Siva Reddy
TL;DR
Knowledge-grounded dialogue systems can generate fluent but unsupported statements, motivating better resources for faithful response generation and evaluation. FaithDial addresses this gap by editing hallucinated Wizard of Wikipedia responses into knowledge-grounded conversations and using them to train critics and generators. The resulting resource improves faithfulness and abstractiveness in-domain and during zero-shot transfer.
Problem
Knowledge-grounded dialogue models can produce fluent but unverifiable or factually incorrect statements, known as hallucinations.
Method
FaithDial edits hallucinated WoW responses to make them faithful to corresponding knowledge snippets and acknowledge ignorance when necessary.
Results
FaithDial improves faithfulness and abstractiveness in dialogue generation, with benefits verified in-domain and during zero-shot transfer to TopicalChat and CMU-DoG.
Takeaways & Limitations
FaithDial provides a resource for training hallucination critics and dialogue generators while preserving conversational quality.
Takeaways & Limitations
The FAITHDIAL training split is smaller than WoW because of limited budget, and Q2-NLI may not robustly detect partial hallucinations.
Abstract
from arXiv · showhide
The goal of information-seeking dialogue is to respond to seeker queries with natural language utterances that are grounded on knowledge sources. However, dialogue systems often produce unsupported utterances, a phenomenon known as hallucination. To mitigate this behavior, we adopt a data-centric solution and create FaithDial, a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark. We observe that FaithDial is more faithful than WoW while also maintaining engaging conversations. We show that FaithDial can serve as training signal for: i) a hallucination critic, which discriminates whether an utterance is faithful or not, and boosts the performance by 12.8 F1 score on the BEGIN benchmark compared to existing datasets for dialogue coherence; ii) high-quality dialogue generation. We benchmark a series of state-of-the-art models and propose an auxiliary contrastive objective that achieves the highest level of faithfulness and abstractiveness based on several automated metrics. Further, we find that the benefits of FaithDial generalize to zero-shot transfer on other datasets, such as CMU-Dog and TopicalChat. Finally, human evaluation reveals that responses generated by models trained on FaithDial are perceived as more interpretable, cooperative, and engaging.
1 Introduction
Knowledge-grounded dialogue models can produce fluent but unsupported statements, so FaithDial addresses hallucination through data-centric correction of Wizard of Wikipedia responses. The resulting benchmark improves faithfulness while supporting hallucination criticism and dialogue generation.
- Motivation: Hallucinated or factually incorrect statements undermine trustworthy deployment and can cause harm or enable disinformation.The problem is especially consequential in high-stakes domains.
- Data-centric solution: FaithDial edits hallucinated Wizard of Wikipedia responses to align with the available knowledge and acknowledge ignorance when the knowledge is insufficient.The approach preserves conversational cohesiveness while contrasting original and corrected wizard responses.
- Benchmark quality: 94.4% of FaithDial utterances are faithful, compared with 20.9% in WoW.FaithDial contains around 50K turns across 5.5K conversations.
- Dialogue generation: Models trained on FaithDial become more faithful while also improving cooperativeness, creativity, and engagement.The benchmark’s benefits also transfer to CMU-DoG and TopicalChat in zero-shot settings.
- Hallucination criticism: FaithDial provides positive and negative supervision for hallucination critics that achieve state-of-the-art performance on BEGIN in zero-shot evaluation.Positive examples come from FaithDial and negative examples from WoW.
2 FAITHDIAL: Dataset Design
FAITHDIAL defines faithful knowledge-grounded dialogue as utterances entailed by at least one available knowledge item and constructs the dataset by correcting hallucinated or uncooperative WoW responses. Its annotation process preserves conversational coherence while encouraging truthful, natural, and non-copying responses.
- Faithfulness definition: Faithfulness requires that at least one non-empty subset of the available knowledge semantically entail the utterance.An utterance may be grounded in multiple facts, but not in none.
- Dialogue roles: The WIZARD must provide source-attributable information conversationally and acknowledge ignorance when the available knowledge lacks an answer.The bot is restricted to the available knowledge, unlike the more flexible human SEEKER.
- Dataset construction: FAITHDIAL corrects problematic existing dialogue turns rather than creating a benchmark from scratch, preserving useful information and enabling larger-scale annotation.The same dialogue turn can be contrasted in hallucinated and faithful forms.
- Data selection: FAITHDIAL uses WoW because it contains fewer full hallucinations than CMU-DoG and TopicalChat, making correction more practical.The source study reports full hallucinations in 19.7% of WoW, compared with 61.4% in CMU-DoG and 46.8% in TopicalChat.
- Annotation guidelines: Workers remove unsupported information, replace it with supported paraphrases, and avoid copying source segments to retain creativity.When the knowledge cannot answer the seeker, the wizard should acknowledge that limitation while continuing with grounded content.
- Coherence preservation: Edits to a wizard response can require editing the seeker’s next utterance so the revised conversation remains coherent.Workers modify subsequent seeker turns when the original continuation no longer fits the edited wizard response.
3 Dataset Quality
FAITHDIAL’s quality-control process used worker qualification and pilot procedures before final validation of edited responses.
- Crowdworker Quality Control: Workers were screened with a qualification test and a pilot round, with poor-quality annotators removed.They also received corrective guidance when pilot errors were observed.
- Human validation: Three new workers re-edited 500 responses, while three others judged faithfulness by majority vote.The validation also measured agreement on BEGIN and VRM labels.
WoW FAITHDIAL
The WoW-to-FAITHDIAL example shows how unsupported claims are replaced with knowledge-grounded, conversational responses that acknowledge the wizard’s limitations.
- WoW FAITHDIAL: The original exchange includes a partial hallucination label for a response that combines unsupported personal experience with knowledge-grounded content.The seeker’s subsequent turn is marked as incoherent with the freshly edited response.
- WoW FAITHDIAL: The edited response replaces the claim “I absolutely love to surf” with an explicit acknowledgment that the virtual bot cannot surf.It then responds to the shark concern and provides supported information about surfing waves.
- WoW FAITHDIAL: The dialogue example pairs seeker questions with knowledge snippets and annotator labels for response attribution and speech acts.Table 3 identifies hallucinated content and the BEGIN and VRM annotations used during editing.
4 Dataset Analysis
FAITHDIAL substantially improves faithfulness while preserving conversational speech acts and expressing knowledge more abstractively than WoW.
- Dataset statistics: FAITHDIAL contains 5,649 dialogues and 50,761 utterances, with workers editing 84.7% of wizard responses but only 28.1% of seeker responses.The limited seeker editing was intended to preserve conversational cohesiveness.
- 4.2.1 Faithfulness: 94.4% of FAITHDIAL responses were faithful, compared with 20.9% in WOW, whose responses included 71.4% hallucination.The FAITHDIAL estimate comes from human validation, while the WOW figures come from a large-scale audit.
- 4.2.1 Faithfulness: FAITHDIAL uses diverse faithful speech acts, including edification, acknowledgments, follow-up questions, and attributable opinions.In WOW, faithful strategies were mostly limited to edification, which reduced naturalness.
- 4.2.2 Abstractiveness: FAITHDIAL responses have similar coverage but consistently lower density than WOW, indicating greater abstractiveness.Density measures copied-span length, whereas coverage measures the proportion of response words appearing in the source knowledge.
- 4.2.2 Abstractiveness: The dataset’s abstractive strategies include inference, rewording, syntactic reshaping, abridging, and adding connectives.These strategies present knowledge without repeating long source fragments.
- Unanswerable questions: Fallback responses were needed in 48% of sampled conversations, with 33% of wizard responses per such conversation edited on average.Fallbacks addressed personal questions, objective questions, and opinions while keeping conversations moving with source knowledge.
5 Experiments
Experiments evaluate FAITHDIAL for hallucination criticism and dialogue generation, including contrastive training, automated metrics, human judgments, and zero-shot transfer. Across these settings, FAITHDIAL-trained models improve faithfulness while retaining or enhancing abstractiveness and dialogue quality.
- 5.1 Task I: Hallucination Critic: FAITHCRITIC substantially outperforms DNLI and DECODE in zero-shot transfer on MNLI and BEGIN.The comparison uses accuracy results reported in Table 4.
- 5.2 Task II: Dialogue Generation: 42.2% lower hallucination and 4.3% higher Q2-NLI result when T5 trains on FAITHDIAL rather than WOW.The comparison supports the importance of data quality despite FAITHDIAL being one third the size of WOW.
- 5.2.2 Automatic Evaluation: 7% higher BERTScore with nearly unchanged word-overlap F1 indicates FAITHDIAL increases semantic similarity without inducing additional extractiveness.The finding is especially relevant because extractive responses can improve faithfulness at the expense of creativity.
- 5.2.2 Automatic Evaluation: T5-INFONCE achieves 1.4 Critic hallucination and 55.8 F1 extractiveness, combining faithfulness with abstractiveness.The contrastive objective distinguishes faithful responses from hallucinated candidates using perturbed knowledge and annotated hallucinations as negatives.
- 5.2.3 Human Evaluation: Human evaluation finds 32.6% less hallucination alongside higher interpretability, cooperativeness, engagingness, and abstractiveness for FAITHDIAL-trained models.T5 is evaluated because it performs best on the automated metrics; T5-INFONCE reaches 77.4% faithfulness and the highest dialogue-quality scores.
- 5.2.3 Human Evaluation: T5-INFONCE answers unanswerable questions properly in 83.2% of cases versus 33.3% for T5-LOSSTRUNCATION trained on WOW.The result comes from manual evaluation by three annotators with Krippendorff’s alpha of 0.9.
6 Related Work
Prior work addresses hallucination through metrics, causal analyses, consistency or control mechanisms, retrieval, and dialogue-specific benchmarks. FAITHDIAL instead contributes highly faithful curated information-seeking dialogue to counter noise in existing training data.
- Hallucination in Natural Language Generation: Hallucination research spans data-to-text, translation, summarization, question answering, and dialogue, with work targeting detection metrics and possible causes.Reported causes include out-of-domain generalization, noisy training examples, and exposure bias from maximum-likelihood training.
- Hallucination in Dialogue Systems: Dialogue systems have addressed hallucination with control tokens, token-level critics, retrieval modules, loss functions, and consistency constraints.These approaches modify generation or add grounding mechanisms rather than primarily curating the training corpus.
- Hallucination in Dialogue Systems: More than 60% of three popular dialogue benchmarks contain hallucinations, which existing faithfulness-oriented models may reproduce or amplify.FAITHDIAL is presented as the first highly faithful curated dataset for information-seeking dialogue.
- Hallucination Evaluation: BEGIN, DialFact, Conv-FEVER, and AIS provide dialogue knowledge-grounding evaluation settings, while Q2 has prompted debate about reliable automatic hallucination metrics.These benchmarks and metrics frame how hallucination-free dialogue systems are assessed.
7 Conclusions
FAITHDIAL is a curated benchmark built by editing hallucinated and uncooperative WoW responses into faithful information-seeking dialogue. It supports hallucination criticism and generation, with improvements in faithfulness and abstractiveness that extend to zero-shot transfer.
- 7 Conclusions: FAITHDIAL contains manually edited WoW examples in which responses are made faithful to gold knowledge and suitable for conversational information seeking.The edited responses address hallucinated and uncooperative behavior in 79.1% of the original dataset.
- 7 Conclusions: FAITHDIAL trains a hallucination critic that reaches a new state of the art on BEGIN and supports several dialogue generation models.The critic discriminates whether utterances are faithful to the available knowledge.
- 7 Conclusions: Intermediate fine-tuning on WOW and an auxiliary contrastive objective leverage noisy and cleaned data while improving generated-response faithfulness and abstractiveness.The reported validation combines automated metrics with human evaluation.
- 7 Conclusions: Faithfulness and abstractiveness improvements hold both in-domain and in zero-shot transfer to TopicalChat and CMU-DoG.The conclusion reports this pattern across automated and human evaluations.
A AMT Instructions
The annotation protocol identifies unsupported content, edits hallucinated or generic responses to use supported knowledge, and preserves relevance, cooperation, and paraphrastic expression. Faithful responses are separately checked for cooperation with the seeker.
- Hallucination and faithfulness checks: Annotators first identify unsupported facts, opinions, feelings, advice, or other information and whether the seeker triggered that content.They also record any supported content remaining in a hallucinated response.
- Editing hallucinated responses: Hallucinated responses are edited to use only facts from K, remain informative, paraphrase rather than copy K, and stay relevant to the previous utterance.The instructions operationalize faithfulness as supported, non-copying, contextually relevant content.
- Generic-response checks: Responses with neither supported personal content nor supported factual content are flagged as generic and rewritten to be supported by K.This rule applies when both corresponding support questions receive a negative answer.
- Cooperativeness checks: Faithful responses are evaluated for cooperativeness, meaning they answer the seeker’s question without ignoring it or acting unhelpfully.Noncooperative faithful responses are modified to be relevant and cooperative.
B Pay Structure
Workers receive a base pay of $1.7 per HIT plus bonuses for successfully submitted HITs, equivalent to $17–$18 per hour.
- $17–$18 per hour is the stated equivalent compensation for workers completing 10 HITs hourly.Workers spend an average of 6 minutes per HIT, while the base pay is $1.7 per HIT and bonuses total $35–$40 per 100 successful HITs.
C Abstractiveness strategies
FaithDial responses use multiple abstractiveness strategies that preserve or derive meaning while changing wording, syntax, or length.
- Inference: Inference adds information derived through intermediate reasoning, implicatures, presuppositions, deductions, or commonsense knowledge.Examples include inferring that unfinished work was not all completed or that Elvis is a person.
- Rewording: Rewording replaces knowledge-source expressions with synonymous, more general, or more specific wording while preserving truth.The strategy includes synonymization, generalization, and specification.
- Restructuring: Restructuring changes syntactic formulation through passivization, reordering, ellipsis, or converting declarative statements into questions.Questioning restructures declarative statements as questions.
- Abridging: Abridging removes modifiers or optional complements while preserving the entailment relationship between knowledge and response.Examples include removing adjectives, adverbs, and independent clauses.
- Bridging: Bridging adds words or phrases that connect or introduce parts of an utterance.Examples include connective phrases such as “So...” and “In other words, ...”.
D Implementation Details
The implementation uses Hugging Face Transformers for critics and Transformers with PyTorch Lightning for generation models, with fixed epoch, batch, optimizer, and learning-rate schedules.
- Critic: Critics train for 10 epochs with batch size 32, Adam at 1 × 10^-5, 6% warmup, and linear learning-rate decay.
- Generation Models: Generation models train for 10 epochs with batch size 32, four-step gradient accumulation, Adam at 6.25×10^-5, 4% warmup, and linear decay.Models are evaluated twice per validation epoch, with the best model saved for testing and early stopping enabled.
- Evaluation Materials: Table 9 reports possible FaithDial abstractiveness strategies based on manual analysis of 200 responses.
- Evaluation Materials: Table 10 provides examples from T5-FAITHDIAL tested on the out-of-domain datasets TopicalChat and CMU-DoG.