Source-linked AI summary
Evaluating Quality of Chatbots and Intelligent Conversational Agents
Nicole M. Radziwill, Morgan C. Benton
TL;DR
Chatbot quality research lacks a unified way to organize quality attributes and assess them across varied implementations. This paper reviews the literature, synthesizes attributes and assurance approaches, and proposes an AHP-based evaluation method. The approach supports comparison between systems or across time, while the resulting attribute list remains suggestive rather than prescriptive.
Problem
The paper addresses limited guidance for identifying chatbot quality attributes and evaluating quality across varied implementations.
Method
The paper conducts a literature review and synthesizes quality attributes into an Analytic Hierarchy Process for comparative evaluation.
Results
The AHP approach compares chatbot implementations or versions over time using selected quality attributes and associated metrics.
Takeaways & Limitations
The quality attributes can serve as a checklist, while AHP can evaluate competing implementations and changes in adaptive systems.
Takeaways & Limitations
The identified quality attributes are suggestive rather than prescriptive because implementations may prioritize different attributes at different lifecycle phases.
Abstract
from arXiv · showhide
Chatbots are one class of intelligent, conversational software agents activated by natural language input (which can be in the form of text, voice, or both). They provide conversational output in response, and if commanded, can sometimes also execute tasks. Although chatbot technologies have existed since the 1960s and have influenced user interface development in games since the early 1980s, chatbots are now easier to train and implement. This is due to plentiful open source code, widely available development platforms, and implementation options via Software as a Service (SaaS). In addition to enhancing customer experiences and supporting learning, chatbots can also be used to engineer social harm - that is, to spread rumors and misinformation, or attack people for posting their thoughts and opinions online. This paper presents a literature review of quality issues and attributes as they relate to the contemporary issue of chatbot development and implementation. Finally, quality assessment approaches are reviewed, and a quality assessment method based on these attributes and the Analytic Hierarchy Process (AHP) is proposed and examined.
Methodology
The review systematically searched multidisciplinary academic and selected industry literature to identify quality research on chatbots and conversational agents. Articles were screened and narrowed to a relevant set for detailed review.
- Search strategy: Searches of Google Scholar, JSTOR, and EBSCO Host covered publications from 1990 to 2017 across five academic and industry-related domains.The search used combinations of terms concerning chatbots, conversational agents, quality, quality attributes, and quality assurance.
- Selection criteria: Selection emphasized scholarly references, supplemented by industry publications from 2016 and 2017, and required quality to be a major research aspect.Articles also needed at least one search term in the title or abstract; highly technical programming and engineering articles were excluded.
- Screening: 7,340 articles formed the initial sampling frame, which was refined to fewer than 300 for review and ultimately 36 scholarly articles and conference papers.Additional terms included evaluation, assessment, quality metrics, and metrics.
Quality Attributes
The reviewed literature organizes chatbot quality around usability-related effectiveness, efficiency, and satisfaction, with attributes spanning performance, functionality, humanity, and affect. Sources largely agree on these attributes but disagree about whether human likeness should be prioritized.
- Quality Attributes: Quality attributes were extracted from 32 papers and 10 articles, grouped by similarity, and aligned with ISO 9241 usability concepts.Effectiveness concerns accurate and complete goal achievement, while efficiency concerns how resources are applied; Table 1 organizes the attributes accordingly.
- Quality Attributes: Performance attributes include graceful degradation, robustness to manipulation and unexpected input, damage control, and escalation channels to humans.These attributes address how systems behave under adverse or unsuitable interactions.
- EFFECTIVENESS: Functionality attributes include accurate interpretation and output, task execution, transaction support, ease of use, problem solving, and breadth of knowledge.These attributes describe the system’s ability to perform requested conversational and operational functions.
- SATISFACTION: Humanity and affect attributes cover transparency, natural and satisfying interaction, themed discussion, personality, conversational cues, emotional expression, and mood responsiveness.The literature includes both passing and not passing the Turing Test among humanity-related positions.
- SATISFACTION: The literature consistently identifies quality attributes, but researchers disagree on whether acting human should be a quality priority.Most researchers prioritize human-like interaction, while others reject human likeness as a valid quality attribute; the resulting list is suggestive rather than prescriptive.
Quality Assessment
The paper addresses quality assurance by synthesizing prior approaches and recommending a composite technique based on Saaty’s Analytic Hierarchy Process.
- Quality Assessment: The review scanned the literature for quality-assurance approaches and synthesized the findings into a composite method based on the Analytic Hierarchy Process.The proposed approach is attributed to Saaty (1990).
Previous Approaches
Previous quality-assessment studies address different aspects of chatbot quality but collectively indicate a need for goal-driven evaluation incorporating users’ subjective experiences. They propose useful metrics without explaining when each metric should be applied.
- Previous Approaches: Seven references focused specifically on quality assurance, with quality serving as the central unifying theme.The reviewed studies came from academic and conference literature spanning several years.
- Previous Approaches: The studies assess different aspects of quality, ranging from response effectiveness for individual questions to customized, goal-oriented evaluation.Their main points are summarized in Table 2.
- Previous Approaches: The reviewed results support goal-driven assessment that incorporates users’ subjective experiences with chatbots.They also suggest that an absolute quality scale is unlikely across the wide variety of chatbot and conversational-agent implementations.
- Previous Approaches: Prior studies suggest many useful metrics but provide no guidance for determining when each metric should be applied.The paper therefore motivates a systematic comparison approach rather than a universal absolute scale.
Synthesized Approach: Analytic Hierarchy Process (AHP)
The proposed AHP approach organizes chatbot quality attributes hierarchically, compares their relative priorities, and evaluates alternative implementations using measured metrics and consistency checks.
- AHP framework: AHP combines qualitative and quantitative considerations by structuring quality attributes into a hierarchy and comparing their relative priorities.The hierarchy progresses from the overall goal to quality categories, selected attributes, and alternative chatbot versions.
- Alternative evaluation: The example evaluates OLD and NEW chatbot versions by comparing their performance for each of nine prioritized quality attributes.Measured values from Table 3 are used to construct lowest-level comparisons, including the Escalation attribute.
- Priority matrices: Pairwise comparisons quantify how much more important one quality category or attribute is than another, using reciprocal values for the reverse comparison.For example, a comparison of 7 in one direction corresponds to 1/7 in the other direction.
- Priority calculation: The first principal eigenvectors of the comparison matrices are combined to generate priority metrics for attributes and categories.The resulting priorities support selection of the alternative that best satisfies the quality hierarchy.
- Example result: In the example, the OLD chatbot received 66.2% overall weight compared with 33.8% for the NEW chatbot under the selected priorities.The example therefore favored OLD despite the NEW version being more effective for Escalation.
- Validation: Consistency scores are checked to identify discrepancies in individual assessments, with values below 10% ideal and values below 20% often acceptable.High inconsistency values prompt inspection for assessment or data-entry problems.
Conclusions
The paper reviews quality attributes and assurance approaches for chatbots, then proposes AHP as a practical way to organize evaluation and compare systems or their evolution over time.
- Contribution: The review gathers and articulates chatbot quality attributes while synthesizing quality-assurance approaches for future practice.The literature covered academic publications since 1990 and industry articles since 2015.
- Practical applications: Practitioners can use the quality attributes as a checklist to address key issues during chatbot implementation.
- Practical applications: AHP supports comparisons between multiple conversational systems and between system versions at different points in time.The latter use is particularly relevant for adaptive systems that learn from additional participants and topics.
- Assessment approach: The example demonstrates a goal-oriented evaluation of two chatbot implementations using pairwise comparisons that can accommodate different metrics.The approach is also described as adaptable to evaluating implementations over time.
Appendix: Analytic Hierarchy Process (AHP)
The appendix describes a three-step workflow for performing and validating the AHP analysis with software or a web-based tool.
- AHP workflow: Prepare a YAML file containing all pairwise comparisons.
- AHP workflow: Generate and validate the hierarchical model.
- AHP workflow: Perform the analysis while checking consistency scores and the overall weightings of alternatives.
Step 1: YAML
The YAML file defines alternatives, a quality-assessment goal, and pairwise preferences across criteria and chatbot options.
- YAML is a column-sensitive data serialization mechanism used in some programming environments.
- The Alternatives section lists the chatbot options and their quantitative or qualitative attributes.
- The Goal section models chatbot quality assessment as a decision process comparing old and new chatbots.
- Pairwise preferences compare quality categories including Performance, Humanity, Affect, and Accessibility.
- Lower-level criteria compare OLD and NEW chatbots for unexpected input, escalation, transparency, themed discussion, entertainment, meaning intent, and social cues.
Step 2: Generate and Validate the Hierarchical Model
The hierarchical model is generated by pasting the complete YAML into the AHP web application and validating it through visualization.
- Paste the entire YAML file into the AHP application’s Model tab, beginning in column 1 and omitting extra trailing characters or blank lines.
- Click Visualize to verify that the decision hierarchy appears and that the YAML model is valid.
Step 3: Perform the Analysis
The analysis processes the AHP data, checks consistency, and identifies the alternative with the largest first-row weight.
- Click Analyze to process the AHP data and inspect the consistency column.
- Consistency values should be no greater than 20%, and the alternative with the largest first-row weight best meets the specified goals.