Source-linked AI summary

Challenges and Contributions in Quality of AI-Based Software: A Systematic Mapping Study

Maryum Hamdani, Mateen Ahmed Abbasi, Marko Jäntti, Markku Tukiainen

arXiv:2608.26215v1cs.SE

TL;DR

AI-based software quality is difficult to define and assure because these systems are data dependent, probabilistic, and context sensitive, while existing evidence remains fragmented. This paper conducts a systematic mapping study of primary research from 2020 to January 2026, identifying recurring challenges, proposed contributions, and evidence maturity. The most prominent challenge is the limitation of existing quality assessment models, while the evidence base remains largely early-stage and weakly validated.

  • Problem

    AI-based software quality is difficult to specify, verify, and evaluate because such systems are data dependent, probabilistic, and context sensitive, while evidence remains fragmented across research scopes.

  • Method

    The paper conducts a Systematic Mapping Study using SMS guidelines to synthesize challenges, contributions, and evidence maturity across selected primary studies.

  • Results

    Limitations in existing quality assessment models are the most frequently reported challenge, covering 40% of studies (13 out of 33).

  • Takeaways & Limitations

    The findings call for collaboration among researchers, industrial practitioners, and standardization organizations to develop comprehensive quality assessment models and measurement methods.

  • Takeaways & Limitations

    External validity is limited by the publication period, selected electronic data sources, English-language restriction, and focus on software engineering.

Abstract

from arXiv · show

Artificial Intelligence (AI) is increasingly embedded in modern software systems, raising important questions about how its quality should be defined, assessed, and assured. This paper presents a Systematic Mapping Study (SMS) on the quality of AI-based software. The study synthesizes primary studies published between January 2020 and January 2026 and selected from five electronic data sources. A total of 33 primary studies were included after automated search, screening, and snowballing. The results identify six recurring challenge categories, with the most prominent being limitations in existing quality assessment models, followed by issues in non-functional requirement management, quality-aware development, and quality assurance. The findings suggest a call for collaboration of researchers and industrial practitioners with standardization organizations, that could possibly devise comprehensive quality assessments and their measurement methods.

1 Introduction

AI-based software embeds AI components into software functionality, making quality harder to address because behavior depends on data, context, and interaction. This study maps fragmented evidence on quality challenges, proposed responses, and evidence maturity through January 2026.

  • AI-based software incorporates AI components as integral parts of system functionality, predominantly using machine learning and deep learning.
  • Data dependence, probabilistic behavior, and context sensitivity complicate quality specification, verification, and evaluation.
  • Reported AI incidents increased concerns about quality, with 233 incidents recorded in 2024, a 56.4% increase over 2023.
  • Existing research covers quality attributes, evaluation methods, frameworks, reviews, mapping studies, and taxonomies, but evidence remains fragmented across subtopics and terminologies.
  • The SMS synthesizes challenges, proposed contributions, and evidence maturity across primary studies published from 1 January 2020 to 15 January 2026.

2 Research Methodology

The study uses SMS guidelines to investigate the state of the art in AI-based software quality through structured searching, screening, snowballing, and data extraction. It combines database evidence with inductive thematic analysis and established classifications of research and contribution types.

  • The study formulates three research questions covering quality challenges, proposed contributions, and classifications by contribution type, research method, and research type.
  • 2.1 Search Process: The SMS follows a two-phase process comprising automated database searching and manual backward and forward snowballing.
  • 2.1 Search Process: Search terms were developed with the PICO framework and refined through multiple pilot-search iterations.
  • 2.1 Search Process: 3,628 records were retrieved from five electronic data sources covering studies published from 1 January 2020 to 15 January 2026.
  • 2.1 Search Process: After screening, 20 primary studies remained, and snowballing added 13 studies for a final corpus of 33 primary studies.
  • 2.2 Data Extraction and Mapping Process: Inductive thematic analysis extracted reported challenges and contributions, while established schemes classified research types, contribution types, and methods.

3 Results

The study identifies six recurring challenge categories in AI-based software quality, alongside eleven contribution categories addressing them. Challenges concern quality assessment models, requirements, development, assurance, quality management, and cross-domain issues.

  • Challenges: Limitations in existing quality assessment models are identified as a recurring challenge for AI-based software.Reported limitations include restricted quality attributes and missing metrics in ISO/IEC 25059:2023.
  • Challenges: Non-functional requirement management is difficult because existing specification, definition, scoping, and prioritization methods may not fit non-deterministic ML-enabled software.The category covers requirements management issues for ML-enabled software.
  • Challenges: Quality-aware development approaches remain limited, particularly for integrating privacy, fairness, and reliability into ML-based software development.The studies also report that engineering practices for ML-intensive systems lag behind those for conventional software.
  • Challenges: Quality assurance studies report a lack of standardized approaches, practitioner agreement, and adequate quality measurement techniques for AI-based software.The reported concerns include code quality and the absence of sufficient assurance techniques for ML-based systems.
  • Challenges: Quality management frameworks are limited because traditional processes do not adequately address AI-specific attributes such as fairness and bias.This category is represented by one study in the mapped corpus.

4 Discussion

The mapped evidence centers on gaps in quality assessment models and highlights uneven empirical grounding across AI-based software quality research. It calls for stronger assessment tools, industrial validation, and collaboration with standardization organizations.

  • Challenges: 40% of studies (13 out of 33) reported limitations in existing quality assessment models.These models are considered important but insufficiently responsive to the rapid evolution of AI-based software.
  • Contributions: Quality model enhancement was the leading contribution category, with 6 out of 33 studies, followed by quality model development with 5 out of 33.Many contributions remained non-empirical or early-stage, and case studies were predominantly conducted in simulated rather than industrial settings.
  • Evidence maturity: NFR management was comparatively more empirically grounded than other challenge categories because evaluation research and field studies were more dominant there.The cross-facet finding indicates greater investigation of NFR-management practices in industrial settings.
  • Evidence maturity: Only 2 out of 33 studies contributed assessment tools for AI-based software quality.This indicates a substantial gap in practical tool support for quality assessment.
  • Implications: The findings call for collaboration among researchers, industrial practitioners, and standardization organizations to devise comprehensive quality assessment models and measurement methods.The proposed direction also includes operationalizing AI-specific quality attributes, strengthening empirical validation, and integrating quality concerns more systematically.

5 Threats to validity

The study addressed search and classification bias through pilot searches, multiple-author review, predefined coding criteria, and discussion-based resolution. External validity remains constrained by the study’s selected period, sources, language, and software-engineering focus.

  • Threats to validity: Pilot searches broadened synonyms and refined the search string, which was applied to titles, abstracts, and keywords rather than full text.The non-full-text search may have introduced a risk of missing relevant studies.
  • Threats to validity: Multiple authors reviewed extracted data and category assignments, resolving disagreements through discussion based on predefined coding criteria.A complete study protocol and replication package were provided to support transparency and reproducibility.
  • Threats to validity: External validity may be limited by the selected publication period, electronic data sources, English-language restriction, and software-engineering focus.Relevant studies from adjacent fields or other indexing services may therefore have been missed.

6 Conclusion

This paper maps quality research on AI-based software by synthesizing reported challenges, proposed contributions, and evidence maturity. It finds that quality assessment models are the most prominent challenge, while the field still needs stronger operationalization and industrial validation.

  • Study contribution: 33 primary studies were selected from 3,628 records retrieved from five electronic databases published between 2020 and 2026.Selection used systematic screening and snowballing.
  • Findings: The most frequently reported challenge was limitations in existing quality assessment models, followed by NFR management, quality-aware development, and quality assurance.The evidence base was also dominated by early-stage and weakly validated research with limited tool support and industrial validation.
  • Conclusion: The field has a strong conceptual foundation but requires further work to operationalize AI-specific quality attributes and validate proposed approaches in real-world industrial contexts.The conclusion identifies empirical validation and practical operationalization as continuing needs.
Loading 2608.26215v1…