Source-linked AI summary
Evaluation and Measurement of Software Process Improvement -- A Systematic Literature Review
Michael Unterkalmsteiner, Tony Gorschek, A. K. M. Moinul Islam, Chow Kian Cheng, Rahadian Bayu Permadi, Robert Feldt
TL;DR
The review examines how software process improvement initiatives are measured and evaluated, addressing the difficulty of obtaining relevant, valid information for decisions. It finds that measurement focuses on project and product quality, while return on investment is infrequently used and broader business benefits require longer-term indicators.
Problem
Software measurement is essential for SPI, but establishing a measurement program that provides relevant and valid decision information remains difficult.
Method
The paper presents a systematic literature review investigating evaluation strategies and measurements used for SPI initiatives in realistic settings.
Results
Return-on-investment appeared in 22 studies (15%) as a success indicator, while the review found a focus on process and product quality measurements.
Takeaways & Limitations
Improvement measurement should extend beyond projects to confirm short-term project or product measurements and capture business benefits at corporate level.
Takeaways & Limitations
The study quality assessment judged reporting rather than the underlying quality of the reviewed studies, limiting conclusions about how validity threats were addressed.
Abstract
from arXiv · showhide
BACKGROUND: Software Process Improvement (SPI) is a systematic approach to increase the efficiency and effectiveness of a software development organization and to enhance software products. OBJECTIVE: This paper aims to identify and characterize evaluation strategies and measurements used to assess the impact of different SPI initiatives. METHOD: The systematic literature review includes 148 papers published between 1991 and 2008. The selected papers were classified according to SPI initiative, applied evaluation strategies, and measurement perspectives. Potential confounding factors interfering with the evaluation of the improvement effort were assessed. RESULTS: Seven distinct evaluation strategies were identified, wherein the most common one, "Pre-Post Comparison" was applied in 49 percent of the inspected papers. Quality was the most measured attribute (62 percent), followed by Cost (41 percent), and Schedule (18 percent). Looking at measurement perspectives, "Project" represents the majority with 66 percent. CONCLUSION: The evaluation validity of SPI initiatives is challenged by the scarce consideration of potential confounding factors, particularly given that "Pre-Post Comparison" was identified as the most common evaluation strategy, and the inaccurate descriptions of the evaluation context. Measurements to assess the short and mid-term impact of SPI initiatives prevail, whereas long-term measurements in terms of customer satisfaction and return on investment tend to be less used.
1 INTRODUCTION
Software Process Improvement requires measurement and evaluation, yet the field lacks agreement on what to measure and how to assess initiative success. This review addresses how SPI benefits are evaluated, which measures are used, and which perspectives guide assessment.
- Motivation: Software measurement is presented as necessary for SPI programs because unmeasured efforts may address the wrong issue.
- Motivation: Measurement makes SPI outcomes visible, supports justification of improvement efforts, and enables assessment of strategies and tactics.Relevant and valid measurement is difficult to establish, and missing systematic approaches may contribute to improvement failures.
- Research objective: The review asks how SPI initiative success is determined and whether evaluation differs by initiative.
- Research objective: The review also identifies measures and classifies assessment perspectives as project, product, or organization.
- Method: The study uses a Systematic Literature Review based on Kitchenham’s approach to synthesize research and practical experience.
2 BACKGROUND AND RELATED WORK
SPI initiatives seek better software quality while reducing time-to-market and costs, using measurement to guide and verify process change. Prior reviews addressed software measurement broadly, whereas this review focuses specifically on measuring and evaluating SPI impact.
- SPI background: SPI commonly aims to increase product quality while reducing time-to-market and production costs.
- SPI background: Measurement supports the improvement cycle by assessing changed processes and confirming whether goals are achieved.
- SPI approaches: Improvement frameworks commonly compare actual processes with best practices, while bottom-up approaches derive changes from organizational knowledge.
- Related work: Earlier reviews examined software measurement generally, including what, how, and when to measure, and trends in software metrics research.
- Review contribution: This review differs by focusing on SPI measurement and examining how measures are used to evaluate and analyze process improvement.
3 RESEARCH METHODOLOGY
The study follows a structured systematic literature review process covering search design, protocol development, study selection, and data extraction. It selected empirical studies that reported how SPI initiatives were assessed.
- Review process: The review process defines research questions, develops and evaluates a protocol, searches databases, selects studies, extracts data, and assesses quality.The protocol was intended to reduce researcher bias and support future replication.
- Search strategy: The search covered Compendex, Inspec, and Google Scholar using terms for SPI and systematic reviews.
- Research questions: The review distinguishes evaluation strategies that demonstrate process-change impact from appraisals that assess conformity to maturity standards.
- Study selection: Primary-study inclusion required empirical data showing how an SPI initiative was assessed; non-empirical discussions and anecdotal evidence were excluded.
- Study selection: From 10,817 retrieved papers, screening and availability constraints reduced the full-text pool to 362 papers and yielded 148 primary studies.
3.4 Data extraction
Data extraction classified studies by research method, context, SPI initiative, success indicators, metrics, measurement perspective, and confounding factors. The review interpreted measurements in context because identical metric labels can represent different attributes or intentions.
- Research-method classification: The review classified studies into case study, industry report, experiment, survey, action research, and not stated categories using explicit criteria.
- Study context: Industry studies were further characterized by company size, customer and product type, application domain, project count, and staff size.
- SPI-initiative classification: SPI initiatives were grouped into frameworks, practices, and tools, with frameworks subdivided by how established or tailored they were.
- Metrics and indicators: The success-indicator scheme emerged during data extraction rather than being defined a priori.
- Metrics and indicators: A success indicator is an attribute of a process, product, or organization used to evaluate improvement, and metrics were assigned according to their measured object.For example, peer-review defects represent process quality, whereas post-shipment defects per KLOC represent product quality.
- Metrics and indicators: Metric categorization depended on study context and sometimes remained unresolved when descriptions lacked enough information.
- Measurement perspective: Measurement perspectives describe which entities are measured: projects and resources, delivered products, or the organization as a whole.
3.5 Study quality assessment
The review assessed study quality to guide interpretation of synthesis findings, but the assessment primarily judged reporting rather than underlying study quality.
- The quality assessment was used to guide interpretation of synthesis findings and determine the strength of inferences.
- The assessment judged how validity threats were reported, not the extent to which studies actually addressed them.
- 98% of publications provided sufficient context and background information, while all clearly stated their research aims and objectives.
3.6 Validity Threats
The review identified threats involving publication bias, study identification, review procedures, and inconsistent extraction, with several mitigations applied. Low search precision and difficulty identifying confounding factors remained validity concerns.
- Publication bias was considered moderate because positive outcomes may be more likely to be published, although the review did not target a specific SPI initiative comparison.
- The review excluded grey literature to balance broad coverage against the need for reliable information.
- Search recall reached 100% after refinement, but inconsistent terminology could still have caused relevant articles to be missed.
- 2.2% precision produced substantial screening effort and represented a moderate validity threat, as improving recall usually retrieves more irrelevant items.
- The researchers added SCOPUS and publisher sources after discovering indexing gaps for earlier studies in Software Process: Improvement and Practice.
- 234 unavailable full texts were considered a minor threat because their low expected relevance implied approximately five relevant studies.
- A protocol, parallel extraction, cross-checking, and consensus discussions were used to reduce researcher bias and improve consistency.
- Only slight or poor agreement occurred for properties P5 and P7; confounding factors were especially difficult to identify, potentially biasing extraction.
4 RESULTS AND ANALYSIS
The review analyzed 148 studies addressing measurement and evaluation of software process improvement initiatives before presenting results for its research questions.
- 148 studies discussed measurement and evaluation of software process improvement initiatives.
4.1 Overview of the studies
The reviewed literature was predominantly recent, industry-based, and focused on frameworks, especially CMM. Incomplete context reporting limits judgments about transferring initiatives to different settings.
- Publication period: 55 papers (37%) were published between 2005 and 2008, following an earlier increase of 35 papers (24%) between 1998 and 2000.
- Research methods: Case studies accounted for 66 papers (45%) and industry reports for 53 (36%), while 12 papers (8%) lacked enough methodological description for categorization.
- Study settings: 126 papers (85%) used industry settings, indicating that the review largely reflects realistic environments.
- Study context: About 50% of industry studies omitted organization size, weakening judgments about whether their SPI initiatives are feasible elsewhere.
- Study context: Future SPI research should adopt established guidelines to improve context documentation.
- SPI initiatives: Frameworks comprised 91 studies (61%), compared with 29 (20%) for practices and 9 (6%) for tools.
- SPI initiatives: Practices, tools, and practice-tool combinations totaled 42 studies, suggesting their impact was measured and evaluated less often than frameworks.
- Established frameworks: CMM was the most reported framework, appearing in 42 studies (44%), while SPICE and BOOTSTRAP were underrepresented.
4.2 Types of evaluation strategies used to evaluate SPI initiatives (RQ1)
The review identified seven evaluation strategies for SPI initiatives, with Pre-Post Comparison predominating despite limited attention to causal validity. Strategies were generally generic and adapted to organizational context and measurement needs.
- Results: 72 papers (49%) used Pre-Post Comparison, while 23 (15%) used Statistical Analysis.Pre-Post Comparison requires baseline values before improvement and comparison values afterward.
- Results: 21 papers (14%) provided relevant data but did not report an identifiable evaluation strategy.These papers remained in the review because they contributed evidence to other research questions.
- Analysis and Discussion: Pre-Post Comparison validity was rarely discussed in terms of whether outcomes were causally related to the SPI initiative.Reasonable baseline values can be difficult to establish, especially in newly created measurement programs.
- Analysis and Discussion: Most evaluation strategies were generic rather than specifically designed for SPI outcomes, so organizations applied them to different success indicators and contexts.The Philip Crosby Associates’ Approach was identified as an exception because it explicitly specifies what to evaluate.
- Statistical Analysis: Statistical methods included descriptive and inferential analysis, regression, correlation, chi-square tests, and repeated analyses over time.CUSUM was presented as a way to detect significant long-term changes in productivity or other process and product attributes.
- Other Strategies: Surveys, Cost-Benefit Analysis, the Philip Crosby Associates’ Approach, and SPAM represented additional evaluation strategies.Surveys gathered employee or customer data; cost-benefit methods assessed resource justification; SPAM modeled productivity across observed phenomena.
4.3 Reported metrics for evaluating the SPI initiatives (RQ2)
SPI evaluations measured quality most often, followed by cost and schedule, but emphasized project-oriented and short- to mid-term indicators over customer satisfaction, return on investment, and consistently defined measures.
- Success Indicators: Process Quality (57, 39%), Estimation Accuracy (56, 38%), Productivity (52, 35%), and Product Quality (47, 32%) were the most frequent success indicators.Product Quality included studies from the relevant success-indicator tables.
- Product Quality: Reliability was the most frequently measured product-quality characteristic, followed by Maintainability and Reusability.Reliability measures were often based on customer-reported product failures, whereas Usability was seldom assessed.
- Overall Measurement Patterns: Quality appeared in 92 papers (62%), cost in 61 (41%), and schedule in 27 (18%).Quality combined Process Quality, Product Quality, and Other Quality Attributes; cost combined Effort and Cost.
- Financial Benefits: Return-on-investment appeared in 22 papers (15%), although accurate financial-benefit calculation requires considering quality, cost, and schedule together.The review links ROI to communicating SPI results to stakeholders.
- Customer Satisfaction: Only 20 papers (14%) used qualitative customer-satisfaction measures and 7 (5%) described quantitative measures.Questionnaires provide broader views but require customer cooperation; reported failures require relation to other variables for validity.
- Estimation Accuracy: Schedule (37, 25%) and Cost and Effort (34, 24%) dominated estimation-accuracy measures, while quality estimation was uncommon.Examples of quality-estimation metrics included actual/estimated quality-assurance reviews and defects removed per development phase.
- Validity of Measurements: Metric instances often measured the same attribute with different units, and basic-measure definitions varied considerably across studies.The review associated these inconsistencies with the lack of shared measurement terminology and doubts about measure validity.
4.4 Identified measurement perspectives in the evaluation of SPI initiatives (RQ3)
SPI initiatives were evaluated predominantly from the project perspective, with product and organizational perspectives used less often. This created a mismatch between organization-wide improvement aims and how outcomes were assessed.
- Results: 98 papers (66%) used only the Project perspective, followed by Project and Product in 30 (20%) and Project, Product, and Organization in 8 (5%).Project-level measurement was the dominant approach across the reviewed evaluations.
- Frameworks: Organizational measurement was identified mainly in CMM-based initiatives, while product-only measurement was not identified within established SPI frameworks.Project and product perspectives were commonly combined in those frameworks.
- Analysis and Discussion: The project perspective dominated partly because project-level measurement may involve fewer confounding factors.Exclusive project measurement can nevertheless make evaluation across several projects more difficult.
- Stakeholder Needs: Project, product, and organizational measurements serve different stakeholder information needs.Corporate stakeholders need visible business benefits, whereas developers and project or product managers prioritize effects on a particular project.
- Frameworks: 77 of 91 framework-supported initiatives (85%) were evaluated from project and/or product perspectives.This contrasts with framework aims that can include organization-wide improvement, such as organizational issues at CMM level 3.
- Tools and Practices: Tools and practices generally emphasized project measurement, although tools can affect projects, product quality, and organizations.Only a small number of tool- and practice-related initiatives considered the organizational perspective.
4.5 Confounding factors in evaluating SPI initiatives (RQ4)
The review found little explicit consideration of confounding factors in SPI evaluations, limiting confidence in generalizing findings or attributing outcomes to improvement initiatives. Identifying and controlling these factors depends heavily on study context and researcher judgment.
- Extent of Consideration: The review identified only a few indications that confounding factors were explicitly considered in SPI evaluations.This suggests that such factors were seldom addressed directly.
- Extent of Consideration: Only 19 of 148 studies discussed potential validity problems involving confounding factors.The limited evidence made it difficult to generalize assumptions or connect findings to particular evaluation strategies.
- Identification and Control: Confounding factors were often described abstractly or generally without remedies for overcoming them.Some studies cautioned that relationships between process actions and product-quality improvement must be examined under specified conditions.
- Control Techniques: Random allocation, homogeneous experimental groups, blocking, and matching were described as ways to control or compensate for confounding effects.These designs aim to make groups comparable with respect to confounding variables.
- Practical Constraints: Random sampling of projects or subjects is seldom available for evaluating improvement initiatives.Consequently, recognizing potential confounding factors remains necessary for selecting compensatory techniques.
- Practical Constraints: No systematic method exists for identifying all confounding variables, because identification depends on context and researcher background knowledge.This makes complete elimination or control difficult and dependent on assumptions and logical reasoning.
5 CONCLUSION
The review finds that SPI impact evaluation is weakened by incomplete context reporting, limited treatment of confounding factors, and inconsistent measurement definitions. It also finds that measurements emphasize project- and product-level quality, while longer-term organizational benefits are less represented, motivating structured guidance for practitioners.
- The review investigates how SPI initiatives are measured and evaluated across realistic settings, identifying evaluation strategies and measurement approaches used in the field.It is based on a systematic literature review intended to characterize how improvement impact is assessed.
- 75 of 148 studies did not or only partially describe their study context, limiting readers’ ability to judge whether findings can be reused or transferred.Context includes the process change and its environment.
- In more than 50% of evaluated studies, “Pre-Post Comparison” was used alone or with another method, while confounding factors were discussed in only 19 of 148 studies.The review therefore questions the accuracy of evaluation results, especially when context descriptions are also unsatisfactory.
- Inconsistent metric definitions hinder improvement by aggravating the comparison and communication of results.The review links this problem to difficulty identifying and using appropriate measures for SPI evaluation.
- Measurements focus on process and product quality, but longer-term indicators such as customer satisfaction and return on investment tend to be used less often.The review calls for extending measurement beyond projects to assess broader business benefits.
- In 129 studies, or 87%, no discussion of potentially confounding factors was identified, underscoring the validity threat in SPI evaluation.The review notes that no good conceptual model or framework currently supports such discussion.
- The findings encourage research on structured guidelines to help practitioners measure, evaluate, and communicate the impact of improvement initiatives.