Source-linked AI summary
Guidelines for the Search Strategy to Update Systematic Literature Reviews in Software Engineering
Claes Wohlin, Emilia Mendes, Katia Romero Felizardo, Marcos Kalinowski
TL;DR
Software-engineering SLRs may become outdated, yet there was no standard search strategy for updating them. The paper compares search strategies, formulates guidelines, and evaluates them on a separate SLR. The evaluation supports single-iteration forward snowballing with Google Scholar using the original SLR and its primary studies as seeds.
Problem
Many software-engineering SLRs may be outdated, and no standard proposal existed for searching for new evidence when updating them.
Method
The authors compare search strategies in an effort-estimation SLR, its update, and two replications, then evaluate the resulting guidelines on a separate SLR update.
Results
The evaluation supports single-iteration forward snowballing with Google Scholar, using the original SLR and its primary studies as the seed set.
Takeaways & Limitations
The authors recommend adopting these guidelines for updating software-engineering SLRs.
Takeaways & Limitations
The evaluation may be biased because the selected SLR had inclusion criteria focused primarily on specific words, favoring database or indexing-service searches over citation searches.
Abstract
from arXiv · showhide
Context: Systematic Literature Reviews (SLRs) have been adopted within Software Engineering (SE) for more than a decade to provide meaningful summaries of evidence on several topics. Many of these SLRs are now potentially not fully up-to-date, and there are no standard proposals on how to update SLRs in SE. Objective: The objective of this paper is to propose guidelines on how to best search for evidence when updating SLRs in SE, and to evaluate these guidelines using an SLR that was not employed during the formulation of the guidelines. Method: To propose our guidelines, we compare and discuss outcomes from applying different search strategies to identify primary studies in a published SLR, an SLR update, and two replications in the area of effort estimation. These guidelines are then evaluated using an SLR in the area of software ecosystems, its update and a replication. Results: The use of a single iteration forward snowballing with Google Scholar, and employing as a seed set the original SLR and its primary studies is the most cost-effective way to search for new evidence when updating SLRs. Furthermore, the importance of having more than one researcher involved in the selection of papers when applying the inclusion and exclusion criteria is highlighted through the results. Conclusions: Our proposed guidelines formulated based upon an effort estimation SLR, its update and two replications, were supported when using an SLR in the area of software ecosystems, its update and a replication. Therefore, we put forward that our guidelines ought to be adopted for updating SLRs in SE.
1. Introduction
The paper addresses the lack of systematic guidance on searching for new evidence when updating software-engineering SLRs. It formulates guidelines by comparing search strategies and evaluates them on a separate SLR update.
- 1. Introduction: More than 430 software-engineering SLRs were published from 2004 to 2016, but only 20 updates were identified from 2006 to 2018.Many SLRs may therefore be out of date, while their primary studies may also rely on obsolete technologies.
- 1. Introduction: The paper fills a gap because prior work had not systematically compared search approaches for identifying new evidence when updating SE SLRs.An update starts from an existing SLR, requiring a choice between replicating its search strategy and using information from the original review.
- 1. Introduction: The authors compare database search and forward-snowballing strategies using an effort-estimation SLR, its update, and two replications.Different authors were sought for replications to reduce bias when applying alternative methods for identifying primary studies.
- 1. Introduction: The proposed guideline recommends one forward-snowballing iteration with Google Scholar, using the original SLR and its primary studies as seeds.The authors note that conducting the compared studies at different times was not optimal, although simultaneous replication could also introduce bias.
- 1. Introduction: A separate SLR update was conducted to evaluate the guidelines, and its outcome was compared with a published update.The evaluation was intended to test whether the guidelines provided a better alternative than database searching.
2. Related work
Earlier work addressed updating processes, visual evidence selection, search sources, and practical lessons, but not a detailed comparison of search strategies. This paper contributes such a comparison and evaluates its recommendations on a separate SLR.
- 2. Related work: Existing SE research developed processes for updating SLRs but did not focus on how to search for new evidence.The paper distinguishes process support from recommendations about the search strategy itself.
- 2. Related work: Visual Text Mining supports selecting update evidence from the original SLR’s studies, whereas this research compares alternative search strategies.The paper’s recommendations are based on detailed comparisons rather than a visualization-based analytical technique.
- 2. Related work: Google Scholar was judged adequate for supporting secondary-study updates, while IEEE Xplore was judged insufficient to identify most studies.That prior work compared sources, but not search approaches or their effects on included-study sets and conclusions.
- 2. Related work: Prior lessons included reusing the original protocol, documenting the SLR, involving original authors, and using software tools.This paper similarly reports lessons learned but primarily provides evidence-based search recommendations.
- 2. Related work: The paper claims no previous SE study had combined a detailed search-strategy comparison with evaluation on a separate SLR and update.This comparison-and-evaluation design is presented as the work’s primary distinction from earlier studies.
3. Research design
The research design formulates guidelines from an SLR update and two replications, then evaluates them by applying the guidelines to a separate SLR and comparing the resulting update with the published one.
- 3. Research design: The formulation study selected an effort-estimation SLR with an update and two replications using different search strategies.The three update papers were treated as updates of the same original SLR.
- 3. Research design: The study defined five research questions covering retrieval, inclusion, seed sets, search processes, and differences in SLR conclusions.The questions compare each update with the union of papers found across the update and replications, called the superset.
- 3. Research design: The superset was defined as the union of papers found by the original update and its two replications, separating search effects from primary-study selection.This design investigates why the three updates did not include identical studies.
- 3. Research design: The resulting guidelines were evaluated in a separate SLR study rather than only in the evidence used to formulate them.The evaluation was designed to apply the guidelines and compare the new update with a published update.
- 3. Research design: The evaluation required an original SLR and database-search update, identifiable primary studies, and an area familiar to the researchers.The authors also preferred an SLR whose original and updated papers had not been co-authored by the guideline developers.
- 3. Research design: The selected evaluation SLR was updated by analyzing citations to the original SLR and its primary studies, then comparing included papers with the published update.The evaluation asked how many papers each update identified and how much their results overlapped.
4. Selection of SLRs for formulating and evaluating the guidelines
The guidelines were formulated from an effort-estimation SLR, its update, and two replications, then evaluated through a separate software-ecosystems SLR and update selection.
- 4.1 SLR for formulating the guidelines: The original SLR’s studies were organized by reported findings, including statistically significant differences and inconclusive results lacking significance tests.Inconclusive results could reflect missing data or information.
- 4.1 SLR for formulating the guidelines: The effort-estimation SLR, its update, and two replications produced a 15-study superset for comparing search strategies.The studies concerned cross-company versus within-company effort estimation and were published through the end of 2013.
- 4.1 SLR for formulating the guidelines: The replication results differed from the update and motivated further investigation of how search strategies identify evidence when updating SLRs.The paper separately distinguishes search strategy from primary-study selection when analyzing these differences.
- 4.1 SLR for formulating the guidelines: The update used the original search string and protocol, identified 11 primary studies, and missed studies S22–S25.The update was published in 2014 and included two of the eventual paper’s co-authors.
- 4.2 SLR for evaluating the guidelines: The evaluation required an original SLR and database-search update, excluded an SLR already using forward snowballing, and selected SLR5 after additional eligibility criteria.The selected SLR concerned software ecosystems.
- 4.2 SLR for evaluating the guidelines: The selected software-ecosystems SLR covered 2007–2012, while its later update covered 2007–2014, requiring filtering to the 2012–2014 search span.The original SLR contained 90 primary studies, and the update’s inclusion criteria were retained.
- 4.2 SLR for evaluating the guidelines: Evaluation applied criteria requiring English papers containing “software ecosystem” or “software ecosystems,” while excluding books, short articles, keynotes, and extended abstracts.The evaluation also assumed that criteria from the original SLR remained applicable where the update did not restate them.
5. Formulating the guidelines
Comparisons of search strategies, seed sets, and screening outcomes supported guidelines for updating SLRs in Software Engineering. The recommended approach combines Google Scholar forward snowballing from the original SLR and its primary studies with independent screening by multiple researchers.
- 5.1 Results: Using only the original SLR as a seed set retrieved all studies except S19, whereas adding its primary studies enabled complete retrieval.The sole-original-SLR seed set was therefore insufficient to retrieve all 15 superset studies.
- 5.1 Results: The compared search strategies produced broadly aligned SLR conclusions, despite some conclusive studies being missed by individual seed sets.Four conclusive studies were missed by at least one seed set, but other included studies supported the same overall interpretation.
- 5.1 Results: Only one conclusive study differed between the two alternative seed-set approaches, suggesting no substantial change in SLR results within this context.The study was S18, which was neither included by Alternative 1 nor retrieved by Alternative 2.
- 5.1 Results: The update and replications omitted three non-inconclusive studies, including S22, one of only two studies reporting superior accuracy for cross-company models.SLR-update-R1 and SLR-update-R2 included S22, whereas the SLR update retrieved but did not include it.
- 5.1 Results: Compared with the original SLR, the 15-study superset preserved similar proportions of within-company superiority, similar accuracy, and inconclusive results, while adding two cross-company-superiority results.The original SLR had 40%, 40%, and 20% for the first three categories; the superset had 39%, 39%, 13%, and 9% for four categories.
- 5.2 Recommended guidelines based on a comparison of results: The only search strategy retrieving all 15 studies used single-iteration forward snowballing with Google Scholar and a seed set containing the original SLR and its primary studies.The original SLR papers were also included in the seed set.
- 5.2 Recommended guidelines based on a comparison of results: Large retrieval volumes increased the likelihood of false negatives during screening, especially when many titles and abstracts had to be reviewed or only one researcher screened them.The study explicitly links these conditions to likely false-negative results.
- 5.2 Recommended guidelines based on a comparison of results: The proposed guidelines are to use the original SLR plus primary studies as seeds, search Google Scholar, apply one forward-snowballing iteration, and involve multiple researchers in initial screening.The researchers recommend independent screening by at least two experienced SLR researchers when resources permit.
6. Evaluating the guidelines
The evaluation process is organized into four remaining steps, each presented in its own subsection.
- The evaluation reports the remaining four steps separately.Each step is presented in one subsection.
- The process follows the design described earlier in the paper.
- This section presents results from the evaluation process rather than introducing its research design.
6.1 Preparation
Preparation standardized the Google Scholar search and aligned inclusion decisions before distributing citation assessment among researcher pairs.
- Google Scholar searches used the original SLR title with quotes and patents unticked, following a common procedure for distributed authors.
- After discussion, decisions were resolved for 13 of the 17 initially disputed papers.The remaining four cases were settled through further discussion or procedural checks.
- The assessment ultimately included 38 papers citing the original SLR and exposed reviewer mistakes, supporting multiple researchers in screening.
- Citation assessment for the 90 primary studies was assigned to mixed teams of two researchers to reduce team and individual bias.
- Each citing paper was investigated only once using Google Scholar link-color changes, reducing duplicate work but preventing agreement measurement.
6.2 Collecting data – Forward snowballing
Forward snowballing produced 100 new eligible papers after team-based screening, while staggered Google Scholar searches introduced small result differences and duplicates.
- Reviewer teams reached high agreement when assessing 15 primary studies each, resolving disagreements by revisiting criteria and discussing papers.Most differences resulted from reviewer mistakes or venue-eligibility interpretation.
- Google Scholar returned slightly different results across teams, potentially increasing discovery but requiring duplicate removal.The searches covered papers published between 2012 and 2014.
- 100 new papers citing at least one seed study fulfilled the inclusion criteria after removing overlap with the original SLR.The resulting collection contained papers not included in the original SLR.
6.3 Results – Analysis of the SLR updates
The evaluation compared the guideline-based update with a published database-search update, finding substantial overlap but better topic coverage from the guideline-based approach.
- The evaluation used a separate SLR from the guideline-formulation study to assess the recommendations’ usefulness.
- 99 papers were common to both updates, corresponding to 72% overlap against 138 papers and 75% against 132 comparable papers.Eight papers from the published update were excluded because they did not satisfy the inclusion criteria.
- At least 29 papers in the guideline-based update could not have been identified through the databases searched by Manikas.Seven of ten examined papers should have been available, while three were unavailable at the relevant time or appeared later.
- Among 33 papers found only by Manikas, no paper received agreement from both evaluators as contributing to software-ecosystems research.Three papers received support from at least one evaluator, while 30 received none.
- Among 39 papers found only by the guideline-based update, 21 were agreed to contribute to software-ecosystems research, while eight had partial evaluator support and ten had none.
- The guideline-based update identified more research concerning software ecosystems than repeating the original SLR’s database search strategy.
6.4 Reporting – Observations from the evaluation
The evaluation found that the proposed citation-based search guidelines performed favorably against database searching, while overlap between updates suggests broadly similar conclusions.
- 6.4 Reporting – Observations from the evaluation: Using the original SLR and its primary studies as a seed set generated approximately as many papers as the published SLR update.The authors attribute this to citations contributing more relevant papers than searching for specific database wordings in this case.
- 6.4 Reporting – Observations from the evaluation: The guidelines outperformed the published update in identifying papers focused on software ecosystems, although papers outside that focus could skew the findings.The two updates had 70–75% overlap, indicating that their conclusions ought to be similar.
- 6.4 Reporting – Observations from the evaluation: Forward snowballing was judged an excellent alternative to repeating database searches over a different time interval.The authors also suggest complementing or replacing database searches with snowballing when seeking a broad research sample.
7. Threats to validity
The authors identify threats involving conclusion validity, citation coverage, publication bias, and researcher bias, while describing design choices intended to mitigate them.
- 7. Threats to validity: The guidelines rely on evidence from one original SLR, one update, and two replications, creating a threat to conclusion validity despite a separate evaluation.The authors chose this evidence-based formulation instead of a formal experiment with a simulated scenario.
- 7. Threats to validity: Author familiarity with the original SLR could have improved snowballing effectiveness, although papers retrieved by database searches were also retrieved through forward snowballing.The latter observation contradicts the proposed higher-effectiveness explanation based solely on author knowledge.
- 7. Threats to validity: Forward snowballing may miss new papers that cite neither the original SLR nor its primary studies, especially when the topic has changed substantially.The authors recommend a validation set and treating the work as a new SLR if several known papers cite neither source.
- 7. Threats to validity: Selection and evaluator bias could favor the guidelines, but predefined SLR-selection criteria and different research teams were used to reduce these risks.The selected SLR’s word-focused criteria were expected to favor database or indexing searches rather than citation searching.
- 7. Threats to validity: Publication bias may affect all search approaches, although the authors judge it least sensitive for forward snowballing and report that snowballing found papers absent from major databases.They conclude that any publication bias in this evaluation favors the proposed guidelines over other approaches.
8. Conclusions
The paper proposes four guidelines for updating software-engineering SLRs and evaluates them across distinct SLR contexts. The evaluation supports the guidelines while motivating further investigation.
- 8. Conclusions: The recommended procedure uses the original SLR and primary studies as seeds, Google Scholar, non-iterated forward snowballing, and multiple screeners.The screening recommendation aims to reduce false negatives during initial selection.
- 8. Conclusions: The evaluation supports the guidelines across both a narrow-scope SLR with few primary studies and a broad-area SLR with numerous primary studies.The formulation and evaluation used different SLRs, providing support beyond the original effort-estimation setting.
- 8. Conclusions: The guidelines are based on the expectation that new papers cite either the original SLR or at least one of its primary studies.Using the guidelines outperformed the original update for finding software-ecosystems research papers.
- 8. Conclusions: Further studies, including evidence from additional SLR updates or formal experiments, are needed to broaden understanding of effective and efficient SLR updating.The authors present the findings as a promising way forward rather than a final resolution of the topic.