Source-linked AI summary

Guidelines for including grey literature and conducting multivocal literature reviews in software engineering

Vahid Garousi, Michael Felderer, Mika V. Mäntylä

arXiv:1707.02553v4cs.SEcs.DL

TL;DR

Existing software engineering systematic-review guidelines provide limited coverage for incorporating grey literature and handling multivocal-review-specific phases. This paper synthesizes software engineering SLR guidance, literature from other fields, and the authors’ MLR experience into guidelines covering planning, conducting, and reporting, while recommending that researchers assess whether grey literature adds value before choosing an MLR.

  • Problem

    Existing software engineering systematic-review guidelines provide limited coverage for including practitioner sources and conducting multivocal literature reviews.

  • Method

    The guidelines synthesize existing software engineering SLR guidance, MLR guidance and experience papers from other fields, and the authors’ experience conducting several MLRs.

  • Results

    The resulting guidelines cover planning, conducting, and reporting MLRs in software engineering and are presented as a unified set of guidance.

  • Takeaways & Limitations

    Researchers should assess whether including grey literature adds value before planning an MLR instead of an SLR, then apply and iteratively improve the guidelines through shared experience.

  • Takeaways & Limitations

    Including grey literature is not always straightforward or advantageous, with drawbacks including lower-quality reporting, particularly for research methods.

Abstract

from arXiv · show

Context: A Multivocal Literature Review (MLR) is a form of a Systematic Literature Review (SLR) which includes the grey literature (e.g., blog posts and white papers) in addition to the published (formal) literature (e.g., journal and conference papers). MLRs are useful for both researchers and practitioners since they provide summaries both the state-of-the art and -practice in a given area. Objective: There are several guidelines to conduct SLR studies in SE. However, given the facts that several phases of MLRs differ from those of traditional SLRs, for instance with respect to the search process and source quality assessment. Therefore, SLR guidelines are only partially useful for conducting MLR studies. Our goal in this paper is to present guidelines on how to conduct MLR studies in SE. Method: To develop the MLR guidelines, we benefit from three inputs: (1) existing SLR guidelines in SE, (2), a literature survey of MLR guidelines and experience papers in other fields, and (3) our own experiences in conducting several MLRs in SE. All derived guidelines are discussed in the context of three examples MLRs as running examples (two from SE and one MLR from the medical sciences). Results: The resulting guidelines cover all phases of conducting and reporting MLRs in SE from the planning phase, over conducting the review to the final reporting of the review. In particular, we believe that incorporating and adopting a vast set of recommendations from MLR guidelines and experience papers in other fields have enabled us to propose a set of guidelines with solid foundations. Conclusion: Having been developed on the basis of three types of solid experience and evidence, the provided MLR guidelines support researchers to effectively and efficiently conduct new MLRs in any area of SE.

1 INTRODUCTION

Software engineering reviews have largely focused on formal academic literature, excluding practitioner-produced grey literature. The paper therefore proposes synthesized guidelines for planning, conducting, and reporting multivocal literature reviews that incorporate both forms of evidence.

  • SLRs help researchers and practitioners index evidence and gaps but typically exclude the grey literature produced by software engineering practitioners.
  • MLRs extend SLRs by combining peer-reviewed academic papers with grey-literature sources such as blogs, videos, white papers, and web pages.
  • Including grey literature recognizes multiple voices and can support rigorous identification of emerging software engineering topics originating in industry.
  • The paper develops a synthesized guideline because existing guidance is scattered across disciplines, may conflict, and does not address software engineering’s specific grey-literature sources.
  • The guidelines cover three review phases: planning, conducting, and reporting the multivocal literature review.

2 BACKGROUND

This background defines grey literature and multivocal literature reviews, distinguishes MLRs from other secondary studies, and motivates dedicated SE guidelines. It also summarizes their benefits, emerging use, and limitations.

  • Grey literature: Grey literature is produced by government, academic, business, or industry bodies but is not controlled by commercial publishers.It includes literature not formally published in sources such as books or journal articles.
  • Grey literature: The adapted “shades” model classifies grey-literature sources continuously by producer expertise and outlet control.SE adaptations add Q/A websites such as StackOverflow; examples range from high-control white papers to low-control blogs, emails, and tweets.
  • Secondary studies: An MLR combines the academic peer-reviewed sources of an SLR with grey sources such as blogs, videos, white papers, and webpages.This multivocal approach recognizes multiple perspectives, purposes, and information bases rather than relying only on formally reported academic knowledge.
  • Secondary studies: Because an MLR unites the sources studied by an SLR and a grey-literature review, it is expected to provide a more complete picture of evidence and both state-of-the-art and state-of-practice.The paper positions MLR as one of six systematic secondary-study types differentiated by source types and analysis types.
  • Benefits and need: Grey literature can provide current perspectives, complement formal-literature gaps, and help avoid publication bias, while also introducing experience- and opinion-based evidence.The authors report that excluding grey literature would have omitted test-engineer expertise in the MLR-AutoTest and potentially affected research direction.
  • Benefits and need: Including grey literature is not always advantageous because its reporting quality, particularly for research methodology, may be lower; researchers should assess whether broadening adds value.The paper therefore does not advocate converting all SE SLRs into MLRs.
  • Benefits and need: The paper argues that MLR guidelines are needed because existing SE SLR guidance rarely provides explicit procedures for collecting and conducting MLRs.The authors report approximately nine SE MLRs published between 2015 and 2018, with substantial variation in reviewed-source counts and grey-literature proportions.

3 AN OVERVIEW OF THE GUIDELINES AND ITS DEVELOPMENT

The guidelines were developed from existing SE review guidance, a survey of MLR guidance in other fields, the authors’ review experience, and related practitioner-evidence work. They organize MLR guidance around established review phases while focusing on grey-literature-specific differences and illustrating the process with MLR-AutoTest.

  • Guideline development: Four inputs informed the guideline development: MLR guidance from other fields, SE SLR and SM guidance, the authors’ review experience, and a study using argumentation theory for practitioner evidence.The inputs combine published guidance, prior review practice, and a related approach to analyzing practitioners’ evidence.
  • Guideline purpose: The guidelines address a gap in SE review guidance by emphasizing grey literature and providing concrete, example-supported recommendations for incorporating it into review studies.They are intended to support effective and efficient MLR execution and reporting in software engineering.
  • Guideline development: The survey identified 24 MLR guideline and experience papers, whose recommendations covered decisions to include grey literature, planning, searching, source selection, quality assessment, extraction, synthesis, reporting, and other guidance.The papers were synthesized and adapted to MLRs in software engineering.
  • Survey of guidance: Fourteen of the 24 surveyed papers provided guidance for the MLR search process.Figure 6 classifies the surveyed papers by MLR activity; further classification details are referenced in an online source.
  • Running example: MLR-AutoTest, a review about deciding when and what to automate in testing, serves as the running example for discussing implementation of each MLR step.The authors also identify cases where guidelines should have been applied because the guidelines were developed after some MLR studies were conducted.
  • Guideline structure: The guidelines adopt three SLR phases—planning, conducting, and reporting—but present mainly the steps that differ for MLRs, especially those concerning grey-literature sources.Existing SLR guidance already covers formal literature, and the authors report that integrating formal and grey-literature sources is usually straightforward.

4 PLANNING A MLR

MLR planning adapts the SLR process by establishing the review’s need, defining its goal and research questions, and deciding systematically whether grey literature should be included. The guidelines emphasize audience usefulness, objective and measurable questions, and traceability from questions through searching, extraction, and synthesis.

  • Planning process: The typical MLR process adapts established SLR guidance and can structure an MLR protocol or serve as variation points within a standard SLR protocol.The process is visualized to improve understandability and traceability between process steps and the guidelines covering them.
  • Establishing the need: Planning begins by confirming the need for a systematic review, identifying existing reviews, and defining usefulness for the intended researchers and/or practitioners.The authors recommend considering whether an SLR, grey-literature review, MLR, or mapping-study counterpart is appropriate.
  • Deciding whether to include grey literature: The decision to include grey literature and conduct an MLR should use a well-defined set of criteria or questions, such as the seven-item decision aid.For MLR-AutoTest, all seven criteria received “Yes” answers; the authors state that a larger sum indicates a higher need for an MLR.
  • Defining goals and research questions: The research questions drive the review because searching must identify relevant primary studies, extraction must collect needed data, and synthesis must answer the questions.The example MLR grouped four top-level questions with several sub-questions, while addressing practitioner needs such as factors relevant to deciding when and what to automate.
  • Defining goals and research questions: Research questions should relate systematically to the review goal, match target-audience needs, and be as objective and measurable as possible.The authors recommend using GQM to connect the review goal, questions, and metrics, and suggest explicitly considering varied question types.

5 CONDUCTING THE REVIEW

Conducting an MLR comprises five phases: searching, source selection, study quality assessment, data extraction, and data synthesis.

  • The conducting phase is organized into search, source selection, study quality assessment, data extraction, and data synthesis.

5.1 Search process

MLR searches require distinct strategies for formal and grey literature, explicit choices about grey-literature sources, and stopping rules suited to heterogeneous search results. The paper identifies three possible stopping criteria: theoretical saturation, effort bounded, and evidence exhaustion.

  • Formal literature is searched through academic databases, whereas grey literature requires web search engines, specialized databases or websites, backlinks, and contacting individuals.
  • Search strings should be refined through informal pre-searches because software-engineering terminology, especially in grey literature, lacks standardization.
  • Relevant grey-literature types and producers should be identified early and justified explicitly; MLR-AutoTest included white papers, blog posts, and YouTube videos.
  • Search stopping depends on evidence type, search volume, and evidence quality; MLR-AutoTest received 1,330,000 Google hits and examined the first 100 hits initially.
  • Grey-literature searches can stop at theoretical saturation, an effort-bounded top N, or evidence exhaustion.

5.2 Source selection

MLR source selection is more demanding because grey literature is diverse and less controlled. The guidelines call for fine-grained quality-aware criteria and coordinated selection across grey and formal literature.

  • Grey-literature source selection is especially time-consuming and difficult because sources are more diverse and less controlled than formal literature.
  • Selection criteria should account for source type and use quality considerations to identify sources providing direct evidence about the review questions.
  • Grey-literature inclusion and exclusion criteria should be combined with quality-assessment criteria.
  • Formal- and grey-literature selection should be integrated in a coordinated process, without allowing effort on one source type to reduce effort on the other.

5.3 Quality assessment of sources

Grey-literature quality assessment requires a more diverse and laborious process than formal-literature assessment. The guidelines recommend adaptable criteria, coordinated selection and assessment, and explicit scoring thresholds.

  • Grey-literature quality is more diverse and laborious to assess because its production and publication processes are less controlled.
  • Inclusion and exclusion criteria should be combined with grey-literature quality assessment criteria.Methodology, publication date, or backlinks can support early source selection and reduce later assessment effort.
  • Quality assessment should adapt authority, methodology, objectivity, date, novelty, impact, and outlet-control criteria to the source type.No single checklist fits every grey-literature type; criteria may require reductions or extensions for specific study designs.
  • Quality assessment can use richer scoring schemes, including partial scores, inter-reviewer agreement, and explicit inclusion thresholds.A three-point scale can assign yes=1, partly=0.5, and no=0 before defining a threshold.
  • Using a quality threshold of 10 out of 20 would include sources above the threshold and exclude sources below it.
  • Five example grey-literature sources received total quality scores of 13, 19, 16, 15.5, and 12 out of 20.Their normalized scores were 0.65, 0.95, 0.80, 0.78, and 0.60, respectively.

5.4 Data extraction

MLR data extraction largely follows SLR practice but must address grey literature’s less standardized structure. The guidelines emphasize research-question traceability, operational procedures, and recording source locations and coverage.

  • MLR data-extraction form design is mostly similar to SLR practice, with additional considerations arising from grey literature.
  • Extraction forms should directly trace research questions to attributes, possible values, and selection rules.The MLR-AutoTest systematic map organized research questions, corresponding attributes, possible values, and whether multiple selection was allowed.
  • Traceability links should identify where extracted information appears in grey-literature sources.Comments or verbatim source text help reviewers verify data and locate its precise origin.
  • Without traceability information, peer review and locating the exact evidence underlying extracted data become challenging.
  • Researchers should record each grey-literature document’s purpose, coverage, and relevant assumptions because sources vary in audience and thoroughness.
  • Extraction should collect enough qualitative and quantitative data to answer each research question without requiring repeated review of primary sources.

5.5 Data synthesis

Data synthesis should match the research questions and evidence types in the included sources. Grey literature commonly supports qualitative synthesis, while reporting limitations often constrain quantitative meta-analysis.

  • Synthesis techniques should be selected according to the research questions and the type of primary-study data.
  • Practitioner grey literature commonly provides qualitative, experience-based evidence that requires qualitative analysis techniques.The MLR-AutoTest example used open and axial coding to derive factors for deciding when to automate testing.
  • Grey-literature surveys may support meta-analysis when questionnaires are repeated, but missing standard deviations often make statistical meta-analysis impossible.
  • Argumentation theory can support synthesis of practitioner beliefs and expert opinions from grey literature.Relevant questions address the writer’s expertise, opinion, trustworthiness, consistency, and evidential backing.
  • Synthesis should balance sources with different rigor rather than treating blog posts and research papers as contributing equal evidence weight.
  • Many grey-literature sources suit qualitative coding, while limited evidence depth and reporting rigor constrain meta-analysis.

6 REPORTING THE REVIEW

MLR reporting should serve its intended audience while preserving appropriate transparency and practical usefulness. Online repositories, audience-specific writing, practitioner feedback, and implications sections are emphasized.

  • MLR reporting should address both researchers and practitioners because reviews summarize evidence about a field’s state of the art and practice.
  • Practitioner-oriented reporting can include shortened versions, clear implications, benefits of the review, and feedback from practitioners.
  • Scientific-journal reports should cover the MLR’s planning and search processes, whereas practitioner outlets should generally be shorter and more direct.
  • Researchers should publish an online repository of included review studies, ideally with export, search, and filtering functions.
  • Writing style should match the target audience: concise and practical for practitioners, transparent and methodologically detailed for researchers.

7 CONCLUSIONS AND FUTURE WORKS

The paper presents experience-based guidelines for planning, conducting, and reporting multivocal literature reviews in software engineering, addressing limited existing coverage of practitioner sources. It recommends iterative application and community feedback while acknowledging the need for empirical evaluation and the possibility of researcher bias.

  • Existing SE guidelines provide limited coverage of practitioner sources and multivocal literature reviews.
  • The guidelines draw on SE SLR and systematic mapping guidance, MLR literature from other fields, and the authors’ experience conducting several SE MLRs.
  • The resulting guidance covers planning, conducting, and reporting MLR studies in software engineering.
  • The guidelines still require empirical evaluation, and researcher bias may remain despite efforts to mitigate it.
  • Researchers are encouraged to apply the guidelines and share lessons learned so the guidance can be assessed and improved iteratively.
  • Future work includes developing guidance for different review types and objectives.
Loading 1707.02553v4…