Source-linked AI summary

On the Pragmatic Design of Literature Studies in Software Engineering: An Experience-based Guideline

M. Kuhrmann, D. Méndez Fernández, M. Daneva

arXiv:1612.03583v1cs.SE

TL;DR

Software-engineering literature studies face extensive publication sets and practical gaps in existing guidance for study design, data collection, and selection. This article distills the authors’ experience into a reusable, experience-based guideline and blueprint. It emphasizes early-stage procedures while documenting practical lessons, tools, workflows, and scope boundaries.

  • Problem

    Existing guidelines give insufficient practical advice for handling study design, data collection, selection, and tool-supported workflows, leaving substantial reliance on researcher expertise.

  • Method

    The authors distill experience from systematic reviews and mapping studies into guidance covering study design, data collection, dataset cleaning, study selection, tools, and reusable workflows.

  • Results

    The article provides an experience-based blueprint that streamlines and links early literature-study practices, explaining how and why they can be executed pragmatically.

  • Takeaways & Limitations

    Researchers, particularly those conducting an initial literature study, can use the blueprint and lessons learned to design and enact domain-specific studies.

  • Takeaways & Limitations

    Search reproducibility is constrained because digital-library indexes continuously evolve, while result-set size does not indicate content quality.

Abstract

from arXiv · show

Systematic literature studies have received much attention in empirical software engineering in recent years. They have become a powerful tool to collect and structure reported knowledge in a systematic and reproducible way. We distinguish systematic literature reviews to systematically analyze reported evidence in depth, and systematic mapping studies to structure a field of interest in a broader, usually quantified manner. Due to the rapidly increasing body of knowledge in software engineering, researchers who want to capture the published work in a domain often face an extensive amount of publications, which need to be screened, rated for relevance, classified, and eventually analyzed. Although there are several guidelines to conduct literature studies, they do not yet help researchers coping with the specific difficulties encountered in the practical application of these guidelines. In this article, we present an experience-based guideline to aid researchers in designing systematic literature studies with special emphasis on the data collection and selection procedures. Our guideline aims at providing a blueprint for a practical and pragmatic path through the plethora of currently available practices and deliverables capturing the dependencies among the single steps. The guideline emerges from various mapping studies and literature reviews conducted by the authors and provides recommendations for the general study design, data collection, and study selection procedures. Finally, we share our experiences and lessons learned in applying the different practices of the proposed guideline.

1 Introduction

Systematic literature studies structure reported software-engineering knowledge, but researchers still lack practical guidance for designing, collecting, and selecting evidence at scale. The article contributes an experience-based blueprint and lessons learned for making these processes pragmatic and reusable.

  • 1 Introduction: Systematic mapping studies structure broad fields by classifying and counting publications, whereas systematic reviews analyze evidence from smaller, more specific publication sets in depth.Mapping studies provide a broad picture; reviews additionally integrate knowledge, identify inconsistencies, and reveal areas needing investigation.
  • 1 Introduction: Researchers must manage large publication sets while designing studies, constructing searches, filtering relevance, and coordinating distributed work.The challenges include handling hundreds or thousands of papers and organizing team-based selection and analysis.
  • 1 Introduction: Existing guidelines provide insufficient practical advice because they are generic, method-focused, or limited to selected techniques rather than explaining cost-effective execution and dependencies.Consequently, literature studies continue to depend substantially on researcher expertise and lack adequate tool support, especially for novices.
  • 1 Introduction: The article contributes a detailed blueprint for study design, data collection, and study selection, together with practical lessons learned and reusable supporting material.The contribution is based on the authors’ own systematic reviews and mapping studies.
  • 1 Introduction: The proposed path connects available practices and deliverables so researchers can reuse the blueprint for domain-specific studies and continue with analysis tailored to their research questions.The intended users already possess basic knowledge of general literature-study guidelines.

2 Study Design and Data Collection: An Experience-based Approach

The guideline organizes literature-study work into preparation, data collection and dataset cleaning, and study selection, while relating these stages to later analysis. It emphasizes the early stages as a practical addition to evidence-based software-engineering methods.

  • 2 Study Design and Data Collection: An Experience-based Approach: The guideline covers preparation, data collection and dataset cleaning, and study selection, with inputs and outputs shown for each phase.The later analysis varies according to whether the study is a mapping study, literature review, or combination.
  • 2 Study Design and Data Collection: An Experience-based Approach: Data analysis and visualization may use classification, qualitative visualization, quantitative visualization, or maps depending on the study type and research questions.The analysis is connected to the individual facets selected for the mapping study or literature review.
  • 2 Study Design and Data Collection: An Experience-based Approach: The guideline is presented as a new building block for the methodical instrumentation of evidence-based software engineering because it emphasizes early literature-study stages.Its detailed relation to existing guidelines is discussed separately in the article.

2.1 Preparation

Preparation defines study goals, scope, research questions, databases, search queries, and inclusion or exclusion criteria before collection and selection begin. The guideline combines broad descriptive questions with carefully tested searches and explicit dataset-quality decisions.

  • 2.1 Preparation: Literature studies may collect knowledge broadly through mapping studies or analyze it in depth through systematic reviews, so goals determine the appropriate study design.Mapping studies quantify selected aspects, whereas reviews examine publications in detail.
  • 2.1.1 Research Goals and Research Questions: Generic research questions help describe publication quantity, frequency, topics, research types, contribution types, and trends across a study’s result set.These questions support a demographic overview and help prepare collection and selection procedures.
  • 2.1.1 Research Goals and Research Questions: Research questions should be narrowly defined for efficient planning while being adapted to whether the study is a mapping study or systematic review.The guideline recommends combining generic questions with study-specific questions.
  • 2.1.2 Search Strings and Search Engines: Search strings must be developed for the domain and tested before the actual search because imprecise queries can produce irrelevant overhead or incomplete result sets.Suggested approaches include preliminary snowballing and iterative trail-and-error searches using meta-search engines.
  • 2.1.3 Inclusion and Exclusion Criteria: Search results may contain thousands of publications, making dataset cleaning and explicit inclusion and exclusion criteria necessary for rigorous, reproducible selection.The criteria address relevance, duplicates or out-of-scope contributions, and full-text availability; full text is especially important for in-depth reviews.

2.2 Data Collection and Dataset Cleaning

Data collection combines searches across multiple literature sources with export and integration practices, followed by cleaning duplicates and out-of-scope publications. Word clouds can help inspect and narrow result sets, but require careful use to avoid losing relevant papers.

  • Data collection: Automated searches should be prepared for source-specific query formats and may benefit from multiple overlapping queries.Different sources impose different query constraints, so repeated simple searches can improve practical search execution.
  • Data collection: Searches should cover standard digital libraries and may be complemented by meta-search engines to identify publications absent from primary sources.Meta-search results can be less repeatable and may introduce duplicates or extraneous publications.
  • Data collection: Data should be exported in both a literature-management format and plain or comma-separated text files for integration and later analysis.Different databases use different export formats, making joining and integration time-consuming.
  • Dataset cleaning: Cleaning removes out-of-scope contributions and duplicates, while duplicate handling requires predefined criteria for selecting among cross-indexed or extended publications.A pragmatic option is to retain the downloadable database version, or prefer a journal article over a conference paper when appropriate.
  • Dataset cleaning: Word clouds visualize term occurrence to inspect result-set appropriateness and identify unexpected keywords or papers for removal.They can support later concept classification, but candidates for exclusion require careful inspection because rarely used terminology may identify relevant papers.

2.3 Study Selection

Study selection systematically identifies relevant publications from potentially very large result sets using planned criteria, reviewer procedures, and voting approaches. The guideline contrasts majority voting with relative ratings and emphasizes documenting decisions and agreement for reproducibility.

  • Study selection planning: Study selection requires careful planning because result sets may contain hundreds or thousands of publications.Relevant factors include the number and distribution of researchers, topic familiarity, schedules, infrastructure, relevance criteria, and voting procedures.
  • Study selection planning: Reviewers independently rate cleaned result sets after agreeing on inclusion criteria, voting procedures, and schedules.A kick-off meeting establishes the procedure before each reviewer receives and rates the cleaned result set.
  • Majority voting: Majority voting uses reviewer headcounts to include or exclude papers, with unresolved papers referred to a third reviewer or a voting workshop.Overlapping subsets can collect multiple votes in one run, reducing the amount each reviewer evaluates.
  • Relative ratings: Relative ratings replace binary votes with Likert scales, using measures such as the mean or mode and requiring extra handling for neutral ratings.The example scale ranges from 5 points for highly relevant to 1 point for absolutely irrelevant, with 3 marking neutral or no opinion.
  • Reliability and documentation: Selection reliability should be assessed through documented rating statistics, explicit inclusion criteria, and inter-rater agreement measures such as Cohen’s κ or Fleiss’ κ.Relative ratings can be more differentiated than headcounts but are more complex and had not yet been applied to a complete study by the authors.

2.4 Concluding and Handover to Data Analysis

After collection, cleaning, and selection, the resulting artifacts are assembled and handed over to the in-depth data analysis. The guideline identifies reports and artifacts as deliverables that integrate with research protocols.

  • Handover to data analysis: The final step assembles outputs from earlier stages for the in-depth analysis dictated by the research questions and secondary-study type.These outputs are shipped to the analysis phase as a consolidated handover.
  • Handover to data analysis: An exemplary search and selection report documents the study’s search and selection activities.The report is presented as an artifact to be created during the early stages of the literature study.

3 Example Studies and Lessons Learned

The authors draw on prior systematic reviews and mapping studies to present interconnected practices for planning, searching, collecting, cleaning, and selecting literature. Their lessons emphasize practical dependencies, reproducibility, scope decisions, infrastructure, and quality assurance.

  • Example studies and lessons learned: The guideline synthesizes practices from prior systematic reviews and mapping studies, presenting them as reusable building blocks with context-dependent combinations and constraints.The authors relate lessons learned to studies in Table 5 and provide a blueprint of beneficial practice combinations.
  • General planning: A concrete research protocol and early researcher involvement establish shared interpretations of concepts and classification criteria.The protocol is described as central because literature studies involve substantial effort, duration, and multiple researchers.
  • Technical infrastructure: Version control supports baselines and distributed concurrent work, while mixing spreadsheet applications can create practical inconsistencies.The authors report that version control helps prevent accidental overwriting of results.
  • Search strings and search engines: Integrated search strings can improve precision, whereas multiple shorter queries can bypass database limits and accommodate database-specific syntax.The trade-off concerns query complexity, database limitations, and the need to customize searches across engines.
  • Search strings and search engines: 215 rather than 125 papers were found when testing a replicated search in Scopus, illustrating that evolving indexes can make old result sets irreproducible.The authors recommend reporting search timestamps and storing queries and raw result sets to improve transparency and reproducibility.
  • Data collection and cleaning: Restricting searches to selected venues may reduce overhead but risks information loss, while result-set size alone cannot establish coverage or quality.The authors recommend study-specific scope decisions, validated protocols, appropriate search strings, and detailed inclusion and exclusion criteria.
  • Study selection: Inter-rater agreement supports quality assurance and transparency, but its value depends on the voting procedure and the effort required.The authors emphasize its use in multi-staged procedures while noting limited usefulness with iteratively incomplete result sets.
  • Data collection and cleaning: Social network graphs can be generated early from search results to support pre-selection and highlight cooperation cliques.The graph provides an overview of subjects and their relationships before the main study begins.

4 Related Work and Discussion

Existing literature-study guidelines provide valuable structures and terminology but remain too generic for direct practical application. This article complements them with an experience-based, streamlined approach focused on early-stage decisions, data collection, selection, and reusable operational guidance.

  • Contribution: The guideline addresses study preparation, data collection, classification, and selection while complementing established systematic review and mapping approaches.It positions itself as an extension of existing guidelines rather than a replacement for them.
  • Approaches: Existing guidelines offer common structures and terminology but are often too generic for direct practical application.They provide an umbrella over fine-grained methods, models, advice, and best practices.
  • Experiences: Reviews of existing practice identified gaps including missing advice on self-evaluation, research-question justification, and shared experiences of applying specific guidelines.Petersen et al. compared 10 guidelines, while other work found substantial time demands and quality-assessment difficulties.
  • Tools: Growing literature-study size and complexity make tool support increasingly important for collecting, managing, and evaluating data.Prior work highlights requirements for planning and teamwork, while the article adds practical workflow and procedure guidance.
  • Contribution: The article contributes an experience-based guideline that streamlines early-stage practices and links them into a practical sequence.It emphasizes concrete advice on what information to collect, how to configure descriptive data, and how to run voting procedures.

5 Conclusion

The article concludes that systematic literature studies are valuable but difficult, with success depending substantially on study design, researcher expertise, and reproducible data collection. Its experience-based guideline offers pragmatic workflows and tool-oriented advice, while the authors also identify the need for more fine-grained guidance and sophisticated support.

  • Conclusion: Systematic literature studies structure reported knowledge but require substantial effort and expertise from the researchers involved.The authors connect study success to research-question appropriateness, design accuracy, and reproducible data collection.
  • Conclusion: The main challenges arise during initial data collection and study organization rather than only during later analysis.This scope motivates the guideline’s emphasis on early-stage practical decisions.
  • Contribution: The authors provide an experience-based guideline, practical tool experiences, and reusable blueprint workflows for literature-study design and reporting.The workflows are intended to support young scholars and make method descriptions more efficient.
  • Limitations and future needs: The authors’ studies relied mainly on simple tools such as spreadsheets and plain-text files, despite the rich data available for more extensive support.They therefore identify sophisticated tool support as an ongoing need.

A Study Workflow Templates

The appendix presents reusable workflow templates for literature studies with different team sizes and selection procedures. The templates connect preparation, searching, data cleaning, researcher coordination, and study selection into practical sequences.

  • A Study Workflow Templates: The appendix provides workflow templates inferred from experience for reuse in scientific-paper method descriptions.Each template includes context, an exemplary workflow, and a textual description.
  • A.1 Template 1: 2 Researcher Workshop Model with Snowballing: The 2 Researcher Workshop Model with Snowballing targets small studies with two collaborating researchers and limited selection-procedure options.The authors report applicability for up to approximately 50 papers and for senior or mixed senior-junior pairs.
  • A.1 Template 1: 2 Researcher Workshop Model with Snowballing: Its workflow begins with a snowballing-based preliminary study, uses extracted keywords to construct database queries, and then defines sources and inclusion or exclusion criteria.The preliminary study provides reference papers for incremental snowballing and search-strategy development.
  • A.1 Template 1: 2 Researcher Workshop Model with Snowballing: The two-researcher workflow cleans collected datasets, then uses a kick-off meeting to align criteria, prepare rating, and schedule selection activities.The meeting closes data collection and cleaning while initiating study selection.
  • A.2 Template 2: 3 Researcher Voting-only Model: The 3 Researcher Voting-only Model targets three-person teams using voting-based selection and is described as applicable to many literature-study settings.The model supports mixed and distributed teams but requires at least one senior researcher.
  • A.2 Template 2: 3 Researcher Voting-only Model: Its workflow defines queries, sources, and criteria before data collection and dataset cleaning.The cleaned datasets are integrated stepwise before rating begins.
  • A.2 Template 2: 3 Researcher Voting-only Model: After cleaning, two nominated researchers independently rate the integrated dataset, after which their results are integrated and unresolved items are prepared for further decision-making.The model uses the third researcher within the voting procedure for final decisions.

B Recommended Data Structure

The recommended data structure provides a minimal foundation for storing literature-search data and can be extended with study-specific classifications, inclusion or exclusion documentation, and dynamic metadata. The appendix also illustrates how voting data can be organized for collaborative selection.

  • B Recommended Data Structure: The article recommends a data structure for storing data obtained through manual or automatic literature searches.The recommendation is based on data structures used across several literature studies.
  • B Recommended Data Structure: The voting spreadsheet example organizes combinations of a three-person majority vote involving two reviewers and one extra reviewer for final decisions.The example is color-coded to distinguish voting combinations.
  • B Recommended Data Structure: The minimal data structure should be extended for mapping studies with classification schemas and documented inclusion or exclusion criteria.Extensions include generic or reused facets, study-specific facets, and reasons for each inclusion or exclusion decision.
  • B Recommended Data Structure: The appendix identifies the recommended data structure as a minimal starting point rather than a complete schema for every study.Its contents should be adapted to the study’s scope and research questions.
  • B Recommended Data Structure: Dynamic metadata can be added during the study to enhance the dataset, with Study and Context recommended as baseline dimensions.These dimensions support later classification and analysis of papers.
  • B Recommended Data Structure: The Study dimension captures a paper’s overall research approach and methods, supporting more detailed classification by research and contribution type.Examples include primary, replication, and secondary studies, as well as interview research and grounded-theory analyses.
Loading 1612.03583v1…