Source-linked AI summary
A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications
Despoina Giarimpampa, Roland Meier, Tegawendé F. Bissyandé, Vincent Lenders, Jacques Klein
TL;DR
Expert surveys are difficult to conduct when cybersecurity practitioners are unavailable, creating a need to assess whether LLMs can serve as surrogate participants. This paper evaluates LLM-generated responses across individual, aggregate, and temporal survey settings, finding stable outputs that fail to capture the variability and diversity of real expert judgements.
Problem
Cybersecurity expert surveys often face low participation and small samples, while limited guidance exists on when LLM-generated surrogate responses are reliable.
Method
The paper proposes a reproducible framework evaluating persona-based, aggregate, and multi-year LLM survey responses across methodological settings.
Results
LLM responses are internally stable but show reduced variance, central tendency bias, homogenised opinions, and limited capture of expert judgement diversity.
Takeaways & Limitations
LLMs are useful for piloting, hypothetical scenarios, and research-design support, but should not replace human expertise in security surveys.
Takeaways & Limitations
The study uses zero-shot prompting, leaving multi-shot and multi-turn prompting outside its scope.
Abstract
from arXiv · showhide
Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts, exhibiting reduced variance, central tendency bias and homogenised opinions. This work contributes methodological evidence and practical guidance to the security research community on the appropriate use and limitations of LLM-generated survey responses. We conclude that LLMs are useful for piloting and hypothesis generation but not for replacing expert elicitation, and we discuss implications for researchers using LLM-augmented surveys.
1 Introduction
Cybersecurity research depends heavily on expert surveys and interviews, but low response rates, small samples, recruitment barriers, and confidentiality constraints make practitioner evidence difficult to obtain. This paper proposes a reproducible framework for evaluating LLMs as surrogate survey participants across individual, aggregate, and temporal settings, while examining their reliability and risks to validity.
- Motivation: Cybersecurity studies rely heavily on human expert judgement, including surveys and interviews with SOC analysts and other practitioners.This dependence spans incident response, decision-making, tooling adoption, and organisational maturity research.
- Motivation: Low response rates, small expert samples, recruitment barriers, and confidentiality constraints make cybersecurity survey evidence fragile or unavailable.Online surveys have a weighted mean response rate of approximately 44%, 11–12 percentage points below other survey modes.
- Motivation: LLMs could help pilot instruments, explore counterfactual populations, and expand limited datasets, but synthetic responses may hallucinate, exaggerate agreement, or lack grounding in practice.Prior evaluations also caution that models can smooth disagreement and fail to reproduce human-like response variation.
- Approach: The paper proposes a general, reproducible framework for responsibly using and evaluating LLMs as surrogate participants in expert cybersecurity surveys.It examines how prompt design, sampling, and aggregation shape the credibility of synthetic responses.
- Approach: The evaluation compares persona-based expert simulation, aggregate survey-distribution replication, and temporal robustness using multi-year SOC survey data.The empirical comparison includes real SOC experts and LLM-generated respondents, alongside stability and robustness analyses.
- Implications: The paper identifies reduced variance, central tendency bias, homogenised opinions, and loss of expert variability as systematic risks to valid practitioner representation.These failure modes motivate responsible limits on synthetic respondents in human-centred security research.
2 Related work
Prior research shows promise for AI-generated survey responses and expert-reasoning simulation, but remains fragmented and offers limited evidence or guidance for validating synthetic responses. In cybersecurity, this gap is especially pronounced, motivating a reproducible framework for evaluating LLM surrogate participants against human experts.
- Cross-domain foundations: AI and LLMs have been explored for survey design, response synthesis, analysis, and simulated expert reasoning across social sciences, psychology, and software engineering.These applications aim to generate plausible survey responses, strengthen questionnaire studies, or support predictive tasks.
- Methodological limitations: Existing work provides limited evidence that plausible LLM-generated survey responses align with real expert judgements or can be responsibly validated and benchmarked.The literature therefore offers limited methodological guidance for using synthetic responses in survey-based research.
- Cybersecurity gap: Cybersecurity research has identified LLM applications and initial surrogate-expert studies, but lacks structured expert evaluation and scalable alternatives remain difficult when Delphi studies or interviews are impractical.This gap is described as more pronounced in cybersecurity despite surveys compiling hundreds of LLM-based technical applications.
- Study positioning: Unlike prior cybersecurity work on computational detection or SOC automation, this study examines whether LLMs can serve as surrogate participants for security-survey knowledge elicitation.The study addresses expert-domain reliability through reproducible evaluation emphasizing stability, alignment with human experts, and responsible use.
3 Methodology
The methodology formalises LLM-based survey simulation through structured representation, prompt conditioning, repeated sampling, aggregation, and evaluation. It is designed to support systematic, reproducible assessment of stability, dispersion, and alignment between synthetic and human responses across multiple survey settings.
- Framework scope: The framework queries LLMs as simulated individual experts, aggregated populations, and surveys spanning multiple years.These modes support individual, aggregate, and temporal analyses.
- Survey and prompt design: Survey questions use standardised categorical, multiple-choice, or Likert-scale representations, while prompts define the task, response format, and optional persona or distributional context.Prompts provide no example answers, preserving zero-shot independence and exposing models’ inductive biases.
- Sampling strategy: Because LLM outputs are stochastic, each model is queried repeatedly under identical conditions to estimate stability, dispersion, and run-to-run variability.The framework proposes ten independent samples per model in each setup, motivated by stability gains plateauing after the first 5–10 generations.
- Aggregation: Repeated outputs are aggregated with question-type-specific rules, including plurality for categorical items, option-set or majority rules for multiple-choice items, and medians for Likert scales.These rules produce comparable responses while respecting the structure of each question type.
- Evaluation dimensions: The framework enables systematic, reproducible evaluation by revealing instability, reduced variability, and misalignment between synthetic and human responses across individual, aggregate, and temporal settings.Its evaluation dimensions include inter-human agreement, intra-LLM consistency, inter-LLM agreement, and LLM–human alignment.
4 Experimental design
The experimental design evaluates LLM surrogate experts in three settings—persona simulation, aggregate distribution replication, and temporal robustness—using common prompting, sampling, aggregation, and evaluation procedures. Six LLMs are tested against survey responses from six SOC professionals, with question-type-specific aggregation and agreement metrics.
- Evaluation settings: The framework compares persona-based expert simulation, aggregate distribution replication, and temporal robustness against practitioner survey data.These settings respectively assess individual expert reproduction, population-level distribution matching, and temporal alignment beyond models’ training horizons.
- Model selection: Six LLMs are evaluated: GPT-4o, GPT-4, DeepSeek, Llama 3.1–8B, Llama 3.2–3B, and Gemini Flash.GPT-3.5-turbo is added in the temporal setting because its 2021 training cutoff provides a contrast for temporal generalisation and training-data exposure.
- Human survey: The human comparison population comprises six cybersecurity professionals with diverse roles, expertise levels, industries, and organisational contexts.Participants completed a structured SOC AI-adoption survey covering deployment, workflow automation, operational metrics, and adoption barriers.
- LLM prompting and sampling: Each model simulates six expert personas across ten independent runs, producing 60 simulated responses per model.Prompts consistently combine prompt sandwiching, persona-based framing, and demographic-specific cues without examples, chain-of-thought reasoning, fine-tuning, or few-shot demonstrations.
- Aggregation and evaluation: Responses are aggregated by question type using plurality voting for categorical items, option-level rules for multiple-choice items, and persona-level medians for Likert items.Multiple-choice aggregation uses exact option-set mode, per-option majority, or a Bayesian Beta(1, 1) rule with a 0.5 threshold; Likert means and standard deviations are descriptive summaries.
- Aggregation and evaluation: Agreement and variability are assessed with question-type-appropriate metrics for intra-group, inter-group, and LLM–human comparisons.Categorical responses use exact match rate, while multiple-choice responses use Jaccard and Hamming similarity.
5 Results
The results show that LLM-generated survey responses are internally stable but weakly aligned with diverse human expert responses. Across models and time, simulations reproduce model-specific distributions rather than human variation, limiting their use to exploratory analysis and hypothesis generation.
- Human–LLM response patterns: Human experts showed substantial dispersion and no single consensus, whereas LLM responses were more moderate, generalised, and centrally biased.LLMs often followed generic best practices even with distinct personas, capturing averaged expert knowledge rather than differentiated operational perspectives.
- Reproducibility: Repeated sampling produced only partial overlap in individual responses, although fixed model–prompt combinations generated nearly identical aggregate distributions.This indicates that stochastic decoding does not explain the divergence between LLM-generated and human survey distributions.
- Cross-model agreement: Inter-LLM agreement was uniformly low, and LLM–human similarities were likewise low or usually zero, providing no evidence of meaningful convergence toward expert responses.Observed similarities are consistent with largely independent, noisy response patterns across models rather than shared reasoning or latent expert structure.
- Implications: Persona-conditioned simulations are neither fully reproducible nor human-equivalent, so they are better suited to exploratory analysis and hypothesis generation than replacing individual expert judgements.Individual simulations should be interpreted as stochastic samples from model-specific response distributions.
- Absolute alignment: Pairwise similarity suggested broadly similar distributional shapes, but chi-square evaluation revealed large absolute mismatches with the human baseline for most systems.Models could reproduce global response shape while over-allocating probability mass to rare options, so pairwise closeness did not imply faithful proportions.
- Temporal alignment: Across survey years, LLM–human alignment remained consistently low, with intra-LLM divergence lowest, inter-LLM divergence higher, and LLM–human divergence highest.Year-to-year changes were small and non-monotonic, indicating stable but time-insensitive distributions without meaningful tracking of historical practitioner responses.
6 Limitations
The study’s limitations concern the rapid evolution of LLMs, the absence of comprehensive cybersecurity ground truth, and the use of zero-shot prompting. These constraints define opportunities for future research while preserving the framework’s methodological relevance.
- Limitations: Rapid LLM evolution may require the framework to adapt, although its model-agnostic evaluation targets properties that remain relevant across versions.The framework evaluates methodological properties rather than depending on specific model versions.
- Limitations: Cybersecurity’s lack of definitive ground truth complicates evaluating AI–expert alignment, but comparison with real experts remains practical for assessing surrogate suitability.The study does not seek absolute correctness; it evaluates whether LLMs are suitable surrogates for experts.
- Limitations: The study uses zero-shot prompting to examine models’ inductive biases without human examples or prompt-specific optimisation, leaving multi-shot and multi-turn prompting outside its scope.Providing task-relevant context while avoiding example answers allows evaluation of default model behaviour; alternative prompting may produce more stable or grounded outputs.
7 Ethical considerations
The study addresses ethical risks involving human participation, misrepresentation of synthetic expertise, bias, privacy, and misuse. It frames LLMs as supplementary, exploratory tools requiring transparent reporting and human-centred validation rather than replacements for expert elicitation.
- Human participants: Human participation received IRB approval, was voluntary and uncompensated, collected no personally identifiable information, and reported demographics only in aggregate.The study reduced participant risk by focusing on professional practices rather than sensitive personal information.
- Human expertise: LLM outputs may appear authoritative and be mistaken for expert opinion, so the study rejects using them as substitutes for expert elicitation.The authors frame LLMs as supplementary tools and maintain that human expertise remains indispensable for decision-making.
- Representation and bias: Synthetic responses can obscure minority viewpoints, reduce variability, reinforce dominant perspectives, and mislead when treated as representative of real experts.The study calls for careful validation, transparent reporting, and clear communication that synthesized responses do not represent consensus or population-wide views.
- Privacy and confidentiality: Using LLMs with sensitive prompts may disclose confidential information through third-party services, requiring safeguards and strict avoidance of sensitive inputs.Although this study used no sensitive operational data, the authors stress careful data handling and transparency in applied contexts.
- Responsible application: Direct operational use of LLM insights could cause harm, so practical applications require human-in-the-loop validation and outputs must be disclosed, limited, and non-authoritative.The study also acknowledges possible dual-use and equity risks, including attacker strategies, deceptive content, and marginalisation of non-Western viewpoints.
8 Conclusion
The framework shows that LLM-generated survey responses are stable and reproducible but do not reliably capture practitioner variability, limiting their validity as substitutes for human participants. LLMs remain useful for piloting, hypothetical scenarios, and stress-testing when human expertise and rigorous human-centred validity are preserved.
- LLM responses were internally stable and reproducible across individual, aggregate, and temporal settings but did not reliably capture practitioner variability.
- Synthetic responses reflect model-specific patterns rather than authentic practitioner perspectives, limiting their validity as substitutes for human participants.
- LLMs’ controllability and scalability support piloting survey instruments, exploring hypothetical scenarios, and stress-testing research designs.
- The framework guides responsible integration of LLMs into survey studies while preserving human expertise and rigorous human-centred validity.
Funding
The authors report that the work received no specific financial support.
- No specific financial support was received for this work.
A Survey instrument
The appendix presents the survey instrument used to study AI adoption among cybersecurity professionals in SOC environments. It covered five thematic sections and generally used multiple-choice or Likert-scale closed questions.
- Survey scope: The survey targeted cybersecurity professionals working in Security Operations Centre environments.It was used in the expert study on AI adoption in SOCs.
- Survey structure: The instrument comprised five sections: demographics, SOC characteristics, processes and tooling, AI adoption, and professional perceptions.
- Question formats: Closed questions generally used multiple-choice or Likert-scale formats unless otherwise stated.
A.1 Demographics
The demographics section records participants’ roles, years of experience, training, and willingness to disclose information.
- A.1 Demographics: Participant roles included Security Analyst (Tier 1), Security Investigator (Tier 2), and Product/Project Manager.
- A.1 Demographics: The survey asked about years of experience in the participant’s role.
- A.1 Demographics: Demographic response options included “Prefer not to say” and “Technical/Vocational training.”
A.2 SOC characteristics … A.7 Tool usage
The survey captures SOC characteristics, operational workload, AI use, perceived challenges, professional attitudes, and security-tool usage. Its measures combine factual questions, categorical responses, challenge ratings, and 1–5 Likert-scale assessments.
- A.2 SOC characteristics: SOC characteristics include team size, a preference-not-to-say response option, and whether the SOC operates 24/7.
- A.3 SOC processes and workload: SOC processes and workload cover daily alert volume, capabilities present, and the level of automation, including fully automated rule-based functions.
- A.5 Challenges and perceptions: The survey asks respondents to rate how challenging specified SOC-related issues are on a 1–5 scale.
- A.6 Professional perceptions: Professional perceptions use 1–5 Likert-scale statements about trust, productivity, security effectiveness, risks, and whether human analysts outperform AI.
- A.7 Tool usage: Tool usage includes threat intelligence platforms as a surveyed SOC technology.