Source-linked AI summary

Fairness Hazard Analysis for Socio-Technical Processes: A Multiple-Case Study in Bias-sensitive Organisational Settings

Giovanna Broccia, Lucio Lelii, Roberto Cirillo, Dario Di Nucci, Samuel Fricker, Fabio Palomba, Giorgio O. Spagnolo, Alessio Ferrari

arXiv:2608.22978v1cs.SE

TL;DR

The paper addresses the limited availability of requirements-level methods for identifying and mitigating fairness hazards across socio-technical processes. It refines Fairness Hazard Analysis (FHA) for real organisational workflows and evaluates it through a qualitative multiple-case study. Practitioners found the identified hazards relevant and most mitigations appropriate, while feasibility depended on contextual resources and reusable patterns required adaptation.

  • Problem

    Requirements engineering lacks systematic methods for identifying and mitigating fairness hazards across socio-technical processes before fairness issues accumulate into systemic bias.

  • Method

    The paper refines FHA for real organisational workflows and evaluates it through workflow modelling, independent hazard analysis, and practitioner validation in two recruitment processes.

  • Results

    Practitioners confirmed the relevance of identified fairness hazards and the appropriateness of most proposed mitigations across two organisational recruitment cases.

  • Takeaways & Limitations

    FHA supports collaborative, iterative fairness analysis, with reusable mitigation patterns such as independent review and collective decision-making adapted to organisational context.

  • Takeaways & Limitations

    The study covers two opportunistically selected recruitment organisations, so its findings support analytical rather than statistical generalisation, and mitigations were not implemented or evaluated.

Abstract

from arXiv · show

Fairness is increasingly recognised as a first-class requirement in socio-technical processes, where interactions among human actors, software systems, and AI technologies may lead to unfair outcomes in decision-making workflows. If left unaddressed, fairness hazards may accumulate and reinforce systemic bias, highlighting the need to engineer fairness proactively. Despite growing interest in fairness-aware systems, systematic methods for identifying fairness hazards in socio-technical processes and deriving requirements-level mitigations remain limited. To support fairness-by-design during requirements engineering (RE), Fairness Hazard Analysis (FHA) is introduced as a methodology for systematically identifying, analysing, and mitigating fairness hazards. FHA is first assessed through a proof-of-concept validation conducted via two focus groups. Then, a qualitative multiple-case study involving two organisations examines its applicability in real-world settings. The proof-of-concept validation highlighted the benefits derived from the structured nature of the method, and suggested the need to include iterative, dialogic reflection with domain experts. In the multiple case-study where FHA was applied, the practitioners involved were positively impressed by the results and confirmed the relevance of the identified fairness hazards (spanning up to 27% of the process elements), as well as the appropriateness of most of the proposed mitigations, while noting that contextual factors might hinder their implementation. The evaluation also highlighted mitigation patterns, such as independent review and collective decision-making, which can be transferred to different organisations. This paper contributes a structured and empirically validated methodology for integrating fairness considerations in RE and preventing systemic bias in socio-technical processes.

1 Introduction

The paper addresses the limited support for engineering fairness in socio-technical processes by refining Fairness Hazard Analysis (FHA) for real organisational workflows and evaluating it across two organisations.

  • Fairness issues can arise from human actors, organisational procedures, software, AI, and their interactions, potentially propagating into systemic bias.
  • Requirements engineering lacks operational methods for identifying, analysing, and mitigating fairness risks across socio-technical processes.
  • FHA models processes and supports analysis of fairness hazards, their consequences, propagation, impact, likelihood, and requirements-level mitigations.
  • The study refines FHA for real organisational workflows and evaluates it through a qualitative multiple-case study involving two organisations.
  • The evaluations progress from a constructed AI-assisted recruitment exemplar to recruitment workflows embedded in real organisational settings.
  • The paper extends prior FHA work by incorporating workflow elicitation, clarifying organisational and expert roles, and characterising recurring and context-dependent hazards.

2 Background

The background frames fairness as a contextual property of socio-technical processes and connects bias propagation with fairness debt and hazard-analysis principles.

  • Fairness concerns equitable treatment throughout a socio-technical process, assessed within its organisational and social context rather than only within technical components.
  • Unfairness may arise from AI, organisational procedures, subjective criteria, unequal information access, inconsistent human decisions, and their interactions.
  • Fairness debt describes latent socio-technical liabilities created when fairness issues are not explicitly managed throughout the software lifecycle.
  • Interdependent fairness-debt causes can propagate across lifecycle stages and increase risks of systemic inequities, reputational damage, and regulatory non-compliance.
  • The paper uses fairness-debt root causes as analytical prompts while requiring detailed process knowledge from organisational members and analysts.
  • Hazard analysis anticipates harmful conditions, evaluates causes and consequences, and designs preventive or corrective controls.
  • FHA adapts the qualitative, top-down structure of Preliminary Hazard Analysis to identify, trace, and mitigate fairness hazards across socio-technical processes.

3 Related Work

Related work offers fairness-aware lifecycle and requirements approaches, but explicit translation of fairness risks into actionable controls across human–AI workflows remains limited.

  • Values@Runtime and ReFair operationalise stakeholder values or fairness-requirement elicitation during system development and operation.
  • Existing work indicates that fairness remains a secondary quality attribute and that safety methods can uncover social and ethical risks in machine-learning systems.
  • Goal-oriented requirements approaches model stakeholder goals and conflicts but provide limited support for fairness risks that emerge, propagate, and accumulate across socio-technical processes.
  • FHA combines empirical, organisational, regulatory, ethical, domain, and stakeholder evidence to derive traceable requirements-level mitigations for unfair outcomes.

4 Fairness Hazard Analysis Methodology

FHA is an iterative, collaborative methodology that models socio-technical processes, identifies and assesses fairness hazards, and plans traceable requirements-level mitigations.

  • System Definition: FHA begins by defining and modelling the socio-technical process, its actors, and their interactions.
  • Fairness Hazard Identification: Analysts examine process steps, actors, decision points, and interactions using empirical, organisational, regulatory, domain, stakeholder, and fairness expertise.
  • Fairness Hazard Assessment: Each hazard is assessed for consequences, propagation, impact, likelihood, and risk through collaborative review and consensus.
  • Fairness Hazard Assessment: Impact ranges from none to high, while likelihood ranges from rare to systemic, enabling prioritisation of fairness risks.
  • Fairness Hazard Assessment: FHA records justified differentiation for transparency while distinguishing it from undesirable bias requiring mitigation.
  • Mitigation Planning: Mitigation planning modifies or adds workflow nodes and controls, including human review or consensus at high-impact decisions.
  • Iteration: FHA revisits hazards and mitigations as the socio-technical process evolves or new empirical evidence emerges.

5 FHA Proof-of-Concept Evaluation

The proof-of-concept evaluation used two focus groups to assess FHA’s comprehensibility and usefulness with a constructed AI-assisted recruitment workflow. Participants found FHA structured and useful, while recommending stronger contextualisation, analytical guidance, mitigation support, and iterative reflection.

  • Evaluation setup: FHA was evaluated with two focus groups using a constructed, simplified AI-assisted recruitment workflow.The workflow included data ingestion, AI prescreening, human review, and auditing, and FHA identified hazards across human and technical components.
  • Evaluation setup: Twelve academics participated, with varied prior fairness expertise and representation across two groups.Participants included 33.4% women and 66.7% men; 58% reported no or basic expertise and 42% intermediate to advanced expertise.
  • Findings: Participants generally perceived FHA as structured, understandable, and useful for identifying fairness issues throughout a socio-technical process rather than only within AI components.They particularly valued the method’s structured sequence and its coverage of organisational, human, software, and AI-related sources of unfairness.
  • Methodological refinements: Participants recommended contextualising fairness interpretations and distinguishing undesirable bias from intentional or contextually justified differentiation.They also suggested structured prompts, examples, and checklists while recognising that predefined categories cannot ensure discovery of unknown or context-dependent hazards.
  • Methodological refinements: Participants recommended clearer impact and likelihood definitions, calibration guidance, domain-specific examples, mitigation templates, and documented mitigation rationales.They also valued independent reviews, decision checkpoints, and multi-person consensus for high-impact decisions.
  • Methodological refinements: Participants framed fairness analysis as iterative organisational reflection requiring hazards, assumptions, and mitigations to be revisited as contexts evolve.The proof-of-concept was formative and limited by academic participants and a single constructed case rather than real organisational workflows.

6 FHA Operationalisation in the Multiple-Case Study

The multiple-case study operationalised FHA through an analyst-led, practitioner-informed procedure applied separately to validated organisational recruitment workflows. The procedure reconstructed and validated workflows, independently analysed fairness hazards, planned requirements-level mitigations, and reviewed findings with practitioners.

  • Operationalisation: FHA was operationalised through an analyst-led and practitioner-informed procedure applied separately to each organisation.Organisational participants supplied process knowledge and validated workflows, while two researchers formally conducted the FHA analysis.
  • Phase 1: Workflow elicitation: Phase 1 elicited each organisation’s recruitment workflow from participants with direct process knowledge.The sessions also considered AI outputs, human supervision, transparency needs, and overreliance where AI-supported activities were present or envisaged.
  • Phase 2: Workflow reconstruction: Phase 2 reconstructed workflows as UML activity diagrams representing activities, decisions, roles, technologies, sequences, and information flows.Activity and decision nodes captured process interactions, while control and object flows represented activity order and exchanged information, artefacts, and decisions.
  • Phase 3: Workflow validation: Phase 3 required organisational participants to review and correct reconstructed workflows before FHA was applied.Participants checked missing steps, unclear responsibilities, inaccuracies, and incomplete information flows.
  • Phase 4: FHA application: In Phase 4, two analysts independently examined validated workflow steps for fairness hazards and preliminary mitigations.This phase operationalised FHA Steps B–D, including hazard identification, analysis of consequences and propagation, risk assessment, and mitigation planning.
  • Phase 4: FHA application: Phase 4 produced fairness-hazard lists, hazard-analysis and risk-assessment tables, and linked requirements-level mitigation strategies.Mitigations could modify procedures, decision criteria, responsibilities, review mechanisms, software-supported activities, or their interactions.
  • Phase 5: Feedback and evaluation: Phase 5 presented hazards and mitigations incrementally to practitioners, who assessed realism, omissions, feasibility, and organisational barriers.The feedback process reviewed one fairness hazard at a time in relation to the relevant workflow activities and actors.

7 Multiple-Case Study Design

The study used a holistic qualitative multiple-case design involving two real recruitment processes in contrasting organisational settings. It combined workflow elicitation, validation, analyst FHA application, and practitioner feedback to examine hazards, mitigations, and perceptions of FHA.

  • Study design: The study comprised two holistic cases, each analysing a real recruitment process within its organisational context.The cases were selected for maximum variation while retaining formal hiring processes with multiple actors and decision points.
  • Scope boundary: The study did not provide detailed calibration guidelines or evaluate the reliability and soundness of impact and likelihood ratings.Systematic calibration and empirical evaluation of Step C were deferred to later validation at TRL 6 and TRL 7.
  • Research questions: The research objective was to evaluate FHA’s applicability in organisational environments through questions about emerging hazards, mitigation strategies, and practitioner perceptions.The study addressed both the types of hazards and mitigations identified and FHA’s usefulness, clarity, and applicability.
  • Analysis: Each workflow was analysed separately and then compared through cross-case synthesis to identify recurring and context-specific hazard and mitigation categories.Practitioner feedback was likewise analysed within cases and synthesised thematically across cases.
  • Participants: The study used organisational participants to provide process knowledge, validate reconstructed workflows, and evaluate FHA outputs.The same participants contributed across phases, supporting continuity but leaving a limitation because the small number of participants might hinder complete workflow reconstruction.
  • Case selection: The cases contrasted a well-established global technology manufacturer with a young public higher-education institution.Their differing structures, governance, decision-making practices, and supporting technologies enabled examination across distinct organisational contexts.
  • Data collection: Data collection comprised workflow elicitation, process clarification, FHA application, and practitioner feedback for each case.Elicitation used a two-hour semi-structured individual interview in Organisation A and a two-hour semi-structured group interview in Organisation B, followed by email-based clarification and validation.

8 Results of the Case Study

The case study applied FHA to two organisational recruitment processes, identifying fairness hazards across both workflows and deriving context-sensitive mitigation strategies. Cross-case analysis found recurring hazard and mitigation categories alongside organisation-specific concerns.

  • Organisation A: Organisation A’s validated workflow contained 15 fairness hazards, identified by analysts and retained after consensus, affecting approximately 27% of distinct process elements.The hazards spanned governance, access to opportunities, recruitment content and channels, candidate evaluation, interviews, and selection.
  • Organisation A: Organisation A mitigation proposals were developed to reduce the likelihood and/or impact of unfair outcomes by modifying the recruitment workflow.The revised process diagram incorporated the proposed mitigation strategies.
  • Organisation B: Organisation B’s validated workflow contained 14 fairness hazards, with analyst-specific findings consolidated after discussion and consensus.The hazards covered governance, professorship definition, exceptional recruitment, job descriptions, recruitment content and channels, and candidate assessment.
  • Organisation B: Organisation B’s mitigation strategies addressed all identified hazards through accountability, transparency, independent review, standardisation, diversified channels, AI oversight, and committee safeguards.The controls included multi-perspective review of AI-supported activities and safeguards for balanced committee discussion and candidate ranking.
  • Cross-Case Synthesis: Eight hazard categories recurred across both organisations, while representation-format bias, institutional agenda-setting bias, and collective decision-making bias were context-specific.Recurring categories included governance, access, job requirements, AI-generated content, informational advantages, subjective evaluation, procedural inconsistency, and workload or order effects.
  • Cross-Case Synthesis: Recurring mitigation categories included governance, independent review, transparency, equal access, standardised evaluation, AI oversight, and workload management.The analyses frequently recommended explicit responsibilities, additional review, standardised consequential evaluations, and documented decision rationales.
  • Cross-Case Synthesis: FHA combined preventive, detective, and corrective controls, indicating that recruitment fairness mitigation requires coordinated organisational, procedural, human, and technological controls.Participants interpreted hazard realism in relation to actual workflows, possible conditions, and safeguards already present in each organisation.

9 Discussion

The cross-case findings show that FHA supports context-sensitive identification and validation of fairness hazards while revealing more reusable mitigation patterns. Practitioners also found that FHA can surface overlooked concerns and clarify existing safeguards, although implementation depends on organisational conditions.

  • Cross-case findings: Fairness hazards recurred across organisations but were also shaped by organisation-specific procedures, exceptions, and decision mechanisms.Practitioners therefore treated hazard realism as a contextual judgement rather than a binary property.
  • Cross-case findings: Additional risks emerged outside the formally modelled workflow, showing that exceptions, informal practices, and surrounding interactions require explicit examination and organisational validation.A reusable hazard catalogue can support analysis but cannot replace detailed workflow examination.
  • Mitigations: Mitigation strategies converged more than hazards, including accountability, independent review, transparency, structured evaluation, human oversight, and safeguards for decision-making.These interventions can be organised as reusable fairness-by-design patterns, subject to contextual adaptation.
  • Mitigations: Practitioners generally regarded mitigations as useful or feasible, but cost, workload, existing safeguards, candidate burden, and organisational practices constrained transferability.Acceptance of a hazard did not automatically imply acceptance of its proposed mitigation.
  • Reflective function: FHA prompted reflection on generative AI, alternative tools, and routine practices that already functioned as fairness safeguards.Its reflective function may support organisational learning by connecting familiar procedures with explicit fairness objectives.
  • Implications: The integrated evidence supports FHA at TRL 5 through operationalisation on real workflows, hazard identification, mitigation generation, and practitioner-informed adaptation.The analysis should involve organisational members throughout and combine process knowledge with fairness and methodological expertise.

10 Lessons Learned

The lessons learned emphasise contextual, collaborative fairness analysis rather than reliance on fixed taxonomies or context-free hazard lists. They also identify reusable safeguards while cautioning that AI-related risks and analytical tools require continued human and empirical scrutiny.

  • Contextual analysis: Conceptual bias frameworks are useful prompts, but detailed process knowledge is necessary because taxonomies are not exhaustive hazard catalogues.Frameworks should sensitise analysis rather than replace contextual investigation.
  • Analysis team: Domain familiarity should be complemented by informed outsiders who can question practices and assumptions treated as self-evident.Further studies should examine how analyst number and profiles affect hazard discovery.
  • Risk interpretation: Fairness analysis must distinguish unfair bias from justified differentiation and constrained residual risk arising from role responsibilities, legal obligations, or legitimate organisational interests.Lower-cost safeguards can provide incremental protection, but should not substitute for stronger interventions when hazard severity requires them.
  • Workflow representation: Visual and spatial workflow representations supported systematic inspection, while the study did not establish that UML diagrams are inherently understandable without modelling expertise.The notation and alternatives require empirical examination.
  • Reuse: Recurring hazard categories may support a preliminary library, but such a resource should stimulate investigation rather than function as a complete checklist.Reuse remains dependent on contextual application.
  • AI in FHA: AI may be both a source of fairness hazards and part of their mitigation, requiring human judgement, traceability, and comparison with other evidence.Repeated reliance on the same model or prompting approach may reproduce assumptions or stereotypes, while biased AI judgements can influence and amplify human bias.

11 Threats to Validity

The validity analysis limits interpretation to applicability and perceived relevance of FHA in two reconstructed recruitment workflows. Threats concern analyst interpretation, incomplete process reconstruction, researcher and participant perspectives, and restricted generalisability.

  • Construct validity: Fairness hazards are context-dependent analytical constructs whose formulation may reflect analysts’ interpretations, available process information, and adopted assumptions.Independent analysis and organisational review were used to mitigate this construct-validity threat.
  • Construct validity: The proposed mitigations were assessed for organisational relevance and feasibility but were not implemented or evaluated for effects on recruitment decisions.The findings therefore concern FHA applicability and perceived output relevance, not control effectiveness.
  • Internal validity: Reconstructed workflows may omit activities, safeguards, exceptional paths, or informal practices despite participant review and clarification.Organisation B feedback exposed an existing control and an informal situation that had not been captured correctly.
  • Internal validity: Researcher involvement in FHA development and participants’ responsibility for the processes may have encouraged favourable assessments.The feedback also excluded candidates and other stakeholders affected by recruitment decisions.
  • External validity: The two opportunistically selected organisations are not statistically representative, so findings should not be generalised through statistical inference.The supported form of generalisation is case-based analytical generalisation, with transferability most plausible for comparable formal decision-making processes.

12 Conclusion

The study refines FHA for real organisational workflows and advances its validation from TRL 3 to TRL 5. Across two recruitment cases, FHA revealed recurring and context-specific risks, produced generally relevant mitigations, and encouraged fairness reflection, but mitigation effectiveness remains unevaluated.

  • Contribution: FHA was operationalised through workflow modelling, independent hazard analysis, and iterative practitioner validation, advancing the methodology from TRL 3 to TRL 5.
  • Findings: FHA identified potential bias sources in software, AI, organisational procedures, human judgement, and interactions among process activities and roles.Cross-case findings combined recurring patterns with organisation-specific risks.
  • Findings: Practitioners considered the hazards largely relevant and most mitigations appropriate, while noting that implementation feasibility depends on available resources.The analysis also surfaced underexamined concerns and clarified fairness-related roles of existing practices.
  • Implications: Recurring mitigation patterns included independent review, explicit decision criteria, documented rationales, accountability mechanisms, and collective decision-making.These patterns are preliminary and require adaptation to organisational context.
  • Limitations and future work: The study did not establish whether proposed mitigations improve actual fairness outcomes.Future work should implement and longitudinally evaluate mitigations, include affected stakeholders, and study additional organisations, domains, and regulatory contexts.
Loading 2608.22978v1…