Source-linked AI summary
Fairness and Accountability Design Needs for Algorithmic Support in High-Stakes Public Sector Decision-Making
Michael Veale, Max Van Kleek, Reuben Binns
TL;DR
Algorithmic support in government operates amid hidden discretion, crossed accountability chains, changing data, and context-specific institutional constraints. The paper identifies design opportunities and argues that fairness and accountability interventions must be tested in real public-sector settings through close interdisciplinary collaboration.
Problem
Public-sector algorithmic systems can reproduce historical bias, while institutional constraints and discretionary spaces complicate how public values are exercised.
Method
The paper examines fairness and accountability challenges through practitioner experiences and stresses testing proposed interventions in domain-specific organisational contexts.
Results
The analysis identifies changing data, discretion, human augmentation, practice transmission, and nuanced performance communication as key areas for design and interdisciplinary collaboration.
Takeaways & Limitations
Workable fairness and accountability tools should address managers and frontline decision-makers while incorporating domain knowledge and practical public-sector constraints.
Takeaways & Limitations
Model behaviour can be altered by decision-subject responses, producing feedback effects that may reduce performance or create unforeseen failures.
Abstract
from arXiv · showhide
Calls for heightened consideration of fairness and accountability in algorithmically-informed public decisions---like taxation, justice, and child protection---are now commonplace. How might designers support such human values? We interviewed 27 public sector machine learning practitioners across 5 OECD countries regarding challenges understanding and imbuing public values into their work. The results suggest a disconnect between organisational and institutional realities, constraints and needs, and those addressed by current research into usable, transparent and 'discrimination-aware' machine learning---absences likely to undermine practical initiatives unless addressed. We see design opportunities in this disconnect, such as in supporting the tracking of concept drift in secondary data sources, and in building usable transparency tools to identify risks and incorporate domain knowledge, aimed both at managers and at the 'street-level bureaucrats' on the frontlines of public service. We conclude by outlining ethical challenges and future directions for collaboration in these high-stakes applications.
INTRODUCTION
Algorithmic systems are increasingly used in public-sector decisions while facing concerns about opacity, bias, and accountability. The paper examines whether fairness and accountability tools address the organisational and contextual realities practitioners encounter.
- Public-sector machine learning systems have drawn criticism for being inscrutable and reproducing biases in historical training data.
- Fairness and transparency tools are often developed in isolation from specific users and deployment contexts.
- Practitioners already face immediate value-laden challenges while deploying in-house and vendor-supplied machine learning systems.
- The study interviewed 27 actors across 5 countries’ public sectors, spanning modelling, project management, procurement, and contractor roles.Applications included taxation, child protection, policing, justice, emergency response, and interior security.
- The paper organises interview themes around internal and external users, stakeholders, and decision subjects before discussing interdisciplinary design challenges.
BACKGROUND AND MOTIVATION
Public-sector algorithmic decision-making involves contested values, institutional constraints, and discretionary practices that are difficult to capture through laboratory-oriented research. The paper therefore grounds its analysis in interviews examining how value-laden concerns appear in practice.
- Public policy problems lack the settled agreement on ends and means often assumed in engineering-oriented decision-support research.Predictive policing illustrates competing interpretations of fairness, such as maximising total arrests versus treating crimes equally across areas.
- Public-sector IT projects are resource constrained, path-dependent, and vulnerable to poor initial scoping, while accountability crosses organisational scales.
- Public-service values are exercised in discretionary spaces between rules, and ethical issues may surface only after design, procurement, or deployment choices are embedded.
- The study addresses limited research on how value-laden concerns manifest and are handled in real public-sector decision-support settings.It seeks empirical foundations for a more grounded approach rather than validating young, primarily unverified theoretical frameworks.
- The research draws on open-ended interviews with 27 individuals, recruited through ad hoc sampling across five OECD countries.Interviews involved public servants and contractors working mainly in modelling or project management; conversations were analysed through open coding and a public-sector values framework.
FINDINGS
The findings are organised around two broad categories: internal actors connected to algorithmic systems and external actors such as decision subjects.
- Interview themes concern both internal actors and their relationships with algorithmic systems, and external actors including decision subjects.
Internal actors and machine learning–driven decisions
Internal users need practical ways to understand models, their inputs, and their performance. Organisational buy-in and responsible use therefore depend on transparency, interpretable metrics, and alignment with users’ knowledge and expectations.
- Internal actors require different approaches for clarifying the workings and processes of machine-learning decision-support systems.These actors include strategic managers who commission external contractors or receive models from internal teams.
- Getting individual and organisational buy-in: Organisational pressure for explanation led practitioners to favour more transparent models, including logistic regression and random forests over more complex alternatives.Explaining model logic was associated with better internal buy-in and clearer communication with business users.
- Getting individual and organisational buy-in: Transparency was also pursued by reducing input variables and comparing models with and without sensitive variables so clients could judge appropriateness.One contractor reduced 18,000 variables to 8 and presented the resulting options alongside their performance trade-offs.
- Getting individual and organisational buy-in: Imbalanced outcomes can make accuracy appear near-perfect while precision remains substantially lower and more meaningful to users.A collision-risk modeller reported 100% apparent accuracy from 40 million records but only 20% precision, while users understood a 1/5 accident probability more readily.
- Getting individual and organisational buy-in: Users may judge models by perceived additional insight, efficiency, or inclusion of preferred indicators rather than performance metrics alone.Automated maps could appear disappointing because they resembled analysts’ maps, even while saving several days of work.
Over-reliance, under-reliance and discretion
Practitioners treated model outputs as decision support whose appropriate influence varied by domain, with professional judgement and user discretion shaping adoption and use.
- The paper distinguishes between under-reliance and over-reliance as recurring concerns when linking algorithmic outputs with professional judgement.
- Interfaces let officers review predictions alongside their own intuition rather than simply follow model instructions.
- Model influence varied by domain: some operators were expected to follow directions, while others resisted losing discretion over decisions.
- Ethics guidance instructed police staff to work down rank-ordered lists while preserving professional judgement as the final authority.
Augmenting models with additional knowledge
Practitioners augmented model outputs with contextual and qualitative knowledge, while confronting feedback loops, organisational incentives, legal change, and risks from excessive transparency.
- Intelligence officers enriched predictive maps with local news, reports, and current operational knowledge before decisions were made.
- A human-trafficking model began identifying car washes after increased intelligence attention there, illustrating how new information can redirect future predictions.
- Legal changes can alter the populations entering prison systems, complicating robust recidivism modelling and requiring awareness, communication, and preparation.
- The paper notes that internal gaming and organisational responses receive less attention despite their relevance to value-laden algorithmic decisions.
- Releasing model inputs and weightings could lead auditors to investigate according to perceived model structure rather than actual outputs.
- Auditors simultaneously use decision support and collect future training data, creating incentives that may distort when fraud cases are recorded.
External actors and machine learning–driven decisions
External actors and organisational differences shape how machine-learning systems are deployed, adapted, and understood across public-sector settings.
- Smaller or poorer public organisations often rely on importing and adapting ideas because they cannot absorb the costs and risks of experimentation.
- Model deployment requires investment in working processes, and insufficiently qualified staff may end up performing interpretive roles.
- Pre-trained vendor models may transfer poorly across jurisdictions when purchasing organisations lack capacity to understand, modify, or augment them appropriately.
- Concerns about external actors therefore include both model portability and the internal capacity needed for responsible adaptation.
Accountability to decision subjects
Accountability to decision subjects requires explanations that connect model outputs to individual administrative decisions, while fairness concerns include protected attributes and problematic proxy outcomes.
- Interpretable models were valued when administrative decisions needed explanation to customers or tribunals.
- A tax agency built client-specific explanations after a model correctly flagged deductions that had often been erroneously allowed previously.
- Police organisations found it difficult to discuss equity and accountability with officers focused narrowly on where they could catch someone.
- Practitioners generally avoided directly using protected characteristics, although some faced pressure to include them for predictive power.
- Using conviction as a proxy for offending can produce models that reproduce differences in conviction rates rather than identify offending itself.
Gaming by decision-subjects
Public-sector models can be altered by deliberate or incidental responses to their outputs, changing future data and degrading performance. Practitioners therefore face challenges monitoring shifting data and maintaining accountability across interconnected systems.
- External actors may probe predictive systems for loopholes, making gaming a practical concern in public-sector deployments.
- A caller’s repeated reports made the model infer that children always went missing at 10pm, requiring manual removal of the spurious pattern.
- Deployment responses can create feedback loops: policing patterns contributed to displacement and diffusion, causing prediction accuracy to collapse during trialling.
- Models can change future training data by influencing resource allocation, potentially concentrating collection in already overrepresented areas.
- Managers may miss data shifts because routine monitoring and evaluation reveal only part of the operational picture.
- Concept drift detection offers a possible response, but detecting relevant distribution shifts in complex real-world settings remains difficult.
‘Always a person involved’: Augmenting outputs
Interviewees described decision-support as augmenting rather than replacing professional judgment, with human intervention remaining distributed across the workflow. This makes fairness depend on contextual reliance, discretionary practices, and interventions beyond model training.
- ‘Always a person involved’: Augmenting outputs: Practitioners commonly combine algorithmic outputs with contextual data and manual analysis rather than allowing systems to replace institutional knowledge.
- ‘Always a person involved’: Augmenting outputs: Fairness interventions can occur in training data, model outputs, or the stage between generated results and their dissemination.
- ‘Always a person involved’: Augmenting outputs: Knowledge elicitation is proposed as a way to constrain and augment patterns learned from data, including in routine rather than only exceptional cases.
- ‘Always a person involved’: Augmenting outputs: Trust and reliance on decision-support vary with users, tasks, stakes, and contexts rather than following a single general pattern.
- ‘Always a person involved’: Augmenting outputs: Systematically following or ignoring advice may treat people with protected characteristics differently, so deployment cannot rely solely on bias removal and enforced obedience.
- ‘Always a person involved’: Augmenting outputs: Staff may use discretionary workarounds to inject fairness, making ethical practice partly dependent on how people adapt systems in operation.
‘I’m called the single point of failure’: Moving practices
Fairness and accountability depend on social practices that are difficult to formalize, preserve, and transfer when key personnel move or organisations share models. The paper also argues that technically rigorous metrics must remain practically intelligible and relevant to domain users.
- ‘I’m called the single point of failure’: Moving practices: Statistical fairness methods and transparency interfaces address different needs from the social detection and organisational responses practitioners commonly rely on.
- ‘I’m called the single point of failure’: Moving practices: Maintaining and transferring these practices is difficult because public-sector expertise is vulnerable to staff turnover and uneven organisational resources.
- ‘I’m called the single point of failure’: Moving practices: Informal knowledgebases and virtual communities could help share ethical issues, but relying on spontaneous collaboration is risky in resource-scarce or competitive settings.
- ‘I’m called the single point of failure’: Moving practices: Performance metrics can be hard to explain beyond accuracy, especially for continuous regression and multiple-classification tasks.
- ‘I’m called the single point of failure’: Moving practices: Practitioners valued preferred risk indicators, continuity with prior analysis, and interpretability, sometimes over additional accuracy.
- ‘I’m called the single point of failure’: Moving practices: Metrics should be discussed with both statistical rigour and practical relevance, requiring domain-specific work across interfaces, visualisation, statistics, and knowledge elicitation.
CONCLUDING REMARKS
The paper argues that fair and accountable algorithmic decision-support must be developed with practitioners and affected stakeholders, while confronting institutional and political realities. It calls for real-world, interdisciplinary research that can navigate contested values and practical constraints.
- Researchers should not assume public-sector practitioners are naïve about fairness and accountability challenges in algorithmic decision-support.The paper links this concern to the participatory-design and action-research commitments of HCI and information systems.
- Future collaboration should address changing data, discretion, model-output augmentation, transmission of social practices, and communication of nuanced performance aspects.
- Proposed interventions must be stress-tested in real situations because domain-specific, organisational, and contextual factors shape fairness and accountability efforts.The paper highlights institutional constraints, high stakes, and crossed lines of accountability in public-sector settings.
- Research should move from studying systems in vitro toward examining them in vivo within messy socio-technical contexts involving political, institutional, and infrastructural constraints.The paper identifies political winds, technical lock-in, and ageing infrastructure as conditions interventions must confront.
- Meaningful impact requires close collaboration among disciplines, practitioners, and affected stakeholders to develop workable social and technical improvements.The paper frames these systems as increasingly involved in choices concerning governmental power and vulnerable groups’ rights and freedoms.