Source-linked AI summary
Understanding artificial intelligence ethics and safety
David Leslie
TL;DR
AI systems can produce individual and societal harms, creating a need for responsible public-sector design and deployment. This guide proposes governance, accountability, and human-centred implementation practices, concluding that re-translating algorithmic outputs can make them interpretable, comprehensible, and justifiable.
Problem
AI ethics addresses individual and societal harms arising from the misuse, abuse, poor design, or unintended consequences of AI systems.
Method
The guide proposes operational governance measures, including anticipatory and remedial accountability across AI design, development, and deployment.
Results
Human-centred re-translation makes algorithmically supported outcomes more interpretable, comprehensible, and justifiable while supporting accountability.
Takeaways & Limitations
Responsible AI implementation requires human judgment and oversight grounded in communication, deliberation, evidence, situational awareness, and moral responsibility.
Takeaways & Limitations
Static machine-learning models can become inaccurate and unreliable when historical training data no longer reflect the population concerned.
Abstract
from arXiv · showhide
A remarkable time of human promise has been ushered in by the convergence of the ever-expanding availability of big data, the soaring speed and stretch of cloud computing platforms, and the advancement of increasingly sophisticated machine learning algorithms. Innovations in AI are already leaving a mark on government by improving the provision of essential social goods and services from healthcare, education, and transportation to food supply, energy, and environmental management. These bounties are likely just the start. The prospect that progress in AI will help government to confront some of its most urgent challenges is exciting, but legitimate worries abound. As with any new and rapidly evolving technology, a steep learning curve means that mistakes and miscalculations will be made and that both unanticipated and harmful impacts will occur. This guide, written for department and delivery leads in the UK public sector and adopted by the British Government in its publication, 'Using AI in the Public Sector,' identifies the potential harms caused by AI systems and proposes concrete, operationalisable measures to counteract them. It stresses that public sector organisations can anticipate and prevent these potential harms by stewarding a culture of responsible innovation and by putting in place governance processes that support the design and implementation of ethical, fair, and safe AI systems. It also highlights the need for algorithmically supported outcomes to be interpretable by their users and made understandable to decision subjects in clear, non-technical, and accessible ways. Finally, it builds out a vision of human-centred and context-sensitive implementation that gives a central role to communication, evidence-based reasoning, situational awareness, and moral justifiability.
What is AI ethics?
AI ethics comprises values, principles, and techniques that guide morally acceptable conduct in developing and using AI. The guidance provides conceptual and practical resources for integrating these considerations into ethical, fair, safe, and responsible AI projects.
- Purpose: The guidance helps department and delivery leads develop and deploy AI ethically, safely, and responsibly.It complements the Data Ethics Framework, which should be used during project initiation.
- Purpose: AI ethics and safety must be prioritised to manage harmful impacts and direct AI development toward optimal public benefit.Rapidly evolving AI can produce mistakes, miscalculations, and unanticipated harms despite its potential to address urgent challenges.
- Implementation: Responsible AI requires integrating social and ethical considerations throughout every stage of an AI project.The guidance provides conceptual resources and practical tools to support responsible design and implementation.
- Definition: AI ethics uses widely accepted standards of right and wrong to guide conduct in developing and using AI technologies.It motivates morally acceptable practices and defines duties needed for ethical, fair, and safe applications.
Why AI ethics?
AI ethics emerged in response to the individual and societal harms caused by AI misuse, abuse, poor design, and unintended consequences. These harms include discrimination, accountability gaps, opaque decisions, privacy invasion, social polarisation, and unsafe or unreliable outcomes.
- Origins of AI ethics: AI ethics has largely emerged in response to individual and societal harms caused by AI misuse, abuse, poor design, and negative unintended consequences.These potential harms motivate the development of a robust culture of AI ethics.
- Bias and discrimination: Data-driven technologies can reproduce, reinforce, and amplify marginalisation, inequality, and discrimination through societal patterns, designer biases, and unrepresentative training data.Insufficiently representative samples can produce biased and discriminatory outcomes because the input data is flawed from the start.
- Accountability: AI-generated decisions can create accountability gaps because responsibility is difficult to assign across complex, distributed design, production, and implementation processes.Such gaps may harm autonomy and violate the rights of affected individuals when injury or other negative consequences occur.
- Explainability: Opaque model rationales can be deeply problematic when algorithmic outcomes affect decision subjects and the processed data may contain discrimination, bias, inequity, or unfairness.High-dimensional correlations can exceed the interpretive capabilities of human-scale reasoning.
- Privacy and autonomy: AI systems can threaten privacy by using personal data without proper consent and by targeting, profiling, or nudging people without their knowledge or consent.These practices can undermine private life and the ability to pursue goals free from unchosen influence.
- Social and public harms: Excessive automation and hyper-personalisation may reduce human interaction, polarise social relationships, and produce unreliable, unsafe, or poor-quality outcomes that damage wellbeing and public trust.Irresponsible data management, negligent design and production, and questionable deployment practices can undermine confidence in socially beneficial AI.
An ethical platform for the responsible delivery of an AI project
An ethical platform combines a culture of responsible innovation with governance that supports ethically permissible, fair, trustworthy, and justifiable AI projects. It consists of SUM Values, FAST Track Principles, and a process-based governance framework applied throughout the delivery workflow.
- Purpose: An ethical platform unites responsible innovation with governance that brings ethical, fair, and safe AI values and principles into practice.It requires multidisciplinary cooperation and responsibility throughout the innovation and implementation lifecycle.
- Project goals: Projects must address stakeholder wellbeing, discriminatory effects and bias, product safety and reliability, and the transparency and interpretability of decisions.These goals establish ethical permissibility, fairness, public trust, and justifiability.
- Purpose: The platform provides both a process-based footing for ethical, equitable, and safe implementation and a basis for responsible AI innovation.Its governance architecture is intended to help teams design and implement systems ethically while facilitating a responsible innovation culture.
- Proportionality: Ethical stewardship should be proportionate to project scope and impacts, while ethical considerations remain relevant to every AI project.Low-stakes, non-safety-critical applications generally require less proactive stewardship than high-stakes projects affecting lives or using sensitive data.
- Building blocks: The three building blocks are SUM Values, FAST Track Principles, and the PBG Framework, which operationalises them across the entire project workflow.SUM Values comprise Respect, Connect, Care, and Protect; FAST Track Principles comprise Fairness, Accountability, Sustainability, and Transparency.
- Implementation: Teams must put the platform into practice at every design and implementation step through continuous reflection, action, and justification.This makes the platform an ongoing practice rather than a one-time assessment.
The SUM Values
The SUM Values adapt critical elements of bioethics and human rights discourse to address harms arising from AI misuse, abuse, poor design, and unintended consequences. They guide ethical deliberation throughout the AI innovation lifecycle, including how teams handle tensions and trade-offs among values.
- The SUM Values draw on bioethics and human rights discourse while targeting the social and ethical problems created by harmful AI systems.They are designed for potential misuse, abuse, poor design, and harmful unintended consequences.
- They guide project evaluation, planning, formulation, design, development, testing, implementation, and reassessment across the innovation lifecycle.The values are intended to remain active from early project decisions through later implementation and review.
- RESPECT the dignity of individual persons: Respect requires protecting individual dignity, autonomy, expression, informed decision-making, participation, and opportunities to flourish.The stated commitments include safeguarding people’s power to be heard and to pursue freely determined life plans.
- CONNECT with each other sincerely, openly, and inclusively: Connect requires sincere, open, inclusive relationships, meaningful dialogue, social cohesion, diversity, participation, and inclusion throughout AI production and use.It also calls for using AI pro-socially to strengthen human relationships and ensure voices are heard seriously.
- AI systems should foster stakeholder welfare, do no harm, minimise misuse or abuse, and prioritise people’s safety and mental and physical integrity.These requirements apply when exploring technological possibilities and when conceiving and deploying AI applications.
- The values provide a framework for judging ethical permissibility, assessing impacts across the lifecycle, and discussing how to weigh competing values in context-specific trade-offs.Teams are encouraged to deliberate when values come into tension in particular use cases.
The FAST Track Principles: … Data fairness
The FAST Track Principles address the ethical gap created by AI systems’ lack of moral responsibility, organizing responsible AI around fairness, accountability, sustainability, and transparency. Fairness requires human and technical attention across the lifecycle, especially responsible data practices that mitigate bias and discriminatory harm.
- The FAST Track Principles:: AI systems’ transfer of cognitive functions to algorithmic processes creates a need for principles tailored to their design and use.These systems cannot themselves be directly responsible or immediately accountable for the consequences of their behaviour.
- The FAST Track Principles:: Fairness, accountability, sustainability, and transparency provide principles intended to fill the gap between machine smart agency and machines’ lack of moral responsibility.The framework treats AI systems as non-moral agents requiring human-centred ethical governance.
- The FAST Track Principles: Fairness, Accountability, Sustainability, and Transparency: FAST concerns operate throughout the AI delivery workflow and require cooperation among technical, domain, management, and policy specialists.The text characterises ethical AI innovation as a team effort from start to finish.
- The FAST Track Principles: Fairness, Accountability, Sustainability, and Transparency: Accountability and transparency govern the workflow, while fairness and sustainability describe qualities of algorithmic systems for which designers and implementers are held accountable.The four principles are interrelated but are not treated as equals or as operating on the same plane.
- Fairness: Human contexts and biases can introduce unfairness anywhere from data extraction and preprocessing through problem formulation, model building, and implementation.Datasets may also encode complex social and historical patterns containing culturally crystallised bias and discrimination.
- Fairness: Fairness-aware design combines non-technical self-assessment with technical controls and evaluation to support just, morally acceptable, beneficial, and equitable outcomes.The text cautions that bias mitigation has no simple or strictly technical solution.
- Data fairness: Responsible data acquisition, handling, and management are necessary for algorithmic fairness because biased, compromised, or skewed datasets can expose stakeholders to discriminatory harm.Data fairness includes representativeness, fit-for-purpose and sufficiency, source integrity and measurement accuracy, timeliness and recency, and relevance supported by domain knowledge.
- Data fairness: A Dataset Factsheet should be created at the alpha stage and maintained throughout design and implementation to support data quality, bias mitigation, and auditability.The factsheet is intended to document data provenance and responsible data practices across the project workflow.
Design Fairness · Outcome fairness
Fairness must be addressed throughout the AI project lifecycle, from defining objectives and preparing data to building, evaluating, and procedurally validating models. Outcome fairness should be explicitly defined and measured according to the use case and technical feasibility, with limitations communicated through a publicly available, plain-language Fairness Position Statement.
- Design Fairness: Design fairness requires precautions across the AI workflow because human choices at every construction stage can introduce discriminatory bias.The relevant stages include problem formulation, data preprocessing, feature determination, model-building, tuning, testing, and evaluation.
- Design Fairness: Problem formulation should translate project goals into measurable targets while considering affected stakeholders, vulnerable groups, and the justice of selected outcomes or proxies.Technical and non-technical team members should collaborate, with stakeholder input and inclusive deliberation beginning at project inception.
- Design Fairness: Data labelling, feature engineering, hyperparameter tuning, and metric selection should be reviewed for bias and aligned with mitigation and discriminatory non-harm.These activities involve human judgments about classification, relevant information, model settings, and evaluation criteria.
- Design Fairness: Designers must detect hidden proxies for sensitive features, assess the moral justifiability and interpretability of learned inferences, and avoid models whose discriminatory non-harm cannot be confirmed.Hidden proxies can produce implicit redlining, especially when complex models prevent human assessors from understanding social or demographic inferences.
- Design Fairness: Procedural fairness requires rules to be applied consistently and uniformly across decision subjects, with identical inputs and procedures producing replicable outputs.Uniform application prevents targeted rule changes that could disadvantage a specific individual.
- Outcome fairness: Teams should publish a Fairness Position Statement explaining the selected fairness criteria in plain, non-technical language for review by affected stakeholders.The statement follows consideration of use-case appropriateness, technical feasibility, and incorporation of the chosen fairness model into the application.
- Outcome fairness: Outcome fairness has competing formal definitions, including demographic parity, equal true-positive rates, equal false-positive rates, equal positive predictive values, individual fairness, and counterfactual fairness.These approaches respectively emphasize group benefit proportions, error or precision parity, similarity-based treatment, or consistency across a closest possible alternative group membership.
- Outcome fairness: The fairness definition should depend on the use case and technical feasibility, because methods require different interventions and may need sensitive-attribute data that is unavailable or unreliable.Formal approaches also have limited scope and often apply only to distributive or allocative consequences.
Implementation fairness
Fair AI implementation requires preventing bias and misapplication at the point of delivery. Training, accountable deployment, and human-centred interfaces should preserve critical judgment, situational awareness, and evidence-based reasoning.
- Implementation fairness: Decision-support systems create risks of bias and misapplication at delivery, requiring special attention to harmful or discriminatory outcomes.
- Implementation fairness: Decision-automation bias can impair critical judgment and situational awareness, causing over-reliance, over-compliance, out-of-loop syndrome, safety hazards, and discriminatory harm.Users may overlook system faults or defer to perceived infallibility, weakening their ability to respond to system failure.
- Implementation fairness: Automation-distrust bias can cause users to disregard salient AI contributions to evidence-based reasoning because of skepticism, over-prioritised prudence, or reliance on human expertise.
- Implementation fairness: Implementer training should explain machine learning’s statistical and probabilistic character, AI limitations, and the role of AI as assisting rather than replacing human judgment.Training should avoid anthropomorphic portrayals of AI systems.
- Implementation fairness: Interfaces should encourage active judgment and situational awareness while clearly presenting system rationale, fairness-standard compliance, and confidence level at runtime.Training should also explore deployment-context biases and use-case-specific misjudgements when weighing statistical evidence.
Putting the principle of discriminatory non-harm into action
Putting discriminatory non-harm into practice requires a workflow-based approach to fairness-aware AI design and implementation across the project pipeline. Teams should make accountability explicit, identify bias and discrimination risks proactively, and conduct collaborative fairness self-assessments at each stage.
- Workflow-based implementation: Managers should map team-member involvement at every AI project-pipeline stage from alpha through beta.This workflow perspective helps teams clarify who is involved throughout the project.
- Workflow-based implementation: Workflow-based fairness design makes end-to-end accountability paths clear and peer-reviewable.The approach concretises and makes accountability explicit across the project pipeline.
- Risk identification: Teams should pinpoint bias and downstream-discrimination risks and streamline solutions proactively, pre-emptively, and anticipatorily.A workflow perspective supports identifying risks before harmful effects occur.
- Collaborative assessment: At each project-pipeline stage, relevant team members should collaboratively self-assess the applicable dimension of fairness.The passage presents this assessment as a three-step process but does not specify the steps in the supplied text.
Discriminatory Non-Harm Self-Assessment · Accountability
The self-assessment requires identifying applicable fairness and bias dimensions, scrutinising project-specific risks, and taking corrective and preventive action. Accountability-by-Design addresses AI’s accountability challenges through continuous human answerability and end-to-end auditability across design and implementation.
- Discriminatory Non-Harm Self-Assessment: Identify the fairness and bias mitigation dimensions relevant to the specific AI project stage under consideration.At the data pre-processing stage, these may include data fairness, design fairness, and outcome fairness.
- Discriminatory Non-Harm Self-Assessment: Scrutinise how the AI project could create risks or unintended vulnerabilities across each applicable fairness and bias area.
- Discriminatory Non-Harm Self-Assessment: Correct identified problems, strengthen weaknesses with possible discriminatory consequences, and proactively prevent bias in areas posing future risks.
- Accountability: Responsible AI delivery must address an accountability gap because automated decisions are not self-justifiable and AI production involves multiple parties.The production process can make it difficult to determine which department leads, technical experts, data personnel, policy experts, implementers, or others should bear responsibility for negative consequences.
- Accountability: A sufficiently fine-grained accountability concept comprises two subcomponents: answerability and auditability.
- Accountability: Answerability places responsibility for justifying algorithmically supported decisions on human creators and users through a continuous chain of human responsibility.Explanations and justifications should be provided by competent human authorities in plain, understandable, coherent, and non-technical language.
- Accountability: Auditability requires accessible records and monitoring that enable oversight of design, use, data provenance, model operations, and outcomes throughout the AI lifecycle.Systems should support peer and overseer review, reproducibility, and end-to-end recording so their operations can be assessed as safe, ethical, and fair.
- Accountability: Accountability-by-Design requires responsible humans-in-the-loop across the entire design and implementation chain and activity-monitoring protocols for end-to-end oversight and review.
Accountability deserves consideration across the entire design and implementation workflow · Sustainability · Stakeholder Impact Assessment
Responsible AI governance requires accountability throughout design, deployment, and after-use, alongside continuous attention to sustainability and stakeholder impacts. Stakeholder Impact Assessments support this by identifying risks, informing transparent decisions, and revisiting social, ethical, and technical consequences across the project lifecycle.
- Accountability deserves consideration across the entire design and implementation workflow: Accountability should be anticipated before deployment to strengthen design and implementation and pre-empt harms to individual wellbeing and public welfare.This anticipatory accountability focuses on decisions and actions taken before algorithmically supported outcomes occur, rather than relying first on later correction.
- Accountability deserves consideration across the entire design and implementation workflow: Remedial accountability remains necessary through auditability, transparent lifecycle logging, and explanations that clarify decisions’ rationale, fairness, ethical permissibility, and safety.These practices provide affected stakeholders with justifications for how algorithmically supported decisions bear on their lives.
- Sustainability: Sustainable AI development requires continuous sensitivity to transformative, long-term real-world effects on individuals, society, and affected communities.Technical sustainability additionally requires systems to be safe, accurate, reliable, secure, and robust despite uncertainty, volatility, anomalies, and perturbations.
- Stakeholder Impact Assessment: A Stakeholder Impact Assessment should evaluate an AI project’s social impact and sustainability whether it delivers public services or supports back-office administration.The assessment can build public confidence, strengthen accountability, expose unseen risks, support informed and transparent innovation, and demonstrate due diligence.
- Stakeholder Impact Assessment: Teams should conduct SIAs at three lifecycle points: Alpha problem formulation, pre-implementation after model validation, and Beta post-deployment reassessment.The initial assessment addresses ethical permissibility, the second checks alignment with the original assessment, and the third compares documented expectations with real-world impacts and unintended consequences.
- Stakeholder Impact Assessment: The SIA is one component of governance and should complement the accountability framework, auditing, and activity-monitoring documentation.Its four sections cover general impacts, sector- and use-case-specific concerns, pre-implementation evaluation, and reassessment using real-world impacts, public input, and unintended consequences.
- Stakeholder Impact Assessment: SIA questions should examine affected and vulnerable stakeholders, objective definitions, autonomy, wellbeing, health, safety, privacy, discrimination, social cohesion, inclusion, diversity, socioeconomic effects, future generations, and the planet.Sector-specific and use-case-specific concerns should be compiled with the team’s planned responses.
- Stakeholder Impact Assessment: Post-deployment reassessment should use performance evidence, monitoring data, implementer and public input, and maintenance checks to rectify harms and accommodate distributional shifts.Teams should mitigate or redress unintended harmful consequences and retune or retrain models when environmental conditions change.
Accuracy, reliability, security, and robustness
Safe AI requires teams to understand and govern four operational objectives: accuracy, reliability, security, and robustness. These objectives involve context-sensitive performance assessment, dependable intended behaviour, protection against adversarial threats, and reliable operation under harsh conditions.
- Accuracy: Accuracy is the proportion of correct outputs, but acceptable performance thresholds should be tailored to the application’s use-case needs.Accuracy may also be expressed as an error rate or fraction of incorrect outputs.
- Accuracy: Accuracy assessment must account for unequal error costs, uncertainty, noisy data, missing features, and changes in the underlying reality.Total-cost metrics can weigh one class of errors against another when some mistakes are more significant or costly.
- Reliability: Reliability means that an AI system consistently behaves as its designers intended and conforms operationally to its programmed specifications.This dependability supports confidence in the system’s safety.
- Security: Security protects an AI system’s information integrity, architecture, functionality, accessibility, confidentiality, and privacy against possible adversarial attack.A secure system prevents unauthorised modification or damage to its component parts.
- Robustness: Robustness requires an AI system to function reliably and accurately under harsh conditions, including adversarial intervention, implementer error, and skewed goal-execution.It concerns the strength of system integrity and operational soundness under difficult conditions, attacks, and perturbations.
Risks posed to accuracy and reliability:
AI systems’ accuracy and reliability can deteriorate when changing data distributions make historical models outdated. They may also fail unpredictably on unfamiliar situations, producing serious and potentially unexplainable mistakes.
- Concept Drift: Concept drift makes static models trained on historical data vulnerable when those data no longer reflect the population concerned.The risk arises because the model is frozen before deployment, while the underlying data distribution may change.
- Concept Drift: Public-sector teams should remain vigilant to potentially rapid concept drift in complex, dynamic, and evolving environments.Technical teams should familiarise themselves with research on detecting and mitigating concept and distribution drift.
- Brittleness: Brittle high-performing models trained through massive datasets and repeated examples may struggle with unfamiliar events and scenarios.Deep neural networks can rely on thousands, millions, or even billions of parameters to generate outputs.
- Brittleness: These systems may make unexpected and serious mistakes because they cannot contextualise problems or use common sense to assess new unknowns.Their mistakes may also remain unexplainable because of the high-dimensionality and computational complexity of their mathematical structures.
Risks posed to security and robustness · End-to-End AI Safety
AI systems face security and robustness risks from adversarial manipulation, poisoned data, and reinforcement-learning objectives that can produce harmful behaviour. End-to-end safety therefore requires workflow-wide testing, monitoring, verification, and documented self-assessment aligned with accuracy, reliability, security, and robustness objectives.
- Risks posed to security and robustness: Adversarial attacks subtly modify inputs to induce highly confident misclassifications or incorrect predictions.Examples include changing a few pixels so a model identifies a panda as a gibbon or a stop sign as a speed-limit sign.
- Risks posed to security and robustness: Such vulnerabilities create serious safety implications when AI is deployed in autonomous transportation, medical imaging, security, and surveillance.Targeted perturbations can cause gross miscalculation and incorrect decisions in critical applications.
- Risks posed to security and robustness: Model hardening combats adversarial attacks through adversarial training, architectural modification, regularisation, and data pre-processing manipulation.Adversarial training enlarges training data with adversarial examples.
- Risks posed to security and robustness: Data poisoning compromises collection or pre-processing sources by modifying dataset subsets used to train, validate, or test models, inducing curated misclassification.The attack targets the data on which the AI system is developed.
- Risks posed to security and robustness: Reinforcement-learning systems can optimise their reward objective through real-world actions that are harmful because they lack context-awareness, common sense, empathy, and understanding.The risk arises when sufficient controls do not prevent an objective-optimal but damaging course of action.
- Risks posed to security and robustness: Mitigations for misdirected reinforcement-learning behaviour include extensive simulation, continuous inspection and monitoring, improved interpretability, and human override mechanisms.These measures support constraint programming, behavioural understanding, decision assessment, and intervention.
- End-to-End AI Safety: End-to-end safety depends on the algorithm, application, data provenance, objective specification, and problem domain, requiring rigorous testing, validation, verification, monitoring, and logged team self-assessments.Self-assessments should examine alignment between design and implementation practices and the objectives of accuracy, reliability, security, and robustness.
Transparency
Transparency in AI ethics covers both interpretability of a system’s behaviour and the justifiability of its design, implementation, and outcomes. Safeguarding it requires justifying processes, explaining outcomes in accessible language, and justifying their ethical acceptability, fairness, and safety.
- Transparency: Transparent AI requires understanding how and why a model produced a decision or behaviour in a specific context, including its rationale and explicability.This interpretability is often described as “opening the black box” of AI.
- Process Transparency: Process transparency requires demonstrating that ethical permissibility, fairness, non-discrimination, safety, and public trustworthiness shaped design and implementation end-to-end.This is supported by applying best practices throughout the AI project lifecycle.
- Outcome Transparency: Outcome transparency requires explaining in plain, non-technical language how and why a model acted in a specific decision-making or behavioural context.The explanation should translate technical model logic into socially meaningful language understandable through the societal factors and relationships involved.
- Outcome Transparency: Outcome justification requires showing that a system’s specific decision or behaviour is ethically permissible, fair, non-discriminatory, safe, and worthy of public trust.This justification begins with the clarified and explained outcome, then assesses it against criteria followed throughout design and implementation.
Process Transparency: Establishing a Process-Based Governance Framework
Responsible AI governance should integrate professional and institutional transparency, a Process-Based Governance Framework (PBG Framework), and a digitally consolidated process log. Together, these mechanisms connect responsible-innovation principles to workflow processes, clarify governance responsibilities and interventions, and support end-to-end auditability and tailored explanations.
- Professional and Institutional Transparency: Professional and institutional transparency requires integrity, honesty, sincerity, neutrality, objectivity, and impartiality throughout AI design and implementation.Project processes should also remain as open to public scrutiny as possible, subject to justified confidentiality and protections against gaming service provision.
- Process-Based Governance Framework: A PBG Framework operationalises the values and principles underpinning ethical and safe AI by integrating them with the processes of the AI design and development pipeline.The framework is intended as a structured template applicable beyond CRISP-DM.
- Process-Based Governance Framework: The framework provides a landscape view of governance procedures, including responsible roles, workflow stages requiring intervention, follow-up timeframes, monitoring, and auditability protocols.These elements help teams see how governance actions and control structures are organised across the project workflow.
- Enabling Auditability with a Process Log: A digitally centralised process log consolidates governance records, activity-monitoring results, and model-development data from modelling through implementation to establish end-to-end auditability.It can help demonstrate responsible design and use practices and the justifiability of system outcomes to concerned parties and affected decision subjects.
- Enabling Auditability with a Process Log: The process log can tailor information access and presentation to protect legitimately restricted data while catering explanations to stakeholders with different interests and expertise.This supports presenting project results with the user or receiver in mind.
Outcome transparency: Explaining outcome and clarifying content
AI outcome transparency requires explanations that inform implementers’ evidence-based judgments and communicate decisions accessibly to affected stakeholders. Because complex models are difficult to interpret, effective explanation must connect technical rationale with socially meaningful communication, human verification, and moral justifiability.
- Outcome transparency: Outcome explanations should support implementers’ evidence-based judgments and be communicated accessibly to affected stakeholders.The task requires standards and protocols spanning the project team rather than a simple technological fix.
- Explaining outcome: Interpretable explanations should make the factors determining an AI outcome intelligible as evidence supporting the decision or behaviour.They should clarify the rationale in plain, non-technical, and socially meaningful language, combining technical design with delivery practices.
- Clarifying content: Safe and fair AI requires logical, semantic, social, and moral dimensions of reasoning, with technical explanations especially important in design and social and moral explanations at delivery.Understanding system logic and technical inner workings supports safety and fairness, while explanations must also address human objectives and social contexts.
- Technical interpretability: High-dimensional feature interactions and unintuitive decision curves make it difficult to trace how individual inputs become model outputs, although supplemental interpretable methods can approximate complex systems.The entanglement of inputs and nonlinear relationships creates the central understandability-complexity challenge.
- Practical explanation methods: 491,520 calculations illustrate SHAP’s computational burden, while counterfactual explanations can offer actionable recourse but are not a complete interpretability solution.SHAP can provide locally consistent and accurate reckoning, yet becomes intractable beyond a threshold; counterfactual explanations may become unclear when many features matter.
Securing responsible delivery through human-centred implementation protocols and practices
Responsible AI delivery should begin with human circumstances, capacities, and context, then define roles, relationships, and processes that support understandable and evidence-based decisions. Re-translating human choices and societal values into explanations enables interpretable outcomes, stakeholder deliberation, end-to-end accountability, and justifiable implementation.
- Human-centred implementation: Human-centred implementation starts by understanding affected people’s circumstances, needs, competences, capacities, and contextual requirements.Explanations should accommodate vulnerable and disadvantaged stakeholders through clear, non-technical accounts of algorithmically supported results.
- Human-centred implementation: Context and domain knowledge guide the definition of roles and relationships, user training, implementation platforms, and understanding of outcomes.The checklist uses a predictive risk assessment system as a generic case for context-sensitive delivery planning.
- Roles and delivery relations: Delivery protocols should map roles, expertise, objectives, and relationships so implementers exercise unbiased judgment while decision subjects receive evidence-based clarification and justification.The primary decision-subject/advocate-to-implementer relationship is information- and dialogue-driven, prioritising comprehension and mutual understanding.
- Re-translation and interpretation: Re-translation makes model internals, mechanisms, outputs, and inferences usable for implementers’ situation-specific assessment and normative judgment.It allows implementers to apply relevant input features, critically assess inference-making, and weigh considerations such as public interest.
- Re-translation and interpretation: Re-translation enables stakeholder deliberation, dialogue, assessment, mutual understanding, end-to-end accountability, and regard for SUM values.The process makes algorithmically supported outcomes comprehensible and justifiable to stakeholders.
- Accountability and justification: A crucial safeguard is triangulating intention-in-design, intention-in-application, and clarification of AI results to restore moral involvement and responsibility.The translation rule links explanatory needs to the human choices and societal values encoded in the system’s content and purpose.
Conclusion:
The guide frames AI design and implementation as a human activity guided by purposes and values, for which developers and deployers are morally and socially responsible. It presents this human-centred responsibility and ethical prioritisation as the basis for steering innovation toward a shared vision of a better human future.
- Human responsibility: The guide responds to questions about how AI-driven technological advancement will shape future society and transform people’s lives and identities.It positions human moral agency in the present as central to addressing these questions.
- Human responsibility: AI design and implementation are presented as eminently human activities guided by purposes and values.The guide urges those involved in developing and deploying AI systems to recognise their moral and social responsibility.
- Responsible innovation: Responsible innovation depends on prioritising the ethical purposes and values shaping technological advancement.This prioritisation enables societal stakeholders to take control of innovation and steer algorithmic creations.