Source-linked AI summary

Problem Formulation and Fairness

Samir Passi, Solon Barocas

arXiv:1901.02547v1cs.CY

TL;DR

The paper examines how uncertain translations from strategic goals to tractable data-science problems shape fairness concerns, a process often overlooked by model-centered assessments. Using six months of ethnographic fieldwork with a corporate data-science team, it finds that problem specifications and operationalizations are negotiated and elastic. The authors conclude that normative intervention must attend to the practical work of problem formulation.

  • Problem

    Normative assessments often overlook how translating goals into target variables and proxies shapes ethical concerns in data-science projects.

  • Method

    The paper uses six months of ethnographic fieldwork with a corporate data-science team to trace actors, activities, and negotiations in problem formulation.

  • Results

    Problem formulations are negotiated and elastic, with different target variables and proxies creating different understandings of what a data-science problem means.

  • Takeaways & Limitations

    Normative implications arise in the discretionary translations between high-level goals and tractable problems, making problem formulation an upstream site for intervention.

  • Takeaways & Limitations

    The case is constrained by selection bias and limited data on dealer decisions and approvals in the special-financing setting.

Abstract

from arXiv · show

Formulating data science problems is an uncertain and difficult process. It requires various forms of discretionary work to translate high-level objectives or strategic goals into tractable problems, necessitating, among other things, the identification of appropriate target variables and proxies. While these choices are rarely self-evident, normative assessments of data science projects often take them for granted, even though different translations can raise profoundly different ethical concerns. Whether we consider a data science project fair often has as much to do with the formulation of the problem as any property of the resulting model. Building on six months of ethnographic fieldwork with a corporate data science team---and channeling ideas from sociology and history of science, critical data studies, and early writing on knowledge discovery in databases---we describe the complex set of actors and activities involved in problem formulation. Our research demonstrates that the specification and operationalization of the problem are always negotiated and elastic, and rarely worked out with explicit normative considerations in mind. In so doing, we show that careful accounts of everyday data science work can help us better understand how and why data science problems are posed in certain ways---and why specific formulations prevail in practice, even in the face of what might seem like normatively preferable alternatives. We conclude by discussing the implications of our findings, arguing that effective normative interventions will require attending to the practical work of problem formulation.

1 INTRODUCTION

Data science projects require difficult translations from amorphous goals into measurable, tractable prediction problems. Because target-variable choices can shape disparate impact and broader fairness concerns, ethical assessment must examine problem formulation as well as model behavior.

  • Problem formulation: Problem formulation translates high-level business objectives into questions that data science can answer by predicting measurable variables.The desired outcome or quality is often not self-evident and must be operationalized.
  • Operationalization: Machine learning requires explicit, measurable definitions of qualities such as a “good” employee.Employers may replace difficult-to-measure qualities like personability with easier-to-monitor sales figures.
  • Fairness: Choosing among competing target variables can affect whether a hiring model exhibits disparate impact.Target variables may correlate with protected characteristics, serve as proxies, or create avoidable disparate impact despite rational business aims.
  • Fairness: Biased proxy labels can cause models to reproduce unequal measurement, as arrests may represent crime racially unevenly.When labels function as ground truth, predictions learn the bias embedded in those labels.
  • Fairness: Across three cases, fairness outcomes often depend on what the model is designed to predict.Target-variable selection can both create unfairness and provide a mechanism for avoiding it.
  • Contribution: Everyday problem formulation involves actors and activities that shape how ethical implications arise in applied data science.The paper studies this underexamined work through an ethnographic account of a corporate special-financing project.

2 BACKGROUND

Research on science and critical data studies frames methods, data, and algorithms as shaping the questions and phenomena they can represent. Applied data science therefore involves iterative, discretionary choices whose consequences extend beyond technical model construction.

  • History and sociology of science: Scientific methods influence which questions are asked and how phenomena are defined and measured.Representations are contingent on methods while remaining actionable for analyzing the world.
  • Critical data studies: Data scientists iteratively align algorithms and data while navigating competing formal, organizational, and practical demands.This work is situated and discretionary rather than merely a collection of mechanical rules.
  • Practical work: Data science requires subjective decisions about measurement, research design, data processing, modeling, and interpretation.The passages connect these practical choices to potentially profound ethical implications.
  • Applied settings: Corporate data science projects are collaborative endeavors involving discretion, negotiation, and aspiration alongside data, numbers, and models.Multiple actors jointly make sense of data and algorithmic results.
  • Problem formulation: Problem formulations emerge from both goals and contingent choices about available data, relevant evidence, and methods.The relationship between data and formulated problems is therefore not one-way.

Knowledge Discovery in Databases

Knowledge Discovery in Databases emphasized the process surrounding data mining and developed formal models to make its steps explicit. Yet translating business goals into data-mining problems remains a negotiated, iterative process shaped by actors, methods, instruments, and data.

  • KDD history: KDD emerged partly to formalize human judgment and discretion throughout the data-mining process.Its process models sought to explicate how practitioners progress through data-mining projects.
  • Process models: CRISP-DM breaks a data-mining project into discrete lifecycle steps while simultaneously describing and prescribing them.It became the most widely adopted process model cited here.
  • Process models: Early KDD process models responded to fears of mistakes, missteps, and misapplications rather than only documenting practitioners’ activities.The goal was to identify an appropriate way to conduct data mining.
  • Negotiated translation: Business understanding translates amorphous business goals into questions amenable to data mining.This translation requires conversations among managers, technologists, and analysts.
  • Negotiated translation: Problem formulation is a negotiated translation contingent on discretionary judgments and choices of methods, instruments, and data.The paper treats it as more than a one-time first-phase mapping from goals to computational problems.
  • Iterative practice: Making machine learning return desired results involves substantial manual work and subjective judgment.The paper presents initial formulations as elastic placeholders that can change through project iterations.

3 RESEARCH SITE AND METHODS

The study draws on six months of ethnographic fieldwork at a pseudonymized US corporate data-science organization. Researchers combined interviews, field notes, photographs, and grounded-theory coding, focusing on a salient special-financing case while observing similar dynamics elsewhere.

  • Research site: The research site was DataVector, a multi-billion-dollar US e-commerce and new-media organization with a core data-science team.The team worked with companies across domains including health and automotive.
  • Data collection: During six months, the researchers conducted 50+ interviews and produced 400+ pages of fieldwork notes and 100+ photographs.Participants included data scientists, product managers, business analysts, project managers, and executives.
  • Analysis: Interviews and fieldwork data were transcribed and coded through two rounds of grounded-theory analysis.Coding focused on categories, themes, topics, and their relationships.
  • Research ethics: Organization, personnel, and project names were replaced with pseudonyms to preserve participant anonymity.The case-study organization is therefore referred to as DataVector.
  • Case selection: The analysis centered on one corporate project because problem formulation was especially salient there.The researchers observed similar dynamics across several other projects.

4 CASE-STUDY: SPECIAL FINANCING

CarCorp’s effort to improve lead quality became a negotiated attempt to define and predict dealer-specific financeability using incomplete, inconsistent data. The project ultimately stalled because the chosen formulation focused on distinguishing leads near a credit-score threshold that the available data could not reliably resolve.

  • Defining lead quality: CarCorp sought to improve lead quality so dealers would continue using its services, but stakeholders disagreed about what made a lead good.Proposed definitions included salary data, vehicle inventory, likelihood of purchase, and ability to secure financing.
  • Defining lead quality: The teams narrowed lead quality to dealer-specific financeability: predicting which dealer was most likely to finance each lead.Because dealers used different approval processes, a lead could be financeable for one dealer but not another.
  • Data constraints: Dealer-approval data was scarce, while available lead data came from inconsistent sources and often represented credit scores only as approximate ranges.CarCorp processed nearly two million leads in 2017 but had relatively little information about dealer approvals; exact credit scores were also unavailable without explicit consent.
  • Data constraints: Adding credit-score ranges improved financeability prediction for the roughly 10% of leads with that information, prompting an attempt to predict ranges for the remaining 90%.Traditional statistical analyses had struggled to predict credit-score ranges from other features.
  • Project outcome: The resulting formulation concentrated on separating leads above and below a 500 credit-score threshold, but the 476–525 range mixed both groups and available models performed only slightly better than chance.Errors near the threshold were especially consequential because the business wanted to identify leads just above it.
  • Project outcome: The project was halted after the team could not develop an accurate model or identify actionable progress, with participants attributing the failure variously to the data, expectations, or the problem’s formulation.Ron characterized the challenge as distinguishing different low credit scores within the special-financing population rather than separating high and low scores across the full spectrum.

5 DISCUSSION

Problem formulations are negotiated translations shaped by actors, proxies, practical constraints, and available data, with different formulations producing different normative concerns. Fairness therefore requires examining how goals become tractable targets, not only evaluating finished models.

  • Problem formulation proceeds through negotiated choices among goals, target variables, and proxies rather than a fixed translation.Actors equated lead quality with financeability, financeability with dealer decisions, and dealer decisions with credit-score ranges.
  • Problem formulations evolve alongside normative implications because high-level goals rarely emerge fully specified.In the case study, changing formulations of lead quality changed the project’s normative implications.
  • Dealer decisions and credit-score ranges produced matching and classification formulations with different meanings of financeability.Matching treated financeability as dealer-specific and continuous, while classification treated it as binary above or below 500.
  • The classification formulation excluded below-500 leads from dealer consideration, while the matching formulation could expand opportunities by directing leads to suitable dealers.The two formulations reflect different practical consequences for which leads are surfaced and to whom.
  • Which formulation appears fairer depends on the governing principle: maximizing lending opportunities versus mitigating existing dealer biases.The paper does not identify one formulation as universally normatively preferable.
  • Translations are always imperfect and partial, so analysis should focus on their consequences and the everyday judgments driving them.The authors present this shift as an alternative to criticizing practitioners for failing to find a perfectly faithful translation.
  • Business requirements, proxy choices, algorithmic task design, data availability, analytic uncertainty, and financial cost shaped formulations more directly than explicit ethical debate.The actors did not explicitly debate the ethical implications of their systems, while data constraints and dataset costs materially affected available choices.
  • Normative interventions should address the iterative, less visible work of formulating problems as an upstream site for downstream change.The paper argues that examining this work helps identify and accommodate the normative implications of data science systems.

6 CONCLUSION

The paper examines how questions are posed in real-world data science projects and shows that problem formulation is an uncertain, consequential process. Making goals actionable transforms them in ways that shape both the problems addressed and the appropriate responses.

  • Problem formulation transforms objectives as practitioners make goals amenable to data science and actionable results.These transformations can affect the conception of the problem and how it is handled.
Loading 1901.02547v1…