Source-linked AI summary
Measurement and Fairness
Abigail Z. Jacobs, Hanna Wallach
TL;DR
Fairness research often relies on unobservable constructs whose operationalizations can mismatch their theoretical meanings, creating potential harms. The paper applies measurement modeling and fairness-oriented reliability and validity concepts to test those assumptions, and argues that fairness-definition debates often concern competing theoretical understandings. Its framework is intended to help anticipate harms, identify actionable causes, and clarify when disagreements are about values.
Problem
Fairness systems often operationalize unobservable constructs through observable data, but mismatches between constructs and operationalizations can produce fairness-related harms.
Method
The paper uses measurement modeling, together with fairness-oriented construct reliability and construct validity, to make and test assumptions about constructs and their operationalizations.
Results
The framework shows how measurement modeling can surface mismatches, anticipate fairness-related harms, identify potential causes, and distinguish operationalization disputes from disputes about fairness values.
Takeaways & Limitations
Fairness debates can be clarified by explicitly separating measurement-model assumptions from competing theoretical understandings of fairness.
Takeaways & Limitations
The paper omits inter-rater and inter-item reliability from its construct-reliability conceptualization because relevant fairness-harm examples were unavailable or limited.
Abstract
from arXiv · showhide
We propose measurement modeling from the quantitative social sciences as a framework for understanding fairness in computational systems. Computational systems often involve unobservable theoretical constructs, such as socioeconomic status, teacher effectiveness, and risk of recidivism. Such constructs cannot be measured directly and must instead be inferred from measurements of observable properties (and other unobservable theoretical constructs) thought to be related to them -- i.e., operationalized via a measurement model. This process, which necessarily involves making assumptions, introduces the potential for mismatches between the theoretical understanding of the construct purported to be measured and its operationalization. We argue that many of the harms discussed in the literature on fairness in computational systems are direct results of such mismatches. We show how some of these harms could have been anticipated and, in some cases, mitigated if viewed through the lens of measurement modeling. To do this, we contribute fairness-oriented conceptualizations of construct reliability and construct validity that unite traditions from political science, education, and psychology and provide a set of tools for making explicit and testing assumptions about constructs and their operationalizations. We then turn to fairness itself, an essentially contested construct that has different theoretical understandings in different contexts. We argue that this contestedness underlies recent debates about fairness definitions: although these debates appear to be about different operationalizations, they are, in fact, debates about different theoretical understandings of fairness. We show how measurement modeling can provide a framework for getting to the core of these debates.
1 INTRODUCTION
The paper proposes measurement modeling as a framework for understanding fairness harms arising when unobservable constructs are operationalized through observable data. It also argues that debates over fairness definitions often reflect competing theoretical understandings of fairness rather than merely different operationalizations.
- Computational systems can encode and exacerbate societal biases through data limitations and assumptions made throughout development and deployment.
- Measurement modeling distinguishes theoretical constructs from their operationalizations, making potential mismatches easier to anticipate and mitigate.
- The paper bridges quantitative social science and computer science by introducing measurement modeling to fairness research.
- Fairness is an essentially contested construct with multiple, context-dependent, and sometimes conflicting theoretical understandings.
- Debates framed as disagreements over fairness operationalizations often concern different theoretical understandings and values.
2 MAKING ASSUMPTIONS
Measurement models link unobservable constructs to observable data, but every operationalization requires assumptions about how measurements relate to the construct. Examples involving height, socioeconomic status, teacher effectiveness, recidivism risk, and care needs show how those assumptions can produce consequential mismatches.
- Measurement modeling links latent variables representing unobservable constructs to observable properties, while requiring assumptions that should be made explicit and tested.
- Measuring Height: Even measuring height requires assumptions about what counts as height and how measurement errors behave across people and contexts.
- Measuring Height: Averaging repeated height measurements can converge on true height only when measurement errors satisfy assumptions such as statistical unbiasedness.
- Measuring Socioeconomic Status: Socioeconomic status is inferred from properties such as income, wealth, education, and occupation because it is unobservable.
- Measuring Socioeconomic Status: The Universal Credit system mismeasured monthly income by using a rolling observation window that misaligned wage deposits, leading to incorrect benefit calculations.
- Teacher effectiveness and recidivism risk are operationalized through proxies and models based on test scores, records, interviews, and criminological theories.
- Using care costs as a proxy for care needs can confound need with unequal access to care.
3 TESTING ASSUMPTIONS
The paper develops fairness-oriented construct reliability and validity concepts for testing assumptions in measurement models. These tools are intended to reveal mismatches that out-of-sample prediction can obscure and to identify actionable sources of fairness-related harms.
- Measurement-model assumptions should be made explicit and tested because unexamined mismatches can obscure fairness-related harms.
- The paper adapts construct reliability and construct validity from quantitative social science traditions for fairness in computational systems.
- Assessing these constructs complements computer science’s primary focus on out-of-sample prediction when evaluating measurement assumptions.
- The proposed tools can help anticipate fairness harms and identify potential causes that suggest concrete, actionable mitigation avenues.
3.1 Construct Reliability
Construct reliability asks whether similar inputs produce similar measurements, especially across time when the underlying construct is assumed unchanged. Test–retest failures can reveal mismatches between a construct and its operationalization, but may also reflect genuine changes in the construct or omitted reliability dimensions.
- Construct reliability: Construct reliability asks whether similar inputs to a measurement model yield similar outputs, including when presented at different points in time.It is roughly analogous to precision, or the inverse of variance, in statistics.
- Test–retest reliability: Test–retest reliability measures whether measurements of an unobservable construct remain the same over time, assuming the construct has not changed.
- Test–retest reliability: A teacher-effectiveness score changing from 6 out of 100 to 96 in consecutive years suggests poor test–retest reliability and a mismatch with the purported construct.The example concerns a respected teacher evaluated by a value-added model.
- Test–retest reliability: Checking whether Universal Credit measured claimant income consistently across different one-month rolling-period start dates might have anticipated or mitigated reported harms.
- Caveats and scope: Apparent test–retest unreliability does not always indicate operationalization mismatch, because the construct itself may change; inter-rater and inter-item reliability are also outside this conceptualization.The paper notes that people’s height can decrease with age and omits inter-rater and inter-item reliability from its quantitative fairness framework.
3.2 Construct Validity
Construct validity tests whether an operationalization captures the intended theoretical construct rather than merely producing plausible measurements. The paper applies multiple validity aspects to show how proxy choices and modeling assumptions can create fairness-related mismatches.
- Construct validity examines whether an operationalization measures the intended construct through several complementary aspects.The paper’s framework unites traditions from political science, education, and psychology.
- Face validity: EVAAS MRM scores are an exception to otherwise plausible-looking measurements because their dramatic test–retest variability is implausible.The model also assumes teacher effectiveness is fully captured by students’ test scores and that scores depend on specified teacher effects.
- Content validity: Substantive validity requires including those—and only those—observable properties and related constructs thought to characterize the target construct.Income alone cannot fully represent socioeconomic status because wealth, education, occupation, and cultural capital also contribute.
- Content validity: COMPAS threatens substantive validity by treating arrests within two years as a proxy for crimes committed, excluding false arrests and unarrested crimes.The paper states that arrest data cannot wholly capture the substantive nature of crime.
- Convergent validity: Many value-added models lack convergent validity when their teacher-effectiveness scores conflict with independent evaluations from principals, colleagues, students, and parents.The paper illustrates this with Sarah Wysocki, whose model score was low despite excellent reviews.
- Discriminant validity: Discriminant validity identifies whether measurements inadvertently capture aspects of other constructs, as demonstrated by care-need scores correlating strongly with race.Obermeyer et al. attributed the pattern to unequal care access and costs shaped by structural racism; only 18% of identified patients were Black.
4 FAIRNESS AS A CONSTRUCT
Fairness is an essentially contested, unobservable construct whose competing theoretical understandings shape debates over its operationalization. Measurement modeling reveals how parity-based and individual-fairness definitions can omit substantive dimensions and conceal value choices.
- Fairness has multiple context-dependent, sometimes conflicting theoretical understandings, making explicit operationalization necessary.
- Quantitative, parity-based operationalizations are appealing because they are numerical and remain within a single computational system.Their system-level focus excludes the broader societal context in which systems operate.
- Operationalizations that omit justice cannot fully capture fairness understandings grounded in procedural, distributive, or representational justice.They may also be ineffective at remedying historical injustices, threatening consequential validity.
- Individual fairness and group fairness represent conflicting theoretical understandings, not merely different mathematical operationalizations.Group fairness requires similar treatment across groups, whereas individual fairness requires similar treatment for similar people.
- Similarity itself is contested and difficult to measure, so practical individual-fairness operationalizations may lack content and consequential validity.Measuring similarity can involve subjective or personal aspects of human experience.
- Predictive parity and error-rate balance encode different group-fairness understandings and therefore prioritize different values and consequences.Predictive parity gives scores the same meaning across groups, whereas error-rate balance makes errors comparable across groups.
- Inferring demographic constructs from observable properties is fraught and can cause fairness-related harms even after reliability and validity are assessed.Self-reports would ideally limit harmful errors, but are often unavailable at the individual level.
5 DISCUSSION
Measurement modeling distinguishes theoretical constructs from their operationalizations and provides tools for testing the assumptions connecting them. The paper argues that this process can expose fairness harms and make debates about fairness definitions more transparent and accountable.
- Unobservable constructs such as socioeconomic status, teacher effectiveness, and recidivism risk must be inferred through measurement modeling.Because these constructs cannot be measured directly, inference necessarily relies on assumptions.
- Treating measurements as ground truth collapses constructs with their operationalizations and embeds developmental assumptions throughout society.The paper connects this concern to the view that measures and classifications help create social categories and stratifications.
- Construct reliability and construct validity provide tools for surfacing mismatches between constructs and operationalizations.The paper develops fairness-oriented conceptualizations drawing on political science, education, and psychology.
- These tools can anticipate fairness harms obscured by focusing primarily on out-of-sample prediction and identify potential causes with actionable mitigation avenues.
- Parity-based fairness definitions can lead stakeholders to adopt systems without critical assessment when mathematical compliance is treated as a fairness guarantee.
- Measurement modeling promotes transparency and accountability by making assumptions explicit and testing for mismatches.The paper acknowledges that assessing reliability and validity can be time-consuming.