Source-linked AI summary
"What We Can't Measure, We Can't Understand": Challenges to Demographic Data Procurement in the Pursuit of Fairness
McKane Andrus, Elena Spitzer, Jeffrey Brown, Alice Xiang
TL;DR
Fairness practitioners often lack the demographic data needed to detect bias. Through interviews with 38 practitioners, the paper finds that access is a significant barrier and raises normative questions about how, when, and whether such data should be collected.
Problem
Many algorithmic fairness strategies require demographic attributes or proxies, but practitioners often lack access to the data needed to detect disparities.
Method
The study uses interviews with 38 practitioners from 26 organizations to examine demographic data procurement and use in practice.
Results
Almost every participant described demographic data access as a significant barrier, with gender and age commonly available but race data rarely accessible outside select domains.
Takeaways & Limitations
The findings caution against simply lowering collection barriers and instead foreground normative questions about how, when, and whether demographic data should be collected and used.
Takeaways & Limitations
How legal challenges to algorithmic fairness efforts will play out in court remains uncertain.
Abstract
from arXiv · showhide
As calls for fair and unbiased algorithmic systems increase, so too does the number of individuals working on algorithmic fairness in industry. However, these practitioners often do not have access to the demographic data they feel they need to detect bias in practice. Even with the growing variety of toolkits and strategies for working towards algorithmic fairness, they almost invariably require access to demographic attributes or proxies. We investigated this dilemma through semi-structured interviews with 38 practitioners and professionals either working in or adjacent to algorithmic fairness. Participants painted a complex picture of what demographic data availability and use look like on the ground, ranging from not having access to personal data of any kind to being legally required to collect and use demographic data for discrimination assessments. In many domains, demographic data collection raises a host of difficult questions, including how to balance privacy and fairness, how to define relevant social categories, how to ensure meaningful consent, and whether it is appropriate for private companies to infer someone's demographics. Our research suggests challenges that must be considered by businesses, regulators, researchers, and community groups in order to enable practitioners to address algorithmic bias in practice. Critically, we do not propose that the overall goal of future work should be to simply lower the barriers to collecting demographic data. Rather, our study surfaces a swath of normative questions about how, when, and whether this data should be procured, and, in cases where it is not, what should still be done to mitigate bias.
1 Introduction
Algorithmic fairness strategies commonly require demographic data to measure or mitigate disparities, yet practitioners encounter social, legal, and practical barriers to accessing and using it. This study examines those impediments and practitioners’ concerns about responsible demographic data collection in practice.
- Most algorithmic fairness strategies measure or mitigate disparities across sensitive or protected attributes, typically requiring access to demographic data.The introduction frames demographic data access as central to applying fairness strategies, metrics, and toolkits.
- Demographic data collection and use face social and legal constraints, including heightened protections under regulations such as the GDPR.These constraints create a potential conflict between demographic data availability and efforts to make systems less discriminatory.
- Researchers have proposed bypassing demographic data collection through inference, proxies, training-only use, or computationally identified groups.These approaches seek to address discrimination without always using demographic data directly.
- Through interviews with practitioners concerned with algorithmic fairness, the study characterizes real-world impediments to demographic data use and concerns about responsible collection.The study focuses on how demographic data availability is encountered and handled in practice.
2 Background and Related Research
Prior research has examined demographic-data procurement primarily through legal constraints and increasingly through critiques of demographic categories themselves. However, comparatively little work has addressed how practitioners navigate these procurement and use challenges in practice.
- Legal constraints: Legal scholarship highlights uncertainty over how GDPR and anti-discrimination law constrain demographic-data procurement, storage, use, and inclusion in decision-making.Open questions concern fairness-related carve-outs under data-protection regulations and the permissibility of using demographic attributes under anti-discrimination law.
- Conceptual critiques: Critical scholarship questions demographic categories as foundations for fairness assessment, examining race, gender, and disability and the harms reproduced by categorization infrastructures.This work interrogates how the algorithmic fairness community conceptualizes demographic attributes and how those concepts propagate into notions of fairness.
- Practitioner research: Less research has examined how practitioners confront demographic-data procurement and use, although studies have addressed public-sector and private-sector needs and organizational patterns.These practitioner-focused studies provide groundwork for understanding practitioners, but the passage identifies procurement and use as comparatively understudied.
- Practitioner research: A majority of surveyed respondents considered tools for fairness auditing without individual-level demographic access at least “Very” useful.The finding comes from Holstein et al.’s survey of fairness practitioners and underscores demand for approaches that do not require individual-level demographic data.
3 Methods
The study used semi-structured interviews with 38 practitioners from 26 organizations to examine how demographic data is procured and used in algorithmic-fairness work. Interviews covered participants’ experiences, data-access constraints, approval processes, legal risks, and privacy–fairness trade-offs, and were analyzed through iterative thematic coding.
- Participants: 38 practitioners from 26 organizations participated, with no more than 5 participants from any single organization.Participants were involved in bias detection, familiar with relevant company policies, or otherwise knowledgeable about demographic-data use.
- Participants: 34 of 38 participants worked at for-profit technology companies, 55% came from companies with more than 1,000 employees, and 80% worked in the US.Participants held varied organizational positions and represented diverse sectors.
- Interview procedure: Each interview was a 60–75 minute video call involving verbal recording consent, transcription, audio deletion, and redaction to protect participant and institutional anonymity.Researchers shared consent and confidentiality practices before the calls; in some cases, multiple interviewees from one organization joined the same call.
- Interview scope: Researchers defined demographic data broadly to include regulated attributes, less-protected categories, and potential proxies such as likes or dislikes.Examples included sex, race, national origin, socioeconomic class, and geography.
- Interview scope: Questions elicited concrete successful or unsuccessful bias-assessment cases, desired data and purposes, availability, approvals, legal risks, privacy trade-offs, and participants’ ideal assessment practices.Question framing was adapted for legal or policy participants and for participants identified through public reporting or organizational publications.
- Analysis: Researchers used open and closed coding in MaxQDA, iteratively grouping open codes into thematic networks before applying and expanding closed codes across the transcripts.Open coding initially covered 25% of the interviews, with themes organized around challenges or constraints on demographic-data use.
4 Results Overview
Practitioners widely described demographic-data access as a major barrier to implementing fairness techniques. The analysis organizes procurement and use challenges into external regulatory and organizational constraints and practitioner-facing concerns.
- Almost every participant described access to demographic data as a significant barrier to implementing fairness techniques.This result mirrors responses reported by Holstein et al.
- Virtually all companies that collected or commissioned their own data had access to gender and age.
- The analysis distinguishes external regulatory and organizational constraints from concerns practitioners themselves surfaced or encountered around demographic-data procurement and use.External actors include legal and compliance teams, policy teams, company leadership, external auditors, and regulators.
5 Legal and Organizational Constraints on Demographic Data Procurement and Use
Demographic data procurement is shaped by privacy regulation, anti-discrimination requirements, and organizational priorities, producing domain-specific tensions between data minimization, fairness assessment, and legal compliance. These constraints can limit access to demographic data even when practitioners view it as necessary for detecting bias.
- Privacy constraints: Privacy laws and policies guide data handling through consent and data minimization, while treating pseudonymous data as regulated personal data.GDPR-related practices emphasized collecting only required data and obtaining data-subject consent; pseudonymous identifiers may still fall under regulation.
- Privacy constraints: Race was considered salient for fairness analysis, yet organizations commonly avoided collecting special-category data because of legal, reputational, and trust-related risks.Large-scale access was rare and generally depended on government datasets, self-identification, or explicitly applicable regulation and standards.
- Privacy constraints: Privacy concerns also discourage demographic inference and can leave external auditors, consultants, and vendors unable to access the data needed for their work.Potential data-access obligations for inferred attributes encourage conservative practices, while many outside evaluators must work without seeing organizational data.
- Anti-discrimination constraints: Anti-discrimination law can require demographic collection for some financial products while nearly barring it for others, even as institutions remain subject to oversight.U.S. mortgage lenders must collect demographic attributes, whereas credit-based products nearly prohibit collection; institutions may follow regulatory guidance when direct data is sparse.
- Organizational constraints: Fairness assessments can diagnose disparities beyond current legal requirements, but their feasibility depends on organizational buy-in and willingness to accept collection-related risks.The paper identifies leadership and clientele support as key determinants of constraint severity and assigns fairness practitioners a role in making the case for demographic data.
6 Fairness Practitioners’ Concerns … Collection Procedure Difficulties.
Fairness practitioners face substantial difficulty procuring demographic data because self-reports are often unreliable, incomplete, and shaped by low trust or limited incentives. Even when collection is feasible, organizations must navigate unclear procedures, reputational risks, and inconsistent interpretations of demographic categories.
- 6 Fairness Practitioners’ Concerns: Practitioners expressed caution about procuring and subsequently using demographic data, beyond the business, legal, and reputational pressures elsewhere in their organizations.
- 6.1 Concerns Around Self-Reporting: Practitioners reported concerns about the effectiveness of current demographic-data collection methods.
- Frequency and Accuracy of Responses.: Self-reported demographic data was most commonly described as unreliable or incomplete, partly because individuals often lack strong incentives to respond accurately.
- Frequency and Accuracy of Responses.: Trust was the most frequently discussed reason for low survey response rates, with users questioning how platforms would use demographic information and who would decide what they wanted.
- Frequency and Accuracy of Responses.: Healthcare practitioners viewed demographic-data collection as an effortful form of donation because it requires individuals’ participation and lacks an automated collection method.
- Collection Procedure Difficulties.: Organizations may want demographic data for disparate-impact analysis but lack clear procedures for asking users to provide it.
- Collection Procedure Difficulties.: A platform-wide request for gender identity could generate public and reputational costs that exceed the expense of designing and deploying the collection mechanism.
- Collection Procedure Difficulties.: Even reliably collected data can create downstream problems when respondents choose among predetermined categories whose meanings are not interpreted consistently.
Flexible Categories. · 6.2 Concerns around Proxies and Inference
Practitioners treat demographic categories as flexible and potentially misinterpretable, while proxy-based or inferred attributes raise distinct concerns. Wary data subjects may provide demographic information inaccurately or not at all, complicating fairness assessments.
- Flexible Categories.: Practitioners caution that demographic categories can be misinterpreted even when data is not inferred.One system distinguishes male, female, unknown, and opt-in custom gender; unknown combines nonresponse with explicit refusal.
- Flexible Categories.: Flexible demographic categories require careful interpretation because labels may combine nonresponse, refusal, and self-described identities.The example includes “unknown” and an opt-in custom gender field.
- Flexible Categories.: Practitioners recognize that a single demographic category may encompass many different identities.This motivates caution in interpreting and using the demographic data they can access.
- Flexible Categories.: Wary data subjects may produce low response rates and poor response accuracy when collection purposes or consequences are unclear.The paper suggests framing fairness data collection as “data donation” might increase buy-in and response.
- Flexible Categories.: Poor understanding of collection purposes can undermine the quality of demographic data used for anti-discrimination and fairness assessments.The passage links wariness to both response rates and response accuracy.
- 6.2 Concerns around Proxies and Inference: Practitioners therefore face separate challenges in collecting demographic data and using proxies or inferred attributes for sensitive features.The paper contrasts survey-design pitfalls with concerns specific to proxy and inference practices.
- 6.2 Concerns around Proxies and Inference: Proxy and inferred demographic data create a distinct set of concerns from conventional demographic-data collection risks.These concerns arise especially for sensitive features that oversight teams are unlikely to approve for collection.
- 6.2 Concerns around Proxies and Inference: Sensitive demographic features may be inferred directly from available data when oversight teams do not approve their collection.The passage identifies this practice as part of practitioners’ response to collection constraints.
Prevalence of Proxies and Inferred Categories. · When Inferred Categories are Preferred. · Treating Proxies as a Weak Signal
Practitioners’ use of inferred demographics and proxies varies sharply by domain, with inference often serving as the only available route to assess discrimination. However, inferred categories and proxies raise unresolved questions about suitability, standardization, and the risk of treating weak signals as definitive measurements.
- Prevalence of Proxies and Inferred Categories.: U.S. financial institutions reported using BISG to infer individual race in accordance with Consumer Financial Protection Bureau standards.By contrast, participants in domains with mandated demographic collection rejected demographic inference as unnecessary or inappropriate in their contexts.
- Prevalence of Proxies and Inferred Categories.: In other domains, demographic inference is often the only available option for discrimination assessments, despite few standardized practices.Participants mainly described inference for content data, including skin-tone classification in images and author-gender prediction in text.
- When Inferred Categories are Preferred.: Inferred demographics may be preferred when the task concerns how users perceive one another’s race rather than users’ self-reported identities.EC3 described perceived race as relevant to studying a racial experience gap on a technology platform.
- When Inferred Categories are Preferred.: Practitioners sometimes use accessible attributes, such as subscription tier, to obtain signals about discrimination involving an inaccessible salient attribute.They emphasized that such attributes are proxies rather than direct measures of the target category.
- Treating Proxies as a Weak Signal: Although proxies are frequently criticized in fairness research, they may be practitioners’ only way to make a quantitative case for escalation and deeper analysis.This instrumental use reflects limited access to salient demographic attributes rather than confidence that the proxy is an adequate measure.
- Treating Proxies as a Weak Signal: Proxies become especially risky when normalized, because practitioners may stop examining their inadequacies or apply standard mitigation techniques without defining the proxy’s measurement model.The passage recommends clearly outlining how a proxy represents the desired attribute before using it in fairness analysis.
6.3 Relevance of Demographic Categories · Deficient Standards. · Difficulties with Granular Representations.
Practitioners struggle to determine which demographic groups merit attention and how categories should be defined, especially across regions and under outdated standards. Even expanded categories leave difficult decisions about reasonable limits and the most relevant granular slices for evaluation.
- 6.3 Relevance of Demographic Categories: Fairness analyses often require discrete groups, but practitioners remain uncertain about which groups warrant attention and how to define them.These concerns were especially salient for teams building products and services for regions outside their own.
- 6.3 Relevance of Demographic Categories: Organizations working across regions face particular difficulty determining relevant demographic groups for fairness analysis.The challenge is heightened when teams build and maintain products for regions different from their own.
- Deficient Standards.: Practitioners described standard demographic categories as inadequate, particularly in industries that collect data according to governmental standards.One hiring/HR product manager said their organization was seeking data beyond employment-law categories because those categories were defined decades ago.
- Deficient Standards.: Governmental demographic standards can lag contemporary needs, prompting practitioners to seek additional categories and variables.The reported concern was that existing categories were difficult to use in present-day contexts.
- Deficient Standards.: Expanding categories to improve representation creates uncertainty about how many categories a demographic question should reasonably include.Practitioners noted that lists could grow to 20 or even 100 categories, making a practical limit necessary.
- Difficulties with Granular Representations.: Even after categories are expanded, practitioners must decide which categories or combinations deserve evaluation and deeper inspection.Identifying the most relevant demographic slices was described as very difficult.
What Categories Should Be Focused on? · Regional Differences. · Convenience.
Practitioners struggle to determine which demographic categories matter, especially when legal requirements, regional assumptions, data availability, and intersecting identities do not align with ethical or cultural salience. The paper argues that context-specific categorization is necessary because convenient or inapt schemas can obscure important forms of bias.
- What Categories Should Be Focused on?: Practitioners often lack authority to decide which demographic categories matter, while assessing intersections across multiple attributes is generally impractical.Teams may need to identify bias across intersecting groups while also informing algorithmic, design, and policy changes.
- What Categories Should Be Focused on?: Legal and policy requirements may provide a straightforward categorization rule, but they can conflict with practitioners’ views of ethical or cultural salience.This conflict can impede efforts to make systems fairer when testing reveals poor performance for groups outside mandated categories.
- Regional Differences.: Companies may define legally salient categories through the laws of their home region, producing a U.S.-centric understanding of relevant demographic characteristics elsewhere.Practitioners specifically questioned whether commonly emphasized categories remain relevant across different countries and cultures.
- Regional Differences.: When relevant demographic data is unavailable, practitioners may prioritize categories that are available rather than those that are socially salient.Gender may receive attention because organizations are permitted to collect and act on it, even when broader diversity assessment is desired.
- Convenience.: Inapt categorization schemas can prevent meaningful understanding of bias because different demographic categories may produce dramatically different analyses and conclusions.The choice of categories therefore affects what types of bias become visible in testing.
- Convenience.: A one-size-fits-all approach that prioritizes scalability may miss the most salient forms of possible discrimination.Addressing this risk requires earnest attempts to grapple with specific issues in limited contexts.
6.4 Ensuring Trustworthiness and Alignment · Generating Well-Founded Trust. · Alignment with Data Subjects.
Practitioners described trust and alignment around demographic data as difficult to establish because corporate and community goals may diverge and the public doubts companies’ data practices. They identified reciprocal transparency and giving affected groups meaningful voice as ways to build more trustworthy systems.
- 6.4 Ensuring Trustworthiness and Alignment: Low public trust in technology companies and weak alignment between company and community goals make fairness relationships difficult to realize in practice.
- 6.4 Ensuring Trustworthiness and Alignment: Practitioners in less-regulated industries questioned whether their organizations could meaningfully earn public trust around demographic data collection.Concerns were tied to recent public-relations cycles involving data misuse and mishandling.
- Generating Well-Founded Trust.: Users and clients often doubt that their data will be used with their control or for their benefit.
- Generating Well-Founded Trust.: Trust-building requires organizations to clarify what data they seek and why, while following through by sharing at least summary results with contributors.The proposed relationship is reciprocal: users and clients provide data for performance assessment, and organizations report back on results.
- Alignment with Data Subjects.: Meaningful alignment requires addressing poor system performance for demographic groups by paying attention to and giving voice to affected communities.
- Alignment with Data Subjects.: Americans are largely unclear about how their data are used and unconvinced that companies use them for their benefit, reinforcing the need to center data subjects.
6.5 Mitigation Uncertainty · What Should Be Done About Uncovered Bias. · Uncertainty About Whether Proposed Interventions Will Be Adopted.
Practitioners face uncertainty about whether demographic data will enable effective mitigation, which can deter them from addressing barriers. They also lack clear guidance on matching interventions to harms and on whether organizations will permit or adopt equity-focused measures.
- 6.5 Mitigation Uncertainty: Uncertainty about the effectiveness of using demographic data and acting on it can deter practitioners from overcoming related barriers and concerns.This uncertainty is identified as a direct obstacle to mitigation efforts.
- What Should Be Done About Uncovered Bias.: Fairness research offers potential mitigation techniques, but practitioners reported few applicable real-world examples and limited clarity about which harms each method addresses.The lack of concrete examples and harm-specific guidance makes context-appropriate method selection difficult.
- Uncertainty About Whether Proposed Interventions Will Be Adopted.: Practitioners were unsure what interventions corporate policy would permit, even when they wanted to exceed minimum fairness requirements.Their analyses could outpace issues previously considered by policy teams.
- Uncertainty About Whether Proposed Interventions Will Be Adopted.: Achieving equity may require nuanced treatment of different groups rather than only equalizing performance metrics or removing demographic data.The passage identifies calibration, predictive parity, and ignoring demographic data as possible approaches that may be insufficient for equity.
- Uncertainty About Whether Proposed Interventions Will Be Adopted.: Practitioners questioned whether anti-bias work should include treating historically marginalized groups with more care, not merely ensuring models are unbiased.TC9 advocated carving out space for differentiated treatment as part of a stronger anti-bias stance.
- Uncertainty About Whether Proposed Interventions Will Be Adopted.: Product and research teams still grapple with defining algorithmic fairness, while fairness optimization can raise broader questions about corporations’ role in creating a more just world.The implications caution that these normative questions cannot be left solely to product and research teams.
7 Discussion
The discussion argues that progress should not simply lower barriers to demographic data collection, because doing so raises fundamental questions about whether and how such data should be collected and used. It highlights legal reforms, privacy-preserving arrangements, and meaningful involvement of data subjects as possible directions, while emphasizing that collection alone may perpetuate harm without intervention.
- Normative implications: The paper rejects simply lowering barriers to demographic data collection, emphasizing normative questions about whether and how developers should collect and use it.The discussion frames these questions as central consequences of the challenges practitioners face when pursuing fairness goals.
- Legal provisioning: Legal carve-outs may be needed so demographic data can be used to assess and address algorithmic discrimination.Possible arrangements include official third-party data holders or auditors, or allowing private organizations to perform this work.
- Legal provisioning: Data-protection carve-outs should coincide with stronger protections against algorithmic discrimination, because existing anti-discrimination law may not incentivize demographic-data use.In domains where collection is already legally permitted, practitioner concerns persist but shift toward issues such as government-defined categories.
- Privacy-preserving strategies: Privacy-preserving fairness strategies use trusted third parties to collect data and conduct secure analyses, model training, and knowledge sharing.These approaches are presented as ways to enable collection in low-trust environments while limiting data use to fairness-related purposes.
- Data-subject involvement: Meaningfully involving data subjects matters because demographic-data collection can perpetuate harm without a strong commitment to reversing systemic inequalities.Debates over racial data in COVID response plans illustrate why collection should be connected to actual intervention.