Source-linked AI summary

"I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work

Anna Kawakami, Chloe Qianhui Zhao, Renee Shelby, Fernando Diaz, Haiyi Zhu, Kenneth Holstein

arXiv:2608.22459v1cs.HC

TL;DR

Workers are often asked to use AI without defining what meaningful augmentation should mean or how it should be evaluated. This paper studies worker-driven AI measurement through eight workshops with 19 school social workers, who designed a benchmark for whether LLMs challenge reflection on assumptions and biases. The resulting benchmark aligned strongly with worker judgments and differentiated six state-of-the-art LLMs, while the case study’s scope and validation limit broader conclusions.

  • Problem

    Workers rarely help define meaningful AI augmentation or the evaluations used to judge it, leaving a gap in worker-centered measurement design.

  • Method

    Through eight workshops with 19 school social workers, workers collaboratively identified use cases, defined success, and operationalized their goals into an LLM-as-a-judge benchmark.

  • Results

    The benchmark showed strong agreement with worker judgments and differentiated performance across six state-of-the-art LLMs.

  • Takeaways & Limitations

    Worker-driven measurement can identify and operationalize work-specific goals that existing AI measurement instruments may not represent.

  • Takeaways & Limitations

    The small organization and exploratory workflow limit scope and generalizability, while validation did not test whether high benchmark scores predict better worker performance.

Abstract

from arXiv · show

Workers are increasingly asked to adopt AI systems to assist their work, yet are rarely given a voice in defining what meaningful AI augmentation should look like or how to evaluate for it. In this paper, we propose worker-driven AI measurement---a bottom-up approach to AI evaluation where workers collaboratively shape decisions about which tasks AI should augment, what "successful" augmentation looks like, and how it should be measured. We explore how to support this through a case study with 19 workers from a local school social work organization. Through a series of eight workshops, workers iteratively develop their own measurement goals for AI evaluation, systematize these goals, and then design a benchmark to capture how effectively an LLM can "challenge" them to reflect on their own assumptions and biases in the context of their day-to-day work. Workers collaboratively design and refine an LLM-as-a-judge rubric based on their professional and lived expertise. In validations of the worker-created benchmark, we find that there is strong agreement between worker and LLM judge ratings and that the resulting benchmark can differentiate performance across six state-of-the-art LLMs. Based on our case study, we discuss opportunities for future work to support worker-driven AI measurement as a complementary approach to existing top-down AI evaluation approaches.

1 Introduction

The paper proposes worker-driven AI measurement, in which workers collaboratively decide what AI should augment, what success means, and how to measure it. In a case study with school social workers, eight workshops produced a worker-designed benchmark for evaluating whether LLMs meaningfully challenge workers’ thinking.

  • Workers are rarely given a voice in defining meaningful AI augmentation or how it should be evaluated.
  • Existing AI measurement work increasingly targets occupation-specific tasks, but worker involvement in shaping evaluations remains limited.
  • Across eight workshops with 19 workers, the study explored identifying desired LLM use cases, defining good performance, and operationalizing shared goals into a benchmark.
  • Workers designed an LLM-as-a-judge benchmark for developing reflection questions that help guide meetings with teachers.
  • Workers identified “people challenging” as a measurement goal: prompting reflection on assumptions and biases while expanding perspectives and supporting learning and growth.
  • The worker-created rubric showed strong agreement with worker judgments and differentiated performance across state-of-the-art LLMs.

2 Background

The background frames AI evaluation as a social-science measurement problem and identifies limited worker participation across the measurement lifecycle. The paper addresses this gap through end-to-end, collaborative worker involvement in defining and operationalizing evaluation concepts.

  • 2.1 AI Evaluation from a Social Science Measurement Perspective: AI evaluations rely on assumptions about what matters and how latent concepts should be measured, yet many benchmarks have poor construct validity.
  • 2.1 AI Evaluation from a Social Science Measurement Perspective: Comparatively little work meaningfully involves non-technical stakeholders in designing LLM measurement instruments.
  • 2.2 Supporting Worker Participation in AI Measurement and Evaluation: This paper explores end-to-end stakeholder participation in identifying meaningful concepts, systematizing them, and operationalizing them into concrete instruments.
  • 2.2 Supporting Worker Participation in AI Measurement and Evaluation: Prior worker-involving evaluations often assess autonomous task completion rather than how LLMs support human workers’ activities.
  • 2.2 Supporting Worker Participation in AI Measurement and Evaluation: Existing efforts typically limit workers to isolated activities or individual input instead of collaborative iteration across measurement design stages.

3 Methods

The study used an exploratory, collaborative process with a Pennsylvania school social work organization to support workers’ participation in benchmark design. Workers moved from identifying use cases and test cases to defining measurement goals, refining a rubric, and validating the resulting benchmark, while the case study remained limited in scope and generalizability.

  • Study setting: The collaboration involved a Pennsylvania school social work organization exploring LLM support for workers’ practices.
  • Study setting: Over eight weeks, 19 workers participated in an iterative series of eight workshops on worker participation in benchmark design.
  • Design framework: The design process treated AI evaluation as three phases: identifying meaningful use cases, defining measurement goals, and operationalizing them into instruments.
  • Participation principles: The activities emphasized expressive participation, allowing workers to shape foundational measurement decisions without predefined categories and with opportunities for collaboration.
  • Phase 1: Use cases and test cases: Workers first documented current and desired LLM uses, discussed six candidate use cases, and selected a focus for benchmark development.
  • Phase 1: Use cases and test cases: Workers then designed realistic test cases based on their experience and the situations in which they wanted valuable LLM support.
  • Phase 2: Measurement goals and systematization: By examining model-response pairs, workers collaboratively identified desirable and undesirable properties and synthesized high-level measurement goals.
  • Limitations: The exploratory case study is limited in scope and generalizability because the partner organization was small and the workflow may not generalize to larger organizations.

4 Designing a Benchmark with Workers

Workers designed a benchmark around “people challenging”: an LLM should support reflection, expose assumptions and biases, and promote learning rather than prescribe solutions. Through iterative rubric refinement and validation, they translated this goal into context-specific criteria that aligned strongly with worker judgments and distinguished model performance.

  • Identifying meaningful measurement goals: Workers defined effective “people challenging” as prompting reflection on assumptions and biases while supporting perspective expansion, learning, and growth.They contrasted this with “people pleasing,” in which models give agreeable responses rather than challenge the writer.
  • Designing realistic test cases: Workers designed 16 realistic test cases from day-to-day classroom observations to ground rubric discussions in challenging, context-specific scenarios.The cases included workers’ subjective interpretations and direct reactions because they viewed LLMs as aids for self-reflection rather than prescription.
  • Operationalizing worker goals: The rubric expanded from five criteria to ten as workers examined LLM-judge interpretations and refined definitions, examples, conditions, and edge cases.The final rubric contained twice as many criteria and four times as many words as the initial version.
  • Operationalizing worker goals: The worker-designed rubric differed from sycophancy benchmarks by defining desirable responses in a specific work context rather than only measuring behaviors to avoid.Workers used professional and lived expertise, together with LLM-judge examples, to discover nuances in how they conceptualized each criterion.

5 Discussion

The discussion frames worker-driven AI measurement as a complementary, bottom-up approach that grounds evaluation in workers’ goals, expertise, and organizational contexts. It also identifies potential reflective benefits and directions for improving the design process.

  • Implications: Workers’ direct participation in systematization and operationalization was essential for translating context-specific concepts into usable measurement criteria.The paper argues that workers’ professional and lived expertise captured nuances that would be difficult to operationalize abstractly.
  • Implications: Worker-driven AI measurement complements broader evaluation paradigms by letting workers define what successful augmentation means and how it should be measured.The approach is intended to reflect worker priorities across occupations while supplementing, rather than replacing, other evaluation goals.
  • Potential benefits: Workers expressed greater recognition of the value of their human work and more critical views of AI performance after the workshops.These are preliminary signs of possible benefits beyond the resulting measurement instrument.
  • Limitations: The study cannot support causal claims explaining the observed attitude shifts, so future work should investigate these potential process benefits.The authors distinguish observed changes from evidence about why those changes occurred.
  • Future work: Future interfaces could preview rubric edits, generate edge-case test cases, and stress-test designs while leaving workers in control of refinement decisions.These tools are proposed to make better use of workers’ time without displacing their expertise.
  • Implications: Practitioners should involve workers in identifying measurement goals, systematizing and operationalizing them, sharing instruments, and potentially supporting critical AI literacy.The recommendations position worker participation across the full measurement-design process.

6 Conclusion

The paper proposes worker-driven AI measurement as a complementary evaluation approach in which workers collaboratively define successful augmentation and its measurement. In a school social work case study, workers built a benchmark for whether an LLM meaningfully challenges their thinking in daily work.

  • 6 Conclusion: Worker-driven AI measurement is a complementary approach in which workers collaboratively shape what successful augmentation means and how it is measured.It is presented as a complement to existing AI evaluation practices.
  • 6 Conclusion: In an exploratory school social work case study, workers developed an LLM-as-a-judge benchmark and test cases for operationalizing meaningful challenges to their thinking.The benchmark targeted a specific use case within workers’ day-to-day practice.
  • 6 Conclusion: Workers drew on professional and lived expertise grounded in their organizational context to identify, define, and operationalize measurement goals.The conclusion emphasizes workers’ role in shaping the measurement instrument.

A.1 Phase 1: Selection of use case

The use-case selection prioritized work that mattered to social workers, was frequently used, addressed concerns about response quality, and was sufficiently bounded for studying worker involvement in evaluation.

  • A.1 Phase 1: Selection of use case: The selected use case was intended to be important to workers and frequently used in their practice.This helped ensure that the benchmark addressed a meaningful work activity.
  • A.1 Phase 1: Selection of use case: Selection also considered workers’ concerns about whether GPT understood their roles, avoided prescriptive or overly positive responses, and met the study’s scope requirements.Options involving mainly multi-turn conversations or extensive prior context were eliminated.

A.2 Phase 2: Supporting workers in creating test cases and annotating model responses

Workers created concise, self-contained test cases and evaluated model responses through iterative workshop activities that supported rubric development and validation. The process combined varied model outputs, worker annotation, behavioral principles, and edge-case coverage.

  • Test-case design: Workers were instructed to make test cases brief and self-contained, avoiding dependence on information from earlier messages.This design supported more controlled evaluation of model responses.
  • Response annotation: Workers inspected model responses and annotated desirable and undesirable properties, with the activity shifted from online tools to paper after accessibility and engagement concerns.The change preserved response-review activities while making participation more accessible.
  • Model-response generation: Model responses were generated from varied base models and system-instruction settings to support worker ideation and rubric refinement.The response-generation procedure mixed older and newer models with minimal or customized instructions.
  • Validation: The benchmark validation compared worker judgments with an LLM judge and assessed differentiation across six state-of-the-art LLMs using repeated scored responses.The second validation averaged five LLM-judge scores for each model response.
  • Rubric principles: Workers’ rubric emphasized strengths-based collaboration, reflection rather than directives, multiple perspectives, assumptions, personal experience, role, and potential biases.It also specified supportive tone, deeper learning, and avoidance of people-pleasing or superficial compliments.
  • Rubric principles: The rubric instructed the model to avoid treating others’ biases as the writer’s own and to avoid prompting reflection when the writer already demonstrates awareness.These constraints distinguish appropriate bias reflection from indiscriminate labeling.
  • Rubric principles: Workers supported deeper learning through concrete external resources and maintained collaborative inquiry while treating actors in the situation as equals.The desired stance avoided directive language and over-centering the writer.
  • Rubric refinement: Before rubric refinement, additional test cases were added so each binary rubric assertion had examples both awarding and not awarding a point.This was designed to provide edge-case coverage during iterative refinement.

A.6 Additional participant information

The supplementary participant information covers demographics and workshop participation across the study.

  • Table 5 summarizes the demographics of the 17 workers who responded to the demographics survey.
  • Table 6 reports worker participation across engagements and the three study phases, including workshop length.

B Additional Findings

Additional findings describe how workers refined criteria for an LLM to serve as a reflective aid, including attention to stakeholders’ perspectives and the writer’s role.

  • Workers split reflection on stakeholders’ perspectives into criteria covering relevant perspectives and all possible perspectives.
  • Workers included a criterion requiring the LLM to prompt reflection on the writer’s own role when the original prompt omitted it.

B.2 Additional details on workers’ experiences

Workers’ experiences included learning how peers use LLMs, becoming more critical of AI information, and finding that a fully automated process would engage them less effectively.

  • Workers attended partly to learn how peers use LLMs, after attributing poor AI performance to their own prompting practices.
  • MD = −0.30 in agreement that AI produces correct information, indicating workers became slightly more critical after the workshops.
  • MD = −0.40 in concern about AI replacing workers’ skills, reflecting reduced concern after workshop engagement.
  • The organization’s director reported that a fully automated process would not have engaged workers as effectively.

C Benchmark Components

The benchmark components include worker-designed test cases, an initial rubric, and a final rubric refined through iterative operationalization.

  • The supplementary materials include examples of test cases, the initial worker-developed rubric, and the final iteratively refined rubric.

C.1 Test Case Examples

The test cases present realistic school social work situations and show how workers’ rubric evaluates whether an LLM challenges assumptions, supports reflection, and avoids superficial or over-centered responses.

  • Test Case Examples: Workers use detailed classroom scenarios involving children’s behavior, teacher relationships, family context, and their own observations as realistic test cases.
  • Initial Rubric: The rubric uses a five-point scale from “Not at all” to “Very much” to assess how strongly a response reflects each criterion.
  • Final Rubric: The rubric asks models to challenge assumptions about people’s intentions and consider alternative interpretations of their actions.
  • Final Rubric: Additional criteria ask writers to reflect on their role, assumptions, and subjective experiences rather than reporting only others’ interpretations.
  • Final Rubric: A core criterion requires the model to identify potential biases in a writer’s interpretation and invite reflection when the prompt contains such biases.
  • Final Rubric: The rubric also discourages superficial compliments, over-centering the facilitator, and unsupported prompts for external resources.
Loading 2608.22459v1…