Source-linked AI summary
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
Kunal Handa, Alex Tamkin, Miles McCain, Saffron Huang, Esin Durmus, Sarah Heck, Jared Mueller, Jerry Hong, Stuart Ritchie, Tim Belonax, Kevin K. Troy, Dario Amodei, Jared Kaplan, Jack Clark, Deep Ganguli
TL;DR
The paper addresses the lack of systematic empirical evidence on how AI systems are integrated into economic tasks. It analyzes millions of privacy-preserving Claude.ai conversations mapped to O*NET occupations and finds usage concentrated in software development and writing, while spanning many occupations and combining augmentation with automation.
Problem
Systematic empirical evidence is lacking on how AI systems are actually being integrated into different economic tasks, despite the importance of anticipating labor-market changes.
Method
The study uses Clio to analyze millions of Claude.ai conversations, classifying them across occupational tasks and skills mapped to the U.S. Department of Labor’s O*NET Database.
Results
AI use peaks in software development and technical writing, reaches at least a quarter of tasks in ∼36% of occupations, and splits between 57% augmentation and 43% automation.
Takeaways & Limitations
The framework provides broad, granular measurements of current AI use across economic tasks and interaction patterns.
Takeaways & Limitations
The findings are an imperfect snapshot because the seven-day Claude.ai sample may not represent longer-term usage, other providers, API data, or non-text AI use.
Abstract
from arXiv · showhide
Despite widespread speculation about artificial intelligence's impact on the future of work, we lack systematic empirical evidence about how these systems are actually being used for different tasks. Here, we present a novel framework for measuring AI usage patterns across the economy. We leverage a recent privacy-preserving system to analyze over four million Claude.ai conversations through the lens of tasks and occupations in the U.S. Department of Labor's O*NET Database. Our analysis reveals that AI usage primarily concentrates in software development and writing tasks, which together account for nearly half of all total usage. However, usage of AI extends more broadly across the economy, with approximately 36% of occupations using AI for at least a quarter of their associated tasks. We also analyze how AI is being used for tasks, finding 57% of usage suggests augmentation of human capabilities (e.g., learning or iterating on an output) while 43% suggests automation (e.g., fulfilling a request with minimal human involvement). While our data and methods face important limitations and only paint a picture of AI usage on a single platform, they provide an automated, granular approach for tracking AI's evolving role in the economy and identifying leading indicators of future impact as these technologies continue to advance.
1 Introduction
The paper introduces a large-scale framework for measuring real-world AI use across economic tasks and occupations. It finds concentrated use in software, writing, and analytical work, broader diffusion across occupations, and mixed augmentation and automation patterns.
- Measurement framework: The framework maps privacy-preservingly analyzed Claude.ai conversations to occupational categories in the U.S. Department of Labor’s O*NET Database.It is presented as an automated, granular, and empirically grounded approach for tracking evolving AI usage.
- Current AI usage patterns: AI use is highest in software engineering, writing-intensive, and analytical occupations, while physical-manipulation roles show minimal use.Examples of high-use roles include software engineers, data scientists, technical writers, copywriters, and archivists; examples of low-use roles include anesthesiologists and construction workers.
- Depth of occupational use: ∼36% of occupations show AI usage in at least 25% of their tasks, while only ∼4% show usage across at least 75% of tasks.The results indicate broad diffusion into task portfolios but deep use across most tasks in relatively few occupations.
- Skills represented in use: Cognitive skills such as Reading Comprehension, Writing, and Critical Thinking are highly represented in human-AI conversations, whereas physical and managerial skills show minimal presence.The reported pattern reflects complementarity between current AI capabilities and particular occupational skills.
- Automation and augmentation: 57% of interactions show augmentation patterns and 43% show automation-focused usage, with most occupations exhibiting a mixture of both.Augmentation includes back-and-forth iteration, while automation involves performing the task directly.
- Limitations: The authors identify limitations because usage data cannot show how outputs are used in practice and static O*NET descriptions cannot capture entirely new tasks or jobs.These limitations constrain interpretation of the framework as AI capabilities and work practices evolve.
2 Background and Related Work
Prior work models or forecasts AI’s labor-market effects, measures productivity in specific domains, or surveys adoption, but does not provide the same broad view of real-world task use. This paper bridges those approaches with large-scale analysis of conversations mapped to economic tasks and occupations.
- Task-based foundations: Economic research commonly uses a task-based framework in which individual tasks can be performed by human workers or machines.This framework underlies analysis of automation and labor-market change.
- Forecasting AI exposure: Forecasting studies use O*NET occupational descriptions and human or machine judgments to estimate future AI exposure or computerization.These approaches predict potential impacts rather than directly measuring current usage.
- Real-world adoption and productivity: Existing evidence also includes controlled productivity studies across domains such as software engineering, writing, customer service, consulting, translation, legal analysis, and data science.These studies provide domain-specific evidence rather than a comprehensive economy-wide account of where and how AI is used.
- Paper’s contribution: The paper differs from exposure forecasts by measuring real-world usage patterns through privacy-preserving analysis of millions of human-AI conversations.The approach is intended to reveal present usage trends and leading indicators of future diffusion.
3 Methods and analysis
The analysis maps privacy-preserved Claude.ai conversations onto O*NET tasks and occupations to characterize where and how AI is used. Usage is concentrated in computational, writing, and analytical work, while task coverage, skill representation, wages, collaboration patterns, and model specialization reveal selective and varied adoption.
- Data and classification: Clio classified millions of Claude.ai conversations across O*NET tasks, occupations, skills, and interaction patterns using aggregated, anonymized data collected in December 2024 and January 2025.The task-mapping process traversed a hierarchical tree created from nearly 20,000 O*NET task statements.
- Task-level analysis: Computer-related tasks had the highest AI usage, followed by writing tasks in educational and communication contexts, with usage grouped into O*NET occupations and broader categories.Computer and Mathematical occupations comprised 37.2% of all queries, while Arts, Design, Entertainment, Sports, and Media occupations comprised 10.3%.
- Task-level analysis: ∼36% of occupations showed AI usage in at least 25% of their tasks, but only ∼4% reached 75% or more, indicating selective rather than comprehensive integration.Examples include task-specific usage for Physical Therapists, Marketing Managers, and Foreign Language and Literature Teachers.
- Occupational skills: Cognitive skills including Critical Thinking, Reading Comprehension, Programming, and Writing were prevalent, whereas Installation, Equipment Maintenance, and Repairing were uncommon.The analysis measured whether skills appeared in Claude’s responses, not whether they were central to the user’s purpose or performed at expert level.
- Wage and barriers to entry: AI usage peaked in the upper wage quartile, especially for computational occupations, while usage was lower at both wage extremes and in occupations requiring advanced training or physical interaction.Computer Programmers and Web Developers were highly represented, whereas waiters and anesthesiologists were among the least represented.
- Automating vs. augmenting users: 57% of interactions showed augmentative patterns and 43% showed automative patterns, with most occupations exhibiting a mixture of both.Directive and Feedback Loop behaviors commonly involved content generation, coding, and debugging; Task Iteration and Learning included development, communication, and education tasks.
4 Discussion
The paper presents a large-scale empirical analysis of AI use across economic tasks and discusses what these usage patterns imply for measurement, occupational change, and future research. It emphasizes task-level diffusion, mixed augmentation and automation, and important limits on interpreting the findings.
- The study provides the first large-scale empirical analysis of how advanced AI systems are used across economic tasks.It uses real-world conversation data rather than only forecasts or controlled productivity studies.
- The sample may not represent longer-window Claude.ai usage, API data, or users of other model providers and modalities.The authors note differences in capabilities, features, user bases, and text-only output.
- The analysis cannot capture new tasks or jobs, because it relies on O*NET’s static occupational descriptions.This scope boundary limits analysis of transformations not represented in the database.
- The data cannot reveal how users incorporate Claude’s outputs into their workflows, preventing definitive judgments about practical uptake.Examples include copying code into development environments, incorporating writing into documents, or fact-checking responses.
- The framework is designed to track emerging AI usage patterns over time and provide earlier visibility into sectors approaching technological inflection points.The authors contrast this with surveys based on self-reported behavior.
- AI use is concentrated in subsets of tasks, suggesting occupations may evolve rather than disappear if this pattern persists.The authors distinguish task-level effects from wholesale automation of occupations.
- Current usage patterns do not definitively establish long-term productivity, displacement, or inequality effects.The authors identify longitudinal tracking of usage and outcomes as a future way to study these relationships.
5 Conclusion
The conclusion summarizes real-world evidence that AI use is concentrated in software development and technical writing but already spans a substantial share of occupations and tasks. It presents dynamic empirical measurement as a foundation for understanding and preparing for AI’s changing role in work, while noting that future capabilities may alter the pattern.
- AI usage peaks in software development and technical writing, with ∼4% of occupations using AI across three-quarters of tasks and ∼36% across at least one-quarter.Usage is split between 43% automation and 57% augmentation of human capabilities.
- The findings are an initial picture of AI’s integration into work, not a complete account of future human-AI collaboration.The conclusion points to video, speech, robotics, more autonomous agents, and newly emerging tasks or occupations as future changes.
- The authors position systematic empirical tracking as crucial for anticipating and preparing for the evolving landscape of work.They frame the challenge as both measuring changes and using that understanding to shape future outcomes.
- The framework analyzes aggregated Claude.ai conversations with Clio, a privacy-preserving system for studying human-AI interactions.The paper uses Clio to identify occupational tasks and characterize interaction patterns.
B.1 Usage Across Tasks and Occupations
The methodology maps Claude.ai conversations to O*NET tasks through a generated hierarchy, then connects tasks to occupations for aggregate usage analysis. The hierarchy enables classification across roughly 20,000 tasks while preserving broad occupational structure.
- The analysis uses a hierarchical classification because O*NET contains approximately 20,000 tasks that cannot fit into a single model context.Conversations are mapped through progressively organized task labels.
- The workflow has three components: build an O*NET task taxonomy, map privacy-preserved conversations to tasks with Clio, and connect tasks to occupations.This separates task organization, conversation classification, and occupation-level aggregation.
- Hierarchy construction embeds task names, clusters them into neighborhoods, proposes higher-level labels, and refines assignments across levels.Neighborhoods help keep task descriptions within the model’s context window, while deduplication and reassignment preserve coverage and distinctiveness.
- The task hierarchy organizes domains including business operations, healthcare, scientific research, and technical analysis.These examples illustrate the range of occupational categories represented in the generated structure.
- The generated hierarchy contains 12 top-level, 474 middle-level, and 19,530 base-level O*NET tasks.These levels organize the task database from broad categories to individual O*NET descriptions.
- For each occupationally relevant conversation, the method traverses the task hierarchy from broad categories toward specific O*NET tasks.The resulting task assignments are then linked to associated occupations for occupation-level analysis.
- Figure 10 reports minimal differences between occupational usage measured by conversation counts and by accounts.This figure compares the two measurement bases directly.
B.2 Conversation-level vs. Account-level Analysis
The paper tests whether conversation-level counting biases occupational usage patterns toward users with many short interactions and compares those patterns with account-level measurement. The reported occupational distributions remain stable across the two approaches.
- The analysis compares total conversations per task with account-based measurement to test whether interaction frequency biases occupational results.The concern is that programmers, for example, may generate many separate debugging conversations.
- AI usage patterns across occupations remain remarkably stable whether measured by conversations or accounts.The authors interpret this stability as evidence that variation in interaction patterns does not affect their findings.
- Occupational usage is related to median occupation wages and O*NET job zones describing preparation requirements.The analysis maps conversation counts to occupation-level wage data and uses five preparation zones.
- O*NET job zones are based on education, related experience, job training, and specific vocational preparation.The five zones range from little or no preparation to extensive preparation.
- The occupational-skill analysis uses 35 O*NET skills and allows each conversation to receive multiple applicable skill assignments.Conversations classified as exhibiting no listed skill are filtered out.
- The augmentation-versus-automation analysis intersects collaboration-pattern classifications with O*NET task mappings.This produces task-level breakdowns of interaction patterns.
- The sample contains approximately 54% Claude 3.5 Sonnet conversations and 46% Claude 3 Opus conversations.These model-usage proportions describe the analyzed conversation sample.
B.7 Validating the Composition of our Dataset
The dataset was screened to assess whether Claude conversations captured occupational tasks rather than predominantly personal, academic, or casual interactions. Non-work and coursework conversations formed minorities, and many remaining non-work interactions still mapped meaningfully to occupations.
- The validation examined conversation types across multiple dimensions to test whether the dataset captured genuine occupational activity.
- 23% of conversations were non-work after occupational task screening.Coursework comprised 5–10% of conversations, depending on the confidence threshold.
- Many conversations labeled non-work still mapped meaningfully to occupational tasks.Examples included personal nutrition planning relating to dietitian work and automated trading strategy development connecting to financial analyst tasks.
C Human Validation
Human validation assessed the task hierarchy and automation-versus-augmentation labels. Classification agreement was high overall, though accuracy declined at more specific task levels and the feedback-based sample may not represent regular Claude usage.
- Validation sample: The feedback-submission sample likely differs from regular Claude.ai usage, although the authors regard it as more reliable than alternative open-source datasets.
- Task hierarchy validation: 95.3% of conversations received an acceptable top-level task assignment.
- Task hierarchy validation: 91.3% of conversations received an acceptable middle-level task assignment.
- Task hierarchy validation: 86% of conversations received an acceptable base-level O*NET task assignment.Accuracy decreased as the hierarchy became more specific.
- Automation versus augmentation: 90.7% of conversations received their optimal automation-versus-augmentation label according to human raters.The labels distinguished Directive, Feedback Loop, Validation, Task Iteration, and Learning interactions.
D.1 Usage Across the Task Hierarchy
AI usage is concentrated in technology-related tasks, with creative and business activities also represented across the task hierarchy. Software development, debugging, and software modification are especially prevalent, while usage varies across occupational preparation levels.
- Top-level tasks: Nearly 50% of conversations involved top-level IT, technology, and related tasks.Creative and cultural work represented approximately 20%, while business management, finance, and customer service represented around 15%.
- Middle-level tasks: 14% of conversations involved software development and website maintenance, the most prevalent middle-level activity.Computer systems programming and debugging followed at roughly 11%, while several other categories accounted for 4–6% or 2–3%.
- Base-level O*NET tasks: Software modification and error correction dominated base-level O*NET tasks.Initial debugging, system administration, and hardware/software troubleshooting followed, with document editing and program analysis also represented.
- Overall pattern: The authors characterize current AI usage as heavily concentrated in technology and content creation roles, with varying use across other business functions.
- Barriers to entry: Job Zone 4 had a representation ratio of 1.50, while Job Zones 1 and 5 had lower ratios of 0.40 and 0.58.Job Zone 4 corresponds to occupations requiring considerable preparation, such as bachelor’s degree-level preparation.
E Analysis Metadata
The analysis uses prompts and metadata to screen conversations, assign occupational tasks, classify collaboration patterns, and preserve privacy through clustered analysis.
- Analysis Metadata: The analyses use specified Claude.ai conversation samples and time spans, including one million conversations for task usage and automation-versus-augmentation analyses.A separate 500,000-conversation sample supports usage-by-skills analysis.
- Screening: Conversations are screened for occupational relevance before classification into occupational tasks or skills.The screening prompt requires a Yes-or-No decision about whether a conversation possibly involves an occupational task.
- Collaboration Patterns: Clio classifies assistant behavior into collaboration patterns such as Directive, Feedback Loop, Task Iteration, Learning, and Validation.The prompt selects the most representative pattern and permits None when context is insufficient.
- Cluster Analysis: The primary analysis assigns conversations to best-fit O*NET tasks, while clustered analysis assigns privacy-preserving conversation clusters to tasks at different aggregation levels.Cluster analysis is presented as producing similar results to direct conversation analysis for many research questions.
G.1 Method and Sample
The study reconstructs O*NET task clusters from Claude conversation data and compares cluster-based assignments with direct conversation assignments across occupational categories and occupations.
- Method and Sample: Approximately 2.8 million Claude.ai Free and Pro conversations were clustered and assigned to best-fit O*NET tasks using the same hierarchical process as the primary analysis.The sample spans November 28 to December 18, 2024 and excludes conversations flagged as abusive or harmful.
- Occupational Categories: 0.95 Pearson, 0.95 Spearman, and 0.82 Kendall’s Tau-b are the minimum correlations between direct and cluster-based occupational-category usage.The comparison evaluates agreement across aggregation levels.
- Occupations: 0.70 Pearson, 0.58 Spearman, and 0.46 Kendall’s Tau-b are the minimum correlations between direct and cluster-based occupational usage patterns.Correlations include only occupations with associated tasks identified by at least one method.
G.4 Task Reconstruction
Task-level reconstruction is less stable than occupational reconstruction because granular O*NET tasks overlap and cluster aggregation recovers only a subset of the task inventory.
- Task Reconstruction: 0.47 Pearson, 0.22 Spearman, and 0.18 Kendall’s Tau-b are the minimum correlations between direct and cluster-based task usage patterns.The correlations cover tasks identified by at least one method.
- Task Reconstruction: Overlapping and highly similar O*NET task descriptions introduce intrinsic noise into task-level comparisons.The paper illustrates this with multiple troubleshooting and debugging task descriptions that can plausibly describe similar conversations.
- Task Reconstruction: Only a small fraction (<20%) of the ∼20K O*NET tasks are recovered in the cluster-based reconstruction.The recovered task sets vary across aggregation levels and overlap with direct-assignment sets.
- Task Reconstruction: Higher aggregation levels assign fewer average tasks per occupation, while direct assignment assigns an average of 4.8 tasks per occupation.Figures 26 and 27 compare task counts across aggregation levels.