Source-linked AI summary
GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
Tyna Eloundou, Sam Manning, Pamela Mishkin, Daniel Rock
TL;DR
The paper asks how LLMs and the software built around them could affect work across the U.S. labor market. It develops an exposure rubric applied by human annotators and GPT-4, finding that most workers have occupations with exposed tasks and that 19% have over half their tasks exposed when LLM-powered software is considered.
Problem
Evidence was limited on the breadth and scale of complementary technologies built around LLMs and their potential effects across occupations.
Method
The study applies a new occupational-exposure rubric to U.S. O*NET data using human annotators and GPT-4 classifications.
Results
19% of U.S. workers are in occupations where over half of tasks are exposed when considering LLMs and complementary tools and applications.
Takeaways & Limitations
LLMs and the software built around them exhibit potential characteristics of general-purpose technologies with broad economic, social, and policy implications.
Takeaways & Limitations
The findings focus on the United States and may not generalize to countries with different industrial, technological, regulatory, linguistic, or cultural contexts.
Abstract
from arXiv · showhide
We investigate the potential implications of large language models (LLMs), such as Generative Pre-trained Transformers (GPTs), on the U.S. labor market, focusing on the increased capabilities arising from LLM-powered software compared to LLMs on their own. Using a new rubric, we assess occupations based on their alignment with LLM capabilities, integrating both human expertise and GPT-4 classifications. Our findings reveal that around 80% of the U.S. workforce could have at least 10% of their work tasks affected by the introduction of LLMs, while approximately 19% of workers may see at least 50% of their tasks impacted. We do not make predictions about the development or adoption timeline of such LLMs. The projected effects span all wage levels, with higher-income jobs potentially facing greater exposure to LLM capabilities and LLM-powered software. Significantly, these impacts are not restricted to industries with higher recent productivity growth. Our analysis suggests that, with access to an LLM, about 15% of all worker tasks in the US could be completed significantly faster at the same level of quality. When incorporating software and tooling built on top of LLMs, this share increases to between 47 and 56% of all tasks. This finding implies that LLM-powered software will have a substantial effect on scaling the economic impacts of the underlying models. We conclude that LLMs such as GPTs exhibit traits of general-purpose technologies, indicating that they could have considerable economic, social, and policy implications.
1 Introduction
The study develops a rubric to measure LLM exposure across occupations and finds broad, uneven technical potential for affecting work. It emphasizes that complementary technologies may substantially expand these effects, while technical feasibility does not guarantee productivity or automation outcomes.
- Implications: Complementary technologies significantly expand the potential impact of LLMs, whose effects may remain pervasive even if new capabilities stop developing.The paper argues that these characteristics support viewing GPTs as general-purpose technologies, while noting that current measurements capture only what is technically feasible now.
- Methods: A new rubric measures tasks’ overall exposure to LLM capabilities and their potential effects on jobs.The measure estimates technical capacity to make human labor more efficient.
- Findings: Approximately 19% of jobs have at least 50% of their tasks exposed when considering current model capabilities.Exposure reflects technical capacity, while social, economic, regulatory, and other factors may prevent productivity or automation outcomes.
- Findings: Most occupations exhibit some LLM exposure, with higher-wage occupations generally showing higher exposure.Exposure varies across different types of work and contrasts with similar evaluations of overall machine-learning exposure.
- Findings: 60 to 72% of variation in the preferred AI exposure measure is associated with earlier technology exposure measures and wage controls.Thus, 28 to 40% of the variation remains unaccounted for by previous technology exposure measurements.
- Industry patterns: Information-processing industries exhibit high exposure, while manufacturing, agriculture, and mining demonstrate lower exposure.The connection between past-decade productivity growth and overall LLM exposure appears weak.
2 Literature Review
The literature review situates LLMs within advances in generative AI, emerging tool integration, and longstanding frameworks for technology’s labor-market effects. It motivates this study’s GPT-4 task-assessment method and its emphasis on complementary technologies and systems-level integration.
- Generative AI advances: Generative AI models have advanced through larger parameter counts, more training data, and improved training configurations for complex language-based tasks.
- Generative AI advances: Fine-tuning and reinforcement learning with human feedback have improved LLM steerability, reliability, utility, and ability to discern user intent.
- Tool integration: LLMs are increasingly framed as versatile building blocks for tools and integrated systems, although implementation may require substantial process reconfiguration across industries.
- Complementary technologies: Out-of-the-box LLMs can remain unreliable because of factual inaccuracies, biases, privacy concerns, and disinformation risks, while specialized workflows can help address these shortcomings.
- Labor-market literature: Research on AI and labor commonly uses skill-biased technological change and task-based automation to analyze how technological progress affects worker demand and occupational tasks.
- Study contribution: This study focuses specifically on LLMs and proposes using GPT-4 to assess task exposure and automation potential, complementing human scoring before aggregating results to occupations and industries.
3 Methods and Data Collection
The study combines O*NET task and activity data with BLS employment information, applies a 50% time-reduction exposure rubric, and aggregates human and GPT-4 annotations into occupation-level measures. Its projections are limited by subjective, potentially inconsistent, and forward-looking judgments about LLM capabilities and task performance.
- Data collection: The analysis uses O*NET 27.2 data covering 1,016 occupations, 19,265 tasks, and 2,087 Detailed Work Activities (DWAs).Tasks are occupation-specific work units, while DWAs are comprehensive actions that may link to multiple tasks.
- Data collection: Employment, wages, worker counts, projections, education, and training data come from BLS sources linked to O*NET through the BLS-recommended crosswalk.The study uses the 2020 and 2021 Occupational Employment series and BLS Labor Force Demographics.
- Exposure rubric: Exposure is defined by whether an LLM or LLM-powered system can reduce task or DWA completion time by at least 50% while maintaining equivalent quality.Direct exposure covers LLM use through ChatGPT or the OpenAI playground; LLM+ exposure includes complementary software that enables the threshold.
- Annotation and measures: The study collects human and GPT-4 annotations, aggregates DWA labels to tasks and then occupations, and constructs three exposure measures: α, β, and ζ.β equals E1 + 0.5*E2, while ζ equals E1 + E2; α corresponds to E1.
- Limitations: Human labeling may be biased because annotators familiar with LLMs were not occupationally diverse and often lacked expertise about the mapped occupations.This limitation can affect judgments of LLM reliability and effectiveness in unfamiliar occupational tasks.
- Limitations: GPT-4 classifications are sensitive to rubric wording, prompts, examples, detail, and definitions, while future applications and capabilities remain difficult to predict.Technological change, emergent capabilities, and shifting human perceptions may alter the projections’ accuracy and reliability.
4 Results
The results indicate that LLMs could substantially affect diverse U.S. occupations, with exposure averaging approximately 15% of tasks under α and exceeding 30% and 50% under β and ζ. Exposure generally rises with wages, education, and intermediate job-entry requirements, while varying little with employment levels and differing by occupational skills.
- Overall exposure: LLMs could significantly affect a diverse range of U.S. occupations, consistent with a key attribute of general-purpose technologies.The assessment focuses on task-level capabilities and does not cover total factor productivity or capital-input effects.
- Overall exposure: Approximately 15% of occupation-level tasks are directly exposed to LLMs under α, rising to over 30% under β and above 50% under ζ.Both human and GPT-4 annotations produce similar estimates, tagging 15% to 14% of total dataset tasks as exposed.
- Wages and employment: Higher wages are associated with increased exposure to LLMs, although low-wage occupations can also have high exposure and high-wage occupations can have low exposure.Human annotations estimate marginally lower exposure for high-wage occupations than GPT-4 annotations.
- Wages and employment: LLM exposure shows little correlation with current employment levels, with neither human nor GPT-4 ratings differing significantly across employment levels.The comparison aggregates overall exposure at the occupation level against the log of total employment.
- Occupational characteristics: Occupations requiring science and critical-thinking skills are less exposed, while programming and writing occupations are more exposed to current LLMs.The reported associations are strongly negative for science and critical-thinking skills and strongly positive for programming and writing skills.
- Barriers to entry: Exposure increases from Job Zone 1 through Job Zone 4 and remains similar or decreases at Job Zone 5 across α, β, and ζ.For occupations with greater than 50% β exposure, the reported worker shares are 0.00% in Job Zone 1, 6.11% in Job Zone 2, 10.57% in Job Zone 3, and 34.5% in Job Zone 4.
5 Validation of Measures
The paper’s LLM exposure measures generally align positively and significantly with prior software, AI, and machine-learning exposure measures, especially those using similar task-level approaches. Regression fit ranges from 60.7% to 72.8%, leaving substantial unexplained variance relative to earlier measurements.
- Methodological validation: The methodology primarily builds on the SML approach by evaluating overlap between LLM capabilities and worker tasks in O*NET.The paper compares its new measures with prior occupation-level exposure measures, including AI Occupational Exposure, Frey–Osborne, and SML measures.
- Cross-measure consistency: Results show generally positive and statistically significant correlations between LLM exposure measures and prior software- and AI-focused measurements.SML exposure scores have significant positive associations with the paper’s measures, indicating cohesion between studies using similar approaches.
- Statistical associations: Software, SML, and routine cognitive scores are positively and significantly associated with LLM exposure scores at the 1% level.Webb’s AI scores are also positive and significant at the 5% level, while the secondary overall-exposure prompt is not significant in columns 3 and 4.
- Interpretation of differences: Low correlations with AI Occupational Exposure and Frey–Osborne measures may reflect differences between ability- or occupation-level scoring and task-level exposure approaches.The paper attributes these differences to how AI capabilities are linked to worker abilities or occupation characteristics versus aggregated from DWA or task-level scoring.
- Model fit and limitations: 28–40% of variance remains unexplained relative to other measurements, with R^2 values ranging from 60.7% to 72.8%.The measure’s explicit focus on LLM capabilities may capture information about future LLM and LLM-powered software progress that earlier efforts lacked.
6 Discussion
The discussion frames LLMs as potential general-purpose technologies because they improve over time, may diffuse throughout the economy, and can generate complementary innovations. It also emphasizes that their effects depend on adoption, human trust and adaptation, policy preparedness, and further research beyond U.S. exposure estimates.
- General-purpose technology: LLMs may qualify as general-purpose technologies by meeting criteria of improvement over time, economic pervasiveness, and complementary innovation.The paper notes that the AI and machine-learning literature supports their improving capabilities, while the other criteria concern economy-wide diffusion and complementary applications.
- Complementary software: 0.42 is the difference in means across all tasks between α and ζ, illustrating exposure potential from complementary software beyond direct LLM exposure.The comparison measures the aggregate within-occupation exposure attributable to tools and software built on top of LLMs.
- Adoption conditions: Widespread adoption depends on bottlenecks including human confidence in model outputs, changed work habits, costs, flexibility, preferences, and incentives.The legal profession illustrates this dependence because usefulness may require professionals to trust outputs without independently verifying documents or conducting research.
- Policy implications: Potential economic disparity and labor disruption from LLM automation underscore the need for societal and policy preparedness.The paper links earlier automation technologies to adverse downstream effects and extends that concern to worker exposure in the United States.
- Limitations and future research: The study’s U.S.-only focus limits generalizability, while future research should examine adoption across sectors and occupations and model capabilities beyond exposure scores.The authors also note that GPT-4 vision capabilities were excluded from the direct LLM-exposure ratings.
7 Conclusion
The study finds that LLMs may broadly affect U.S. occupations, with higher-wage occupations generally having more highly exposed tasks. It frames LLMs as general-purpose technologies whose software-enabled advances could significantly affect economic activity, while broader consequences require further research.
- Most occupations show some exposure to LLMs, with higher-wage occupations generally presenting more tasks with high exposure.
- LLMs may have pervasive impacts across a wide swath of U.S. occupations and economic activities.
- Software and digital tools built on LLMs may significantly extend their effects across a range of economic activities.
- Further research should examine whether LLM advancements augment or displace labor and affect job quality, inequality, skill development, and other outcomes.
- A new rubric gauges tasks’ GPT exposure, particularly in the U.S. labor market.
A Rubric
The rubric classifies occupational tasks by whether an LLM, LLM-powered applications, or image capabilities can reduce completion time by at least half while preserving quality. It assumes an average-skilled worker with access to the LLM, existing tools, and common laptop-accessible technical equipment, but no additional physical tools.
- A.1 Exposure: The model handles text-input/text-output tasks whose context fits within 2,000 words, but cannot retrieve facts from the past year unless they are supplied in the input.These assumptions define the LLM’s baseline capabilities for exposure assessment.
- A.1 Exposure: Workers are assumed to have average role expertise, access to the LLM, existing task-related software and hardware, and common laptop-accessible technical tools.They do not have access to additional physical tools or materials.
- A.1 Exposure: E0 denotes no exposure when direct LLM access cannot halve task time at equivalent quality or when the task requires substantial human interaction or physical embodiment.The rubric specifically treats in-person demonstrations and equipment repair as E0 examples.
- A.1 Exposure: E1 denotes direct exposure when ChatGPT-like access alone can halve task time at equivalent quality, including writing, translation, summarization, document feedback, and document question-answering.The rubric also includes generating questions about a document.
- A.1 Exposure: E2 denotes exposure through LLM-powered applications when the LLM alone cannot halve task time but plausible additional software could do so.Examples include processing long documents, retrieving current facts, and integrating customized or live responses into workplace systems.
- A.1 Exposure: E3 denotes exposure enabled by image capabilities when an LLM system can view, caption, or create images and thereby significantly reduce task time.The image system cannot accept video or reliably extract highly detailed measurements from images.
- A.1 Exposure: The annotations illustrate E1 for applying theoretical expertise when relevant principles fit in text, E2 for reservation automation, E2 for contract negotiation tools, and E2 for medication prescribing with human judgment retained.These examples show how classifications depend on available capabilities, existing automation, tool adoption, and the need for a human final decision.
B O*NET Basic Skills Definitions … Process
The O*NET framework defines basic skills as capacities that facilitate learning or faster knowledge acquisition, alongside content skills for domain-specific work and process skills for acquiring knowledge and skill across domains.
- Basic Skills: Basic Skills are developed capacities that facilitate learning or the more rapid acquisition of knowledge.
- Content: Content skills provide background structures needed to work with and acquire more specific skills across different domains.
- Content: Reading Comprehension involves understanding written sentences and paragraphs in work-related documents.
- Content: Active Listening involves attending fully, understanding points, asking appropriate questions, and avoiding inappropriate interruptions.
- Content: Writing, Speaking, Mathematics, and Science cover effective communication and problem-solving through mathematical and scientific methods.
- Process: Process skills are procedures that contribute to the more rapid acquisition of knowledge and skill across varied domains.
- Process: Process skills include Critical Thinking, Active Learning, Learning Strategies, and Monitoring to evaluate alternatives, apply new information, select appropriate methods, and improve performance.
Cross-Functional Skills
The analysis focuses on Programming as the selected cross-functional skill, based on prior knowledge of LLMs’ coding abilities. Programming is defined as writing computer programs for various purposes.
- Cross-Functional Skills: Programming was the only cross-functional skill selected for analysis because of prior knowledge about models’ ability to code.The selection was methodological and based on the researchers’ existing knowledge of model capabilities.
- Cross-Functional Skills: The selected skill connects the cross-functional-skills analysis to LLM coding capabilities.This point synthesizes the skill selection rationale and the definition of Programming.
- Cross-Functional Skills: Programming refers to writing computer programs for various purposes.
C Industrial and Productivity Exposure
LLM exposure is present across nearly all industries, with substantial heterogeneity and especially high exposure in data processing, information processing, and hospitals. Exposure appears unrelated to productivity growth since 2012, offering little evidence that already fast-growing industries will gain disproportionately.
- Industrial Exposure: Exposure spans nearly all industries, with wide heterogeneity across employment-weighted 3-digit NAICS industries.Human raters and GPT-4 both assess industry exposure using the paper’s exposure rubric.
- Industrial Exposure: Data processing, information processing, and hospitals show high exposure under both human and GPT-4 assessments.The two methods generally agree on relative industry exposures.
- Productivity Exposure: Productivity growth since 2012 appears unrelated to exposure to LLM technologies.This pattern holds for both total-factor and labor productivity growth, with little relationship to current model-rated exposure.
D Occupations Without Any Exposed Tasks
The section identifies 34 occupations with no tasks labeled as exposed by any of the study’s measures. These occupations therefore form the complete set with zero measured task exposure.
- Occupations Without Any Exposed Tasks: 34 occupations had none of their tasks labeled as exposed by any measure.The table covers all occupations meeting this criterion.
- Occupations Without Any Exposed Tasks: The defining criterion was zero exposed tasks across every measure.No measure labeled any task in these occupations as exposed.
- Occupations Without Any Exposed Tasks: Table 11 presents the complete list of occupations meeting this no-exposure condition.The table contains all 34 occupations identified under the criterion.