Source-linked AI summary

Generative AI at Work

Erik Brynjolfsson, Danielle Li, Lindsey Raymond

arXiv:2304.11771v2econ.GNq-fin.GN

TL;DR

The paper addresses limited evidence on the workplace effects of generative AI by studying a staggered conversational-assistant rollout among 5,172 customer-support agents. Using this deployment and agent-level panel data, it finds that AI access raises productivity on average, with heterogeneous effects across workers and additional evidence on learning, fluency, and work experience.

  • Problem

    Because generative AI technologies are only beginning to be used in workplaces, little is known about their impacts.

  • Method

    The study uses a staggered rollout of an AI customer-support assistant and agent-level panel data to compare outcomes before and after access.

  • Results

    15% is the average increase in productivity, measured as issues resolved per hour, after AI assistance becomes available.

  • Takeaways & Limitations

    AI assistance improves worker productivity on average, but its effects vary across workers and may be strongest where baseline human training and experience are limited.

  • Takeaways & Limitations

    The reported resolution-rate increase is economically modest and statistically insignificant against an 82% baseline.

Abstract

from arXiv · show

We study the staggered introduction of a generative AI-based conversational assistant using data from 5,172 customer support agents. Access to AI assistance increases worker productivity, as measured by issues resolved per hour, by 15\% on average, with substantial heterogeneity across workers. Less experienced and lower-skilled workers improve both the speed and quality of their output while the most experienced and highest-skilled workers see small gains in speed and small declines in quality. We also find evidence that AI assistance facilitates worker learning and improves English fluency, particularly among international agents. While AI systems improve with more training data, we find that the gains from AI adoption are largest for relatively rare problems, where human agents have less baseline training and experience. Finally, we provide evidence that AI assistance improves the experience of work along two key dimensions: customers are more polite and less likely to ask to speak to a manager.

1 Generative AI and Large Language Models

Generative AI uses machine learning to generate new content from patterns in existing data. Large language models combine scale, architecture, pre-training, and fine-tuning to perform increasingly broad language and sequential-data tasks, with implications for work and inequality.

  • Generative AI generates new content by analyzing patterns in existing data, including text, images, music, and video.
  • Model quality depends strongly on computing power, parameter count, and dataset size, while fine-tuning adapts general-purpose models to specific applications.
  • Large language models use positional encoding and self-attention to capture long-range semantic relationships while processing text segments in parallel.
  • LLMs are pre-trained on large unlabeled corpora by predicting the next word, allowing them to learn grammatical and semantic relationships without explicit guidance.
  • Generative AI may perform non-routine tasks involving judgment and tacit knowledge, potentially altering relationships among technology, productivity, and inequality.

2 Our Setting: LLMs for Customer Support

The paper studies generative AI in customer support, a setting with substantial productivity variation, high turnover, costly training, and extensive conversational data. The assistant augments agents with response suggestions and technical documentation trained on customer-service interactions and top-performer behavior.

  • The setting has substantial worker productivity variation, high turnover, and costly training, creating practical incentives to use AI tools.Industry estimates cited in the paper put annual turnover at 60% and replacement costs at $10,000 to $20,000 per agent.
  • Customer support combines technical problem solving with managing frustrated customers, making both product knowledge and communication important.
  • Training on top performers is intended to encode subtle practices such as clarifying questions, de-escalation, adaptable communication, and simple explanations.
  • The AI system generates real-time response suggestions and links to internal technical documentation based on historical agent conversations and outcomes.
  • The system augments rather than replaces agents, who see the recommendations and retain discretion over whether to use them.

3 Deployment, Data, and Empirical Strategy

The study analyzes a staggered, individual-level AI rollout among customer-support agents using agent-month data and difference-in-differences methods. The dataset measures productivity, handling, resolution, satisfaction, sentiment, topics, and language, though quality outcomes are available for a smaller subset.

  • 3.1 AI Rollout: AI access was staggered because limited training resources, licensing constraints, and budget limits created variation in adoption timing.
  • 3.1 AI Rollout: Managers assigned workers to different training sessions within teams to maintain customer-service coverage, producing effectively individual-level rollout timing.
  • Data: The sample contains 5,172 agents and 3 million chats, including 1.2 million post-AI chats from 1,636 agents; 89% of agents are outside the United States.
  • Empirical Strategy: The preferred empirical strategy uses difference-in-differences with year-month, agent, location, and tenure fixed effects to estimate the effect of AI access.
  • Data: Resolutions per hour summarizes productivity, while average handle time, chats per hour, resolution rate, and net promoter score measure its components and customer satisfaction.
  • Data Limitations: Quality outcomes are observed only for a smaller subset because subcontractors do not consistently collect call-quality metrics for all agents.

4 Main Results

AI assistance substantially increases customer-support productivity, with effects appearing immediately and persisting through the sample. Gains reflect faster chats and more multitasking, while average resolution quality and customer satisfaction show little economically significant change.

  • 4.1 Overall Impacts: 0.30 chats per hour is the estimated increase in resolutions per hour after controlling for agent and tenure fixed effects.The less-controlled estimate is 0.47 chats, or 23.9% relative to the pre-treatment mean of 1.97.
  • 4.1 Overall Impacts: AI increases productivity immediately in the first deployment month, with a slightly larger effect in the second month that remains stable thereafter.
  • 4.1 Overall Impacts: 15% is the approximate increase in chats per hour relative to a baseline mean of 2.4, alongside a 3.7-minute reduction in average chat duration.The duration decline is 8.5% from a baseline mean of 43 minutes.
  • 4.1 Overall Impacts: The stronger chats-per-hour effect suggests that AI helps agents both speed up chats and handle multiple chats simultaneously.
  • 4.1 Overall Impacts: 1.3 percentage points is the estimated increase in chat resolution rates, but the effect is economically modest and insignificant against an 82% baseline.
  • 4.1 Overall Impacts: Customer satisfaction does not change significantly, with a net promoter score coefficient of -0.12 percentage points.

Appendix Figure A.3 presents the accompanying event studies for additional outcomes. We see

Event-study and pilot evidence indicates that AI assistance increases productivity while leaving resolution rates and customer satisfaction broadly unchanged on average.

  • AI assistance produces immediate reductions in average handle time and increases in chats per hour.Event-study patterns are relatively flat for resolution rate and customer satisfaction.
  • The average productivity gains do not come with negative effects on resolution rates or surveyed customer satisfaction.
  • A pilot analysis involving approximately 50 workers also finds no average impact on resolution rate or customer satisfaction.
  • The pilot treatment group of 22 workers shows significant increases in resolutions per hour and decreases in average handle time.The estimated magnitudes are similar to the main results.
  • Instrumenting individual adoption with company, office, and team rollout timing estimates 0.55 additional resolutions per hour, compared with 0.30 under the main specification.Effects for average handle time and chats per hour are essentially identical to the main estimates.

Appendix Table A.9 finds similar results using alternative difference-in-difference estimators in-

Alternative estimators, clustering choices, weighting schemes, and subgroup analyses generally preserve the paper’s central finding: AI assistance raises productivity, especially among less-skilled and less-experienced workers.

  • Robustness checks: Alternative difference-in-difference estimators generally produce similar or larger estimated effects of AI assistance.These estimators avoid comparisons between newly treated and already treated units.
  • Robustness checks: Standard errors are similar when clustered at the individual, team, or geographic-location level.The analysis clusters at the individual level to reflect individual variation in rollout timing.
  • Robustness checks: Reweighting worker-month observations by the number of customer chats produces results similar to the main estimates.
  • Skill heterogeneity: 36% increase in resolutions per hour occurs for workers in the lowest skill quintile, while the most skilled workers show no productivity increase.
  • Experience heterogeneity: Agents with less than one month of tenure improve resolutions per hour by 0.7 resolutions per hour, with larger effects among less-experienced workers.
  • Experience heterogeneity: Workers with AI access reach approximately 2.5 resolutions per hour after two months, compared with 8 to 10 months for never-treated workers.Access to AI recommendations therefore helps workers move more quickly down the experience curve and reduces ramp-up time.

5 Adherence, Learning, Topic Handling, and Conversational Change

AI assistance is associated with selective recommendation use, durable learning, stronger gains on less routine problems, improved English fluency, and convergence in communication patterns across skill levels.

  • Adherence: 35% of AI recommendations are followed on average, and returns to AI assistance are highest among workers who follow recommendations more closely.Adherence increases over time, especially among initially skeptical and more experienced workers.
  • Learning: Workers exposed to AI recommendations continue to perform better during software outages, with larger effects after more exposure and among closer followers.These patterns provide evidence of durable changes in worker skills.
  • Topic handling: AI access is particularly beneficial when workers encounter less routine issues, where they might otherwise struggle to find solutions.
  • Conversational change: AI assistance improves workers’ English language skills, with the largest gains among those with the lowest initial proficiency.
  • Conversational change: After AI adoption, lower-skilled workers show larger textual shifts and their conversations become more similar to those of higher-skilled workers.The findings suggest convergence toward practices associated with higher-skill and more experienced agents.
  • Adherence: 10% gain in productivity appears among agents in the lowest adherence quintile, compared with an estimated impact close to 25% in the highest quintile.The positive relationship is strongest for average handle time and chats per hour, and noisier for resolution rate and customer satisfaction.

Appendix Figure A.15 shows an example of such an outage, which occurred on September 10,

Outage analyses compare chats with and without direct AI suggestions and indicate that exposure produces increasingly durable productivity benefits, especially for workers who engage with recommendations.

  • Outage setting: During non-outage periods, 30-40% of chats typically lack AI recommendations, while one documented outage increased that share to almost 100%.
  • Outage analysis: 10% to 15% decline in individual chat duration appears during post-adoption periods without reported outages.
  • Outage analysis: During outage periods, estimated chat-duration declines are equivalent to 15% to 25%, but the estimates are noisy.Outages are rare and not necessarily random, so outage-period chats may differ from non-outage chats.
  • Learning: The benefit of prior AI exposure during outages increases with time since adoption rather than declining immediately and remaining stable.After three months of exposure, workers handle chats faster even without receiving direct AI assistance.
  • Learning: Workers with high initial adherence experience rapid declines in chat-processing times during outages, while frequent deviators show no reduction even after prolonged access.
  • Learning: The findings suggest that actively engaging with AI suggestions and observing customer responses contributes to worker learning.AI assistance may therefore supplement existing on-the-job training with specific, real-time, actionable suggestions.

Appendix Figure A.16 reports the distribution of conversation topics in our dataset. Unsur-

AI assistance has its largest productivity effects on moderately rare customer problems, where agents have less direct experience but the AI has sufficient training exposure. The gains are therefore greatest when AI capabilities complement gaps in workers’ baseline skills.

  • Topic rarity: AI assistance has a non-monotonic relationship with topic rarity, producing smaller reductions for very routine and very rare problems.The largest handle-time improvements occur for somewhat uncommon problems.
  • Topic rarity: 4 to 5 minutes: access makes agents faster on the most routine payroll and account-management problems.This is approximately a 10% decrease from the pre-treatment mean duration for these topics.
  • Interpretation: Productivity gains depend on AI capabilities relative to workers’ baseline skills, not only on the AI system’s absolute capabilities.The findings indicate that AI complements or exceeds human capabilities most effectively for somewhat uncommon problems.
  • Agent-specific experience: 15% versus 10%: AI reduces conversation times more for problems agents encountered least often than for those they encountered most often.This comparison holds after controlling for the problem’s overall frequency.
  • Agent-specific experience: AI assistance is most valuable for problems to which a specific agent has the least exposure, holding the AI model’s exposure constant.The marginal value of AI is highest where humans have greater need for AI input.

6 Effects on the Experience of Work

AI assistance improves several dimensions of workers’ experience: customers become more positive, manager requests decline, and attrition falls, especially for newer workers. These benefits are accompanied by heterogeneous effects across worker skill and experience, with the strongest customer-sentiment gains among lower- to middle-ranked agents.

  • Measures: The study measures workplace experience through customer and agent sentiment, managerial-escalation requests, and turnover.Turnover serves as a broad response measure because not all aspects of stress, customer perception, and job satisfaction are directly observed.
  • Customer sentiment: 0.18 points: AI access increases mean customer sentiment, equivalent to half a standard deviation.Agent sentiment changes by only 0.02 points, or about 1% of a standard deviation, from an already high baseline.
  • Customer sentiment: AI significantly improves how customers treat agents across skill and experience levels.The largest effects occur among agents in the lower to lower-middle ranges of skill and tenure, while the highest-performing and most-experienced agents benefit least.
  • Managerial escalation: Almost 25%: AI assistance reduces customer requests to speak to a manager relative to an approximately 6-percentage-point baseline.Requests for escalation decline gradually after AI introduction, with point estimates suggesting larger reductions for less-skilled or less-experienced agents.
  • Worker turnover: 10 percentage points: AI access is associated with a 40% decrease in attrition among agents with less than six months of experience.The comparison uses a baseline attrition rate of 25% for this group.
  • Worker turnover: AI access significantly decreases attrition for all skill groups, without a clear gradient.The authors caution that these results are less definitive because agent fixed effects cannot be included and treatment may target agents more likely to stay.

7 Conclusion

The paper finds that generative AI assistance increases worker productivity and improves aspects of workers’ and customers’ experience, while effects vary across workers and remain subject to important scope and equilibrium limitations.

  • Main findings: 15%: AI assistance increases worker productivity, measured by issues resolved per hour.The authors report that the firm could handle the same number of support issues with 12% fewer worker-hours.
  • Main findings: Productivity gains are larger for lower-skill and novice agents, indicating substantial heterogeneity within customer support work.The paper also reports minimal productivity impacts for workers with more than six months of tenure.
  • Worker experience: AI assistance improves workers’ on-the-job experience, including customer sentiment and confidence, and is associated with reduced turnover.The authors report that improved treatment by customers is associated with the strongest reductions in attrition among newer agents.
  • Limitations: The findings are limited to one AI tool used in one firm and one occupation, with a relatively stable product and set of technical support questions.The authors caution that effects may differ in rapidly changing environments, where recommendations could either synthesize new practices or promote outdated ones.
  • Labor-market implications: AI adoption may affect labor demand in opposing ways: productivity gains could reduce demand when customer demand is inelastic, while improved support experiences may increase demand.The paper also notes possible new roles for agents testing and training AI models.

A.2.5 Language Comprehensibility and Fluency

The paper measures agent comprehensibility and native fluency from customer-service transcripts using LLM-based scores validated against human evaluations. AI assistance is evaluated separately for general clarity and native-like English ability.

  • Measurement: Gemini Pro scores agent comprehensibility and native fluency on 1–5 scales from written customer-service transcripts.Comprehensibility measures clarity and ease of understanding regardless of native-speaker status; native fluency assesses native-like American English.
  • Measurement: The native-fluency assessment considers grammar, vocabulary, idioms, natural phrasing, and cultural reference points in American English.The rubric ranges from definitely not native to definitely native American English speaker.
  • Validation: Human validation found no statistically significant difference between LLM and human native-fluency scores: 4.29 versus 4.22.Two independent human reviewers evaluated 100 randomly selected agent conversations.
  • Interpretation: The comprehensibility measure distinguishes fluent, understandable writing from native-speaker status.For example, an agent can write highly fluent English without using idiomatic language typical of native speakers.
  • Validation: Human validation found no statistically significant difference between LLM and human comprehensibility scores: 4.31 versus 4.40.The validation used a random sample of 100 agent conversations evaluated by human reviewers.

A.2.6 Conversation Topic

The paper classifies customer-support conversations into topic categories using LLM-assisted grouping and labeling, with human validation. Topics are highly concentrated, and the resulting categories support analyses of conversation content and worker skill.

  • Topic distribution: 50% of chat topics are concentrated in payroll and taxes, account access, and management issues.The next five topics account for another 25%, while the top 16 topics account for over 90% of chats.
  • Validation: Human evaluators unanimously agree on a topic in 30% of sampled conversations and reach modal agreement in 75%.The remaining 25% are discordant conversations in which each evaluator selects a different topic.
  • Validation: The LLM matches human topics 87% of the time under unanimous agreement, 66% under modal agreement, and 74% when humans disagree.These comparisons assess agreement with human classifications under three levels of reviewer consensus.
  • Analytic use: Topic categories are ranked by overall frequency and by their frequency for each individual agent.The paper uses these classifications to characterize the subject matter of customer-support conversations and worker skill.

B Average Productivity Effects

The productivity analysis estimates how AI deployment changes several agent and customer-service outcomes using staggered-adoption event studies and difference-in-differences. The main outcome is resolutions per hour, alongside handling time, throughput, resolution, and satisfaction measures.

  • Outcomes: The analysis examines average handle time, chats per hour, resolution rate, and customer satisfaction alongside resolutions per hour.The event-study outcomes include handling duration, multitasking throughput, successful resolution share, and net promoter score.
  • Estimation: Event-study regressions estimate deployment effects with agent, chat year-month, and agent-tenure fixed effects.The regressions use agent-month observations and cluster robust standard errors at the agent level.
  • Pilot RCT: A pilot randomized-control analysis compares 22 AI-treated agents with pre-treatment observations while controlling for agent, time, and tenure effects.The specific control agents are unavailable, so the comparison uses all pre-treatment agents.
  • Robustness: The paper checks robustness using alternative clustering levels and multiple robust difference-in-differences estimators.Alternative specifications include company/team, location, and agent-level clustering, plus several dynamic treatment-effect estimators.
  • Estimation: Difference-in-differences specifications estimate AI deployment effects on productivity and agent performance using agent-month observations weighted by chat counts.The table covers average handle time, chats per hour, resolution-related outcomes, and other performance measures.

C Impacts by Agent Skill and Tenure

AI assistance has uneven productivity effects across workers: impacts vary by pre-AI skill and tenure, with worker experience curves and adoption cohorts providing additional comparisons.

  • Productivity effects are evaluated across average handle time, chats per hour, resolution rate, and customer satisfaction.
  • Skill-based analyses compare AI impacts across worker-skill quintiles while controlling for tenure and fixed effects.
  • Tenure-based analyses compare average handle time, chats per hour, resolution rate, and net promoter score by tenure at AI deployment.
  • Experience curves compare always-treated, never-treated, and initially untreated agents across tenure and productivity measures.

D Adherence to AI suggestions

The section examines how agents follow AI suggestions and how adherence relates to performance, outages, conversational topics, and adoption experience.

  • AI adherence is measured as the share of AI recommendations followed by agents.
  • Adherence distributions are reported overall and after adjusting for company, location, and team fixed effects.
  • Performance impacts are analyzed by initial-adherence quintile using measures including average handle time, chats per hour, resolution rate, and NPS.
  • Adherence over time is compared by initial adherence, agent tenure, and pre-deployment productivity.
  • A documented software outage is used to identify post-treatment chats without AI suggestions and compare chat-duration effects.
  • Chat-duration impacts are also estimated by conversational-topic commonality.

G Language Fluency

The section evaluates AI assistance alongside language fluency, comprehensibility, customer sentiment, manager requests, and attrition.

  • Language outcomes are measured with native fluency and comprehensibility scores for pre-AI, post-AI, and never-AI agent-month observations.
  • Post-AI treatment effects are 0.250 for native fluency and 0.241 for comprehensibility.
  • For US agents, the corresponding effects are 0.159 and 0.134; for Philippine agents, they are 0.251 and 0.234.
  • Work experience outcomes include customer sentiment, agent sentiment, customer sentiment by skill and tenure, manager assistance, and attrition.
Loading 2304.11771v2…