Source-linked AI summary

Sophistication in GenAI Use: Field Evidence from a Large Firm

Nicholas J. Hallman, Zachary T. Kowaleski, Anu Puvvada, Jaime J. Schmidt

arXiv:2608.27364v1cs.AIecon.GN

TL;DR

Organizations need evidence on how sophisticated genAI use varies beyond adoption. This paper measures sophisticated use among back-office employees in a large firm and finds substantial differences across functions, with limited evidence of sustained change over time or after training.

  • Problem

    The paper examines how genAI sophistication varies across seniority levels, functional areas, time, and AI training in a widely adopting firm.

  • Method

    The study analyzes employee-month genAI activity using composite measures of prompt clarity, prompting strategies, and task-use diversity.

  • Results

    Sophistication is highest in Strategy, Digital Innovation, and Project Management, with no clear lasting improvement after formal AI training.

  • Takeaways & Limitations

    The findings indicate that organizations may face difficulty producing sustained changes in sophisticated employee genAI use.

  • Takeaways & Limitations

    The descriptive measures do not establish whether LLM outputs satisfied users’ goals or measure output quality and productivity effects.

Abstract

from arXiv · show

We study how sophistication in generative AI (genAI) use varies among the back-office workforce of a large firm. Using proprietary data, we observe 713,564 employee prompts and their corresponding large language model responses from nearly 4,000 back-office employees across 15 functional areas over eight months in 2025. We document three main findings. First, senior employees exhibit more sophisticated genAI use, consistent with domain expertise complementing genAI capabilities. Second, sophistication varies considerably across functions and is highest in Strategy, Digital Innovation, and Project Management, three groups that share a focus on firmwide strategic initiatives and organizational change. Third, we observe neither improvements in sophistication over time nor lasting improvements following formal AI training, suggesting that sophisticated use can be difficult to change. Together, our study provides measures of and insights into sophisticated genAI use that managers can use to improve outcomes and that researchers can use in future research.

1 Introduction

The paper defines sophisticated genAI use as skilled, informed interaction and develops measures to study how it varies across employees, tasks, and periods. It also provides evidence and benchmarking tools for managers seeking to improve genAI use and assess whether changes persist.

  • Concept and contribution: Sophisticated genAI use combines clear instructions, deliberate prompting techniques, and application across a broad mix of tasks.The definition reflects understanding of how genAI works and where it can be applied.
  • Adoption and use cases: 83.9% of employee-months include use of at least one genAI tool, compared with 32.1% in prior survey evidence.Copilot is used more frequently than aIQ Chat, while broad adoption masks substantial variation in usage intensity.
  • Adoption and use cases: Writing-related tasks appear in roughly three-quarters of aIQ Chat conversations, alongside knowledge retrieval, document understanding, software guidance, coding, and data analysis.These use-case categories form the basis for measuring how broadly employees apply the tool.
  • Concept and contribution: The primary analyses measure sophistication using use-case diversity, deliberate strategy use, and prompt clarity.The three measures correspond to broad task application, deliberate prompting techniques, and clear, specific instructions.
  • Managerial relevance: Because transcript-based measures are costly and rarely available, the paper examines their relationship with metadata-based measures of ambition, persistence, frequency, and flexibility.These observable characteristics capture how much employees ask the tool to do, how much they iterate, how often they use aIQ Chat, and whether they also use Copilot.
  • Managerial relevance: The study provides a customizable measurement process, descriptive field evidence, and benchmarking opportunities to evaluate organizational progress and inform resource allocation.The absence of improvement over time or sustained improvement following training illustrates the difficulty of fostering lasting changes in employee genAI use.

2 Background and Literature Review

Prior research shows that genAI adoption and use are uneven across workers, firms, functions, and tasks, and that organizational benefits depend on how employees engage with the technology. The literature therefore highlights variation in sophisticated use by seniority, training, and functional area.

  • Adoption and diffusion: Employee access to AI tools expanded by 50 percent in one year, yet genAI diffusion remains uneven across workers, firms, functions, and tasks.The cited research documents broad adoption alongside substantial variation in workplace diffusion.
  • Adoption and diffusion: 40 percent of workers used genAI frequently enough to realize reductions in email time and faster document completion in a randomized experiment.Frequent use appears necessary for achieving organizational benefits from genAI.
  • Common workplace uses: Writing, practical guidance, information seeking, software development, and other information-work activities dominate common workplace genAI uses, especially in professional work.Large-scale analyses identify writing, retrieval, analysis, and communication as central workplace applications.
  • Effective use: Information systems research distinguishes technology use from effective use, while prompt variation meaningfully affects outcomes on specific tasks.Organizational benefits depend on how users engage with genAI, not merely whether they access it.
  • Dimensions of sophisticated use: The literature organizes potential variation in sophisticated use around seniority, formal training, and functional area, reflecting mixed evidence on experience and the importance of training and workflow fit.AI adoption and use vary across functions, occupations, and tasks, while prior studies identify training as a lever for adoption and perceived value.

3 Research Design

The study uses proprietary January–August 2025 data from a large firm’s back-office workforce, combining aIQ Chat transcripts with employee-level usage and workforce characteristics. The analysis focuses on aIQ Chat conversations because Copilot interaction content is unavailable, with substantial and skewed observed activity.

  • Data and sample: The dataset covers back-office employees performing administrative functions such as IT, internal audit, strategy, public relations, marketing, accounting, and finance.Data were collected from January through August 2025 at a large professional services firm.
  • Data and sample: aIQ Chat provided access to commercial LLMs from several vendors, with GPT-4o as the default model.Available models included Anthropic’s Claude 3.5 Sonnet, OpenAI’s GPT-4o and o1, and models from Google’s Gemini and Meta’s Llama.
  • Measurement and analysis: The analysis links conversation transcripts and usage measures to employee seniority, functional area, and completion of firm-provided AI training.Conversation-level measures use only aIQ Chat because the content of Copilot interactions is not observed.
  • Data and sample: The sample includes 3,925 (75.6%) of 5,191 back-office employees using aIQ Chat at least once during any month of the sample window.The Copilot dataset covers 4,392 employees, of whom 3,962 (90.2%) use Copilot and 2,861 (65.1%) use both tools.
  • Data and sample: The study observes 158,496 conversations and 713,564 user prompts, averaging 4.5 user prompts per conversation.Usage is highly skewed: the mean active employee has 40 conversations, the median has 14, and the top decile accounts for 51% of conversations.

4 Measures

The study converts aIQ Chat transcripts into structured measures using an LLM metaprompt, covering conversation use cases, prompting characteristics, and indicators of sophisticated use. It aggregates these measures to active employee-months and summarizes group differences with dot plots and appendix statistics.

  • Transcript coding: An LLM applies structured coding rules to each aIQ Chat transcript to identify use cases, prompting strategies, and other characteristics of sophisticated use.The resulting LLM outputs serve as structured measurement data.
  • Use cases: Conversations are classified into eight use-case categories, including WRITING, CODING DATA, TEXT ANALYSIS, KNOWLEDGE, TOOL GUIDE, IDEATION, PERSONAL, and OTHER.Each broad category is further decomposed into subcategories, and workplace conversations may blend tasks.
  • Sophistication measures: Measures capture prompt complexity, conversation structure, prompting sophistication, specificity, format clarity, behavioral indicators, and prompting strategies.First-prompt measures include 1ST PROMPT LANG COMPLEXITY and 1ST PROMPT COMPLEXITY; conversation measures include STRUCTURE, PROMPT SOPH, SPECIFICITY, and FORMAT CLARITY.
  • Employee-month measures: Conversation-level measures are aggregated to active employee-months, defined as months in which an employee has at least one aIQ Chat conversation.Sophisticated genAI use is operationalized as skilled use reflected in clear and specific instructions, deliberate prompting techniques, and a broad mix of tasks.
  • Results presentation: Dot plots report group mean outcomes relative to the overall sample mean, with underlying tables and statistics provided in Online Appendix B.Rows represent months, seniority levels, or functional areas, and dot position shows how far a group lies above or below the sample average.

5 Extent of GenAI Use

GenAI use is widespread in the sample, with Copilot more prevalent than aIQ Chat. Usage varies by month, seniority, and functional area, but does not show rapid growth during January–August 2025.

  • Usage prevalence: 83.9% of employee-months include Any Tool use, compared with 78.0% for Copilot and 43.7% for aIQ Chat.Copilot use is substantially more common than aIQ Chat use at the employee-month level.
  • Usage over time: Copilot use rises from 71.0% in January to 82.8% in August, while aIQ Chat usage remains comparatively flat.Month-to-month changes are modest, with no rapid or accelerating growth during the sample window.
  • Variation across employees: Usage positively correlates with seniority: above-manager employees use genAI more overall, driven by greater Copilot use, while aIQ Chat use is flatter across groups.Across functional areas, Digital Innovation, Communications, Sales & Marketing, and Internal Audit are typically above the sample mean; Accounting & Finance is significantly below it for every reported measure.
  • Comparison with prior research: 83.9% of employee-months use at least one enterprise LLM channel, exceeding the 32.1% of employed respondents and 27.3% using genAI at work in Bick et al. (2026).The prior study surveyed U.S. workers in August and November 2024, shortly before this sample begins.

6 GenAI Use Cases

Workplace genAI use is dominated by writing, while other use cases vary modestly over time and substantially by employee seniority and functional area. These patterns align with employees’ responsibilities and workflow differences.

  • Use-case composition: 73.3% of conversations involve writing, including editing user-provided material (45.8%) and generating new text (34.8%).Knowledge/expertise queries account for 23.0% of conversations, text/document analysis for 10.7%, and software/tool guidance for 9.9%.
  • Use over time: Writing remains the baseline use case throughout the sample window, while coding/data analysis increases somewhat and knowledge retrieval and ideation decrease.The study does not observe a clear monotonic shift in use across months.
  • Variation by seniority: Use cases vary by seniority: staff skew toward writing and personal requests, managers make relatively more coding and data-analysis requests, and above-manager employees are most knowledge-intensive.The seniority pattern is consistent with genAI use reflecting employees’ responsibilities and complementing their domain knowledge.
  • Variation by function: Use-case patterns align with functional workflows, including more personal requests in Administrative Services and Internal Operations, coding and tool guidance in Information Technology, and ideation alongside writing in Communications and Sales & Marketing.These differences are consistent with the expected tasks of each functional area, such as travel and scheduling in administrative groups.

7 Sophistication of GenAI Use

Sophisticated genAI use varies across employees and functions, while deliberate prompting strategies remain uncommon. Higher sophistication is associated with longer initial prompts, more iteration, and current-month training, but not reliably with broader platform use.

  • Descriptive patterns: 5.2% of conversations use any deliberate prompting strategy, while role prompting occurs in 4.4% and each other listed strategy appears in less than 1%.The listed strategies are few-shot examples, explicit chain-of-thought requests, self-checking prompts, and interactive refinement.
  • Functional variation: Project Management is the only functional area significantly above the sample mean on all three sophistication measures.Communications and Sales & Marketing exceed the mean for prompt clarity and strategy use but fall below it for use-case diversity.
  • Functional variation: Digital Innovation and Strategy exceed the mean on prompt clarity and use-case diversity, whereas Administrative Services, Property & Facilities, and Internal Operations fall below the mean on all three measures.Accounting & Finance is also below the mean across all three sophisticated-use measures.
  • Observable correlates: User fixed effects eliminate broad evidence that Copilot use improves sophisticated use, suggesting that users with greater sophistication are more likely to adopt multiple platforms.Figure 4 and Table 4 also show no reliably positive associations between sophisticated use and frequent aIQ Chat use.
  • Observable correlates: Longer first prompts, more iteration, and current-month training remain positively associated with sophisticated-use measures across multivariate specifications.These associations persist in specifications with and without user fixed effects.

8 Discussion and Conclusion

The study documents substantial variation in sophisticated genAI use across seniority levels, functions, and time, while finding little evidence that formal training produces lasting behavioral change. These descriptive findings provide managerial signals and historical benchmarks, but do not establish output quality, productivity effects, or causality.

  • Seniority: Senior employees exhibit more sophisticated genAI use, potentially reflecting domain expertise, delegation proficiency, accountability, or broader professional experience.The evidence is consistent with domain expertise complementing genAI capabilities, but does not identify a single explanation.
  • Functional differences: Strategy, Digital Innovation, and Project Management exhibit the highest sophistication, while Accounting & Finance is below average across all measures.The three highest-sophistication groups share a focus on firmwide strategic initiatives and organizational change.
  • Time and training: Across the eight-month sample, sophistication does not clearly improve over time or show lasting increases after formal AI training.Sophistication rises in the month of training completion, but not in subsequent months.
  • Observable behaviors: Longer initial prompts and more conversational iteration are more consistently associated with sophisticated use than usage frequency.These observable behaviors may provide useful signals when transcript-based measures are unavailable, but none directly measures performance.
  • Limitations: The composite measures proxy for sophisticated use but cannot determine whether LLM outputs satisfied users’ goals, improved output quality, or affected productivity.Rapid model and AI-agent advances may also have changed prompting best practices since the study window ended.
  • Implications: Despite its descriptive design, the study offers useful AI-behavior benchmarks and motivates future research on genAI’s workplace entry and effects.The authors emphasize that the analysis does not provide causal inference.

Appendix A: Variable Definitions

Appendix A defines employee-month measures of sophisticated genAI use, use-case breadth, prompting strategies, conversation characteristics, training, and prompt style. It also specifies binary indicators for six core use-case categories and detailed prompt-quality, strategy, and style measures.

  • Sophisticated use measures: Sophisticated use is measured with CLARITYZ, ANY STRAT, and USE DIV, capturing prompt quality, deliberate prompting strategies, and breadth across use-case categories.CLARITYZ averages z-scored continuous and binary prompt-quality features; ANY STRAT is the share of conversations using at least one of five strategies; USE DIV is normalized Shannon entropy across eight categories.
  • Predictor variables: Employee-month predictors include first-prompt length, conversation turns, aIQ Chat and Copilot usage, and completed training, with several variables defined as natural logarithms of 1+ counts.Predictors also include cumulative training through the beginning of the month, while threshold indicators identify specified minimum levels of turns, usage days, and training.
  • Prompt input-output measures: Strategy indicators identify role play, few-shot examples, step-by-step reasoning requests, self-checking, and collaborative refinement when explicitly requested by the user.Each strategy is coded as 1 when present and 0 otherwise.

Online Appendix B: Supplementary Tables

Online Appendix B provides supplementary tables underlying the paper’s four main figures, covering usage, use cases, evaluation, and bivariate associations. The tables document their estimation samples, centering conventions, significance procedures, and model specifications.

  • Appendix purpose: Four supplementary tables correspond to Figures 1–4 and report usage and intensity, use cases by group, evaluation of use, and bivariate associations.Table B1 corresponds to Figure 1, Table B2 to Figure 2, Table B3 to Figure 3, and Table B4 to Figure 4.
  • Usage and intensity: Table B1 reports usage and intensity statistics computed at the employee-month level, with entries mean-centered relative to the sample mean.The sample mean is reported in the “Sample Mean” row.
  • Use cases by group: Table B2 reports conversation-level use-case proportions by group, allowing row proportions to exceed one because conversations may receive multiple labels.Panels B and C exclude conversations lacking seniority or functional-area data.
  • Evaluation of use: Table B3 reports group averages over active employee-months, defined as months with at least one aIQ Chat conversation.N denotes the number of active employee-months in each group; Panels B and C exclude observations missing seniority or functional-area data.
  • Bivariate associations: Table B4 reports bivariate associations using active employee-month observations with month fixed effects, with Panel B additionally including user fixed effects.Predictors are grouped by construct, and most predictors enter as log(1+x), excluding 1ST PROMPT LEN and threshold indicators.

Online Appendix C: Meta Prompts

Online Appendix C reproduces the metaprompts submitted with each aIQ Chat transcript and is intended for online publication only.

  • Online Appendix C: Meta Prompts: The appendix reproduces the metaprompts submitted with each aIQ Chat transcript.It is intended for online publication only.

Prompt One: Use Cases

The use-case taxonomy classifies full LLM conversations by the action requested, using user prompts and follow-ups as evidence and allowing multiple applicable categories. It spans writing, coding and data analysis, text-based analysis, knowledge and expertise, software guidance, creative thinking, non-work activities, and LLM-capability testing.

  • Classification principles: Conversations are classified by the requested action after reviewing the entire dialogue, with multiple categories assigned when applicable; decisions rely on user prompts and follow-ups.LLM responses are used only to interpret clarifications in subsequent follow-ups.
  • Writing and communication: Writing and communication covers generating new content, editing or improving existing user-provided content, language translation, and other writing tasks.Generation and editing can both apply when a request adds substantial new material to provided content.
  • Coding and data analysis: Coding and data analysis covers executable-code generation or debugging, analysis of provided data, and data cleaning, restructuring, or standardization.Data-analysis and data-cleaning categories require actual data plus analysis or manipulation, respectively.
  • Analysis, expertise, and software guidance: Text-based analysis, knowledge and expertise, and software guidance distinguish document or prose analysis, domain-specific questions, and instructions for using software functionality.Software guidance concerns how to use a tool, whereas content creation or task completion using an LLM is classified by the underlying task.
  • Creative and non-work uses: Additional categories capture brainstorming and ideation, personal or other non-work activities, and testing LLM capabilities.Brainstorming requires generating new ideas, while personal-task indicators include grocery lists, travel, and home projects.

Prompt Two: Strategies

The strategy-classification prompt analyzes entire KPMG employee–LLM conversations for explicitly requested prompting strategies, assigns each identified strategy an intensity score, and returns the results as JSON.

  • Prompting strategies: The prompt identifies five strategies: role-playing, few-shot, chain-of-thought, self-verification, and interactive refinement.Each strategy is defined by linguistic markers and distinctions, such as assigning an identity, providing complete examples, requesting reasoning, checking one’s own output, or seeking collaborative dialogue.
  • Classification rules: The classifier analyzes the entire conversation, permits multiple strategies, codes only explicit requests, and assigns “none” when no strategy applies.LLM responses provide context for follow-up prompts but are excluded from classification.
  • Intensity scoring: Each identified strategy receives an intensity score: Light (1), Moderate (2), or Heavy (3).The levels represent brief or minimal, clear but not elaborate, and extensive or sophisticated use, respectively.

Prompt Three: Other Factors

The analysis evaluates additional conversation-level factors using the entire user–LLM dialogue, emphasizing user instructional language rather than submitted work product or LLM response content. It combines qualitative 1–5 ratings with binary indicators covering prompt clarity, structure, constraints, feedback, and writing conventions.

  • Analysis principles: Ratings use the entire conversation and assess user prompts, while LLM responses serve only to interpret later clarifications.The method evaluates the overall conversational journey rather than only the initial prompt.
  • Qualitative ratings: Qualitative measures use a 1–5 scale for initial language complexity, initial complexity, prompt structure quality, output format clarity, prompting sophistication, and specificity.These measures distinguish linguistic form, task articulation, organization, requested output specifications, prompting skill, and clarity across the conversation.
  • Conversation features: Binary indicators capture document uploads, explicit constraints, structured-output requests, acceptance criteria, and explicit user satisfaction.Constraints must be actionable, structured output must specify formats or schemas, and satisfaction requires positive feedback rather than neutral acknowledgment.
  • Writing conventions: Additional behavioral indicators record frequent typos, all-lowercase writing, and pleasantries, with explicit exclusions for intentional abbreviations, provided content, and other non-user text.Frequent typos require 3+ across the conversation or 2+ in one message; all-lowercase requires every user message to qualify.
Loading 2608.27364v1…