Source-linked AI summary
"ChatGPT, what am I missing?": Designing AI Workflows around Professional Task Structure to Shape Analytic AI Use
Zilin Ma, Suzi Jazmati, Marco Chimenton, Yiyang Mei, Jacqueline Lane, Krzysztof Z. Gajos, Finale Doshi-Velez
TL;DR
General-purpose AI lets users request support but leaves them to structure the professional task, motivating workflows that embed task structure without prescribing engagement. The paper compares completed and user-directed development of the same negotiation scaffold in a four-condition randomized study, finding that scaffolded workflows improve coverage over chat while user-directed development broadens analytic requests and lowers subjective effort.
Problem
General-purpose AI leaves users to construct the workflow needed for professional work, creating a need to preserve flexibility while providing professional scaffolding.
Method
The study compared AI-Prefilled and Co-Evolving interfaces built around the same negotiation scaffold with no-AI Reader and open-ended Chatbot conditions.
Results
Scaffolded workflows produced stronger preparation coverage than Chatbot, while Co-Evolving elicited broader analytic requests and required less subjective effort than AI-Prefilled.
Takeaways & Limitations
Professional AI should structure not only what it produces but also how users direct, inspect, and develop analysis together with it.
Takeaways & Limitations
The two prior workflow comparisons changed the scaffold and the user process simultaneously, limiting what could be inferred about their separate effects.
Abstract
from arXiv · showhide
General-purpose AI lets users choose what support to request, but leaves them to structure the support a professional task requires. We examine how interactive workflows can embed professional task structure without prescribing how users engage with AI. We designed two scaffolded interfaces around the same negotiation scaffold: one presented a completed AI analysis, while the other supported user-directed, incremental development. A four-condition randomized experiment with 800 participants compared these interfaces with no-AI and an AI chat interface. AI-supported conditions improved preparation coverage over unaided work; the scaffolded workflows further improved coverage over chat. Although the scaffolded workflows produced similar coverage, the user-directed workflow elicited a broader repertoire of analytic requests and lower subjective effort. Professional scaffolding therefore depends not only on displayed structure but on how workflows organize users' engagement with it. Effective professional AI must structure how users and AI build analysis together.
1 Introduction
The paper examines how AI workflows can embed professional task structure while preserving users’ ability to direct their engagement. In a randomized study of humanitarian negotiation preparation, scaffolded workflows improved coverage over chat, while user-directed development shaped AI use and effort differently from completed analyses.
- Study and findings: The study used a negotiation scaffold covering agreements, interests, red lines, and bottom lines, with Reader, Chatbot, AI-Prefilled, and Co-Evolving conditions.The four conditions supported comparisons between AI-supported and unaided work, scaffolded workflows and chat, and incremental versus completed scaffold development.
- Study and findings: Scaffolded AI workflows produced stronger preparation coverage than open-ended Chatbot, while AI-Prefilled and Co-Evolving did not differ detectably in overall coverage.The study compared both scaffolded workflows with no-AI Reader and open-ended Chatbot baselines.
- Motivation: General-purpose chat leaves users responsible for recognizing needed support, decomposing the task, formulating requests, and evaluating AI responses.Users may struggle to determine which support to seek and may lack the metaknowledge needed to allocate work productively.
- Study and findings: Co-Evolving elicited a broader analytic repertoire and required less subjective effort than AI-Prefilled despite similar overall preparation coverage.AI-Prefilled presented completed analysis; Co-Evolving let users choose what to pursue while AI developed analysis incrementally.
- Implications: The findings frame professional AI as work design: systems should organize not only outputs but also how users direct, inspect, and develop analysis.The authors argue that effective scaffolding must preserve users’ ability to decide which issues require attention and how AI should support them.
2 Related Work
Prior work shows that open-ended AI shifts workflow construction and cognitive allocation to users, while adaptive and scaffolded systems can move some structure back into the interface. The paper highlights an unresolved comparison: how the same professional scaffold changes work when presented as completed output versus a user-directed workflow.
- 2.1 Configuring AI Decision support: Open-ended generative AI lets users decide what to delegate and how AI contributes, but leaves each user responsible for constructing the supporting workflow.Users must recognize needed support, assess capabilities, formulate executable requests, and evaluate responses.
- 2.1 Configuring AI Decision support: Adaptive, generative, and just-in-time interfaces reduce fixed-tool selection but do not necessarily expose the larger professional task surrounding a user-stated or inferred objective.These approaches still depend on an objective expressed by the user or inferred by the system.
- 2.2 Professional scaffolds as output representations and workflow designs: The same scaffold can organize AI’s completed output or organize how users and AI produce analysis together through selectable interaction targets.The latter approach lets users choose what to examine while AI develops the corresponding analysis.
- 2.2 Professional scaffolds as output representations and workflow designs: Existing studies compared scaffolded with open-ended interfaces but did not compare different uses of the same scaffold, leaving workflow effects unresolved.Prior systems changed the scaffold and the process users followed at the same time.
- 2.2 Professional scaffolds as output representations and workflow designs: Professional scaffolds make recurring components of competent work visible, but representations can create blind spots by emphasizing named considerations and omitting others.Visible structure may improve represented task components without improving omitted ones, and does not ensure users verify AI-generated content.
- 2.2 Professional scaffolds as output representations and workflow designs: AI may redistribute effort from producing answers toward requesting, evaluating, and integrating them, while completed and incremental analyses impose different processing demands.Completed analyses can create reading and integration work; incremental development requires continued interaction but may support smaller-part processing.
3 Method
The study randomized participants to no-AI, chat-based AI, and two scaffolded AI workflows built around a professional negotiation-preparation structure. It evaluated these interfaces in a timed humanitarian negotiation task using shared case materials and preparation assessments.
- Study design: The four-condition experiment compared Reader, Chatbot, AI-Prefilled, and Co-Evolving interfaces on frontline humanitarian negotiation preparation.Reader was the no-AI baseline; Chatbot added a conversational assistant; AI-Prefilled showed a completed scaffolded analysis; Co-Evolving progressively updated the scaffold from participant–AI conversation.
- Interfaces: Both scaffolded workflows used adapted Iceberg, Island of Agreements, and Paths to Agreement frameworks from the CCHN Field Manual.The frameworks were shortened for the study, with AI-Prefilled displaying a completed analysis and Co-Evolving updating the scaffold progressively.
- Measures: A 15-minute preparation phase was followed by a 17-minute assessment covering participants’ positions, the opposing side, overlap and conflict, and scenario planning.Participants then answered ten free-text questions and reported preparation effort and confidence using seven-point scales.
- Preparation task: Participants prepared for a humanitarian negotiation over food assistance using 15 case documents containing conflicting or incomplete information.The 6,977-word corpus required integrating sources, distinguishing documented constraints from assertions, and addressing access, security, distribution, neutrality, and operational independence.
- Assessment procedure: Participants retained notes, chat history, or scaffolded analysis from their assigned interface while completing the assessment.The original case files were unavailable during writing, so retained preparation materials supported the assessment phase.
4 Results
AI support improved preparation coverage over unaided work, and making the professional scaffold visible added coverage over chat. The two scaffolded workflows achieved similar overall coverage, while Co-Evolving required less reported preparation effort.
- Overall preparation coverage: AI support increased average preparation coverage by 3.65 points relative to Reader (95% CI [1.25, 6.05], p= .003).Making the scaffold visible added 2.36 points over Chatbot (95% CI [0.11, 4.61], p= .040, g= 0.16).
- Overall preparation coverage: AI-Prefilled and Co-Evolving did not differ detectably in overall coverage.The estimated Co-Evolving advantage was 3.19 points, with 95% CI [−0.14, 6.52] and p= .060.
- Subjective experience: Co-Evolving participants reported less preparation effort than AI-Prefilled participants: 6.17 versus 6.55 (g= −0.35, pHolm = .005).Other contrasts for preparation effort, answer-writing effort, and frustration were not detectable.
- Subjective experience: AI conditions showed lower psychological ownership than Reader, while decision-agency means remained between 5.04 and 5.12 in every condition.Psychological ownership was also lower in scaffolded conditions than Chatbot; Co-Evolving and AI-Prefilled did not differ detectably.
- Question-level coverage: Scaffolded conditions exceeded Chatbot on the Commander’s redlines and main conflicts, but package-risk coverage showed no detectable change.The corresponding gains were +5.75 points (q= .0066) and +7.91 points (q< .001); package-risk coverage had q= .95.
Appendix Tables 15 and 17).
Scaffolding changed how participants engaged with AI: it shifted use toward analytic work, and Co-Evolving elicited a broader analytic repertoire than AI-Prefilled. Broader analytic engagement was associated with stronger preparation.
- AI-use purposes: Compared with Chatbot, scaffolded workflows reduced reading-support use by 30.1 percentage points and increased counterpart analysis, cross-party synthesis, and strategy or package development.The corresponding increases were 16.6, 13.0, and 11.8 points, respectively.
- Analytic breadth: Relative to AI-Prefilled, Co-Evolving increased each analytic purpose’s incidence by 30.9 to 40.2 percentage points.Co-Evolving also increased coverage of at least two purposes by 52.3 points and all four purposes by 23.5 points.
- Analytic breadth: Co-Evolving increased the expected number of analytic purposes in a three-turn sample to 1.68, versus 0.91 in Chatbot.The difference remained when interaction volume was held constant among participants with at least three turns.
- Request form: Scaffolding shifted turns from passage submission toward participant-authored requests, while explicit verification remained uncommon.Participant-authored requests comprised 25.2 more percentage points of turns and passage-only turns 20.2 fewer points than in Chatbot; only 2.4% of turns requested checking, defense, correction, or support.
2. Change the workflow
Analytic engagement was associated with stronger preparation, while prior experience tracked preparation quality in Reader but not in AI-supported conditions. The exploratory analyses characterize workflow-specific relationships rather than establish causation.
- Analytic engagement: Each additional analytic purpose corresponded to 2.66 more preparation-coverage points, whereas a ten-percentage-point increase in reading-support share corresponded to 0.56 fewer points.The model controlled for condition and log participant-turn count among 481 assistant users.
- Analytic engagement: Own-side analysis, counterpart analysis, cross-party synthesis, and strategy or package development each had positive associations with preparation coverage.The reported associations were +4.77, +4.15, +5.74, and +4.27 points, respectively.
- Analytic engagement: In a joint model, the four analytic-purpose coefficients averaged +2.65 coverage points, with no detectable differences among them.The average had 95% CI [1.59, 3.72] and p< .001; the coefficient comparison had p= .814.
- Prior experience: In Reader, each step on the four-level prior-experience scale corresponded to 2.62 additional coverage points, while slopes in AI conditions were negative and not detectable.The Reader slope exceeded the participant-weighted mean AI slope by 4.33 points (95% CI [1.66, 7.00], p= .002).
- Prior experience: Prior experience did not predict a broader analytic repertoire in open chat among Chatbot participants who used the assistant.Each experience step was associated with −0.06 analytic purposes (95% CI [−0.23, 0.11], p= .49).
A. Quality slope per experience level B. Analytic breadth among Chatbot users
Figure 5 relates prior negotiation experience to preparation quality across conditions and to analytic breadth among Chatbot users.
- A. Quality slope per experience level B. Analytic breadth among Chatbot users: Panel A shows condition-specific changes in preparation-quality coverage for each step on a four-level prior-experience scale.Panel B estimates four-purpose analytic breadth among Chatbot participants who interacted with the assistant; intervals use exploratory OLS with HC3 standard errors.
5 Discussion
AI improved negotiation-preparation coverage, and scaffolded workflows outperformed open-ended chat by embedding task-relevant objectives and organizing analytic engagement. However, these benefits did not uniformly reduce effort or preserve ownership, underscoring the need to design workflows that allocate authorship, verification, and decision responsibilities explicitly.
- Empirical findings: AI-supported interfaces improved negotiation-preparation coverage over unaided work, while pooled scaffolded interfaces improved coverage further than open-ended chat.The four-condition experiment compared AI-Prefilled, Co-Evolving, Chatbot, and no-AI work.
- Expertise and AI use: Negotiation experience predicted preparation quality without AI but not under any AI interface, and experienced participants asked similar kinds of chat questions as novices.A participant with no negotiation experience identified about as much of the case as a more experienced participant under AI support.
- Scaffolding versus chat: Co-Evolving surfaced framework-derived preparation objectives, elicited broader analytic requests, and shifted interaction from document processing toward analytic questions linked to stronger preparation.Chatbot users predominantly requested summaries or explanations, leaving much of the system’s analytic potential unused.
- Scope of scaffold benefits: Scaffolding improved several factual and package dimensions but not package-risk coverage, while participants rarely asked AI to justify or verify its suggestions.These findings indicate that making analytic structure visible does not ensure that all important analytic work is performed or checked.
- Design implications: Effective workplace AI requires workflow design that specifies how human judgments, AI support, authorship, verification, decision rights, and accountability are distributed.The authors recommend revisable scaffolds and lightweight forms of authorship rather than assuming expertise can compensate for an underspecified workflow.
- Effort and workflow pacing: AI redistributed rather than simply removed cognitive work: AI-Prefilled was most demanding, whereas incremental interaction was experienced as less demanding than receiving a complete analysis.Progressive disclosure may let users control the pace of analysis and process it incrementally.
- Agency and ownership: AI preserved decision agency but reduced psychological ownership, especially in Co-Evolving, because users chose the direction while AI generated much of the accumulated content.The authors therefore distinguish choosing what AI does from authoring what the system ultimately produces.
6 Conclusion
General-purpose AI can improve professional preparation, but effective support depends on making task structure visible and organizing how users and AI build an analysis together. In negotiation preparation, user-directed incremental development produced broader analytic engagement and lower effort than receiving a completed analysis, despite similar overall coverage.
- AI support improved case coverage, while workflows that made professional task structure visible added a further coverage gain over open-ended chat.The conclusion identifies this as the main preparation-quality result.
- Completed and incrementally developed versions of the same scaffold produced similar overall coverage but different forms of engagement.The scaffold influenced work through how participants encountered and developed the analysis, not only through its represented structure.
- The co-evolving workflow elicited a broader analytic repertoire and required less subjective effort than the prefilled workflow.This contrast held despite the two workflows producing similar overall coverage.
- Professional AI should not require users to reconstruct domain workflows through prompting or assume that a comprehensive answer alone provides effective support.Designers should instead make professional task structure visible through artifacts while allocating what users and AI direct, author, verify, and retain.
- The central design object is the workflow through which people and AI build analysis together, enabling preparation quality, analytic engagement, manageable effort, and a meaningful relationship to the resulting work as distinct goals.The conclusion frames workflow design—not only prompts or generated outputs—as the basis for pursuing these goals.
B.1 Common Task Shell
All conditions shared the same browser-based reading, timing, note-taking, and case workflow, while AI conditions added conversational or scaffolded support. The scaffolded interfaces differed in whether the preparation artifact began completed or was developed incrementally through conversation.
- Shared task shell: All four interfaces used the same browser application, case corpus, document order, preparation timer, and note-taking area.Participants selected numbered documents, read them in a central pane, and retained notes for the written questions.
- Shared task shell: Participants completed a non-skippable tutorial with neutral practice materials before the real case and 15-minute preparation task began.The practice content was cleared before the experimental case loaded, and the timer started only after the tutorial process ended.
- AI conditions: The Chatbot condition added a conversational assistant that accepted typed questions or selected passages from the reader.Its document-grounded assistant appeared below the reading area and preserved chat transcripts for later use.
- AI conditions: AI-Prefilled presented a completed, cited preparation scaffold covering sides and needs, agreement and conflict, and deal options.The artifact was generated from the case corpus before participants began, with citations linking claims to source passages.
- AI conditions: Co-Evolving began with an empty scaffold that was populated from conversation, while later proposed changes appeared as revisions participants could apply or reject.The workflow also marked an area not yet developed and attached supporting case passages to generated content.
C Written Preparation Questions
The timed preparation task asked participants to analyze both parties, synthesize overlap and conflict, and develop a realistic negotiation package. Its evaluation measured coverage across nine questions, while additional coding captured how participants used AI analytically.
- Written preparation task: Participants answered nine scored questions across four timed sections covering their side, the opposing side, overlap and conflict, and scenario planning.A tenth question elicited clarification questions but was not quantitatively scored.
- Written preparation task: The questions required identifying priorities, objectives, redlines, common ground, conflicts, proposals, package risks, and clarifications.Responses could be concise bullet points or incomplete sentences without penalty.
- Measurement: The professional scaffold represented the first eight preparation dimensions but did not explicitly represent package risks.This distinction matters because risks were still assessed in the written preparation measure.
- Measurement: Overall preparation coverage was the unweighted mean of nine question-level scores, preventing questions with more criteria from receiving greater weight.Questions 1–7 and 9 used atomic case-grounded criteria, while Question 8 used package-issue dimensions and direct option coding.
- AI-use analysis: Participant AI use was coded by request form and purpose, including reading support, party analysis, cross-party synthesis, and strategy or package development.Purpose coding was multilabel except for the exclusive other category, and coders did not infer purpose from assistant responses.
H.1 Proposed-Package Diversity
AI support did not produce a detectable change in the semantic diversity of proposed packages. Across conditions, package dispersion estimates varied numerically, but none of the tested contrasts was detectable.
- Proposed-package diversity: AI-supported preparation did not produce a detectable change in the semantic diversity of proposed packages.The exploratory analysis included 796 non-empty responses from the 800 participants.
- Proposed-package diversity: Mean leave-one-out cosine distance was 0.340 in Reader, 0.313 in Chatbot, 0.331 in AI-Prefilled, and 0.306 in Co-Evolving.Lower leave-one-out cosine distance indicates greater within-condition concentration.
- Proposed-package diversity: Although three contrast estimates indicated greater concentration, none differed detectably from its condition-label randomization distribution.The contrasts were evaluated with exact randomization tests using 49,999 condition-label randomizations.