Source-linked AI summary

AI Soccer Analyst: Stage-Aware and Verifiable Human-AI Collaboration for Soccer Data Analysis

Calvin Yeung, Keisuke Fujii

arXiv:2609.11224v1cs.HCcs.AI

TL;DR

Soccer analysts must translate domain questions into inspectable definitions, computations, and evidence, but prompt-to-report LLM workflows can obscure those decisions. AI Soccer Analyst organizes analysis into revisable stages that retain analyst input and connect claims to execution evidence. In evaluation, 33 of 48 tasks met operational completion criteria, with favorable completed-task perceptions and domain knowledge emerging through interaction.

  • Problem

    Soccer data analysis requires domain-specific definitions and evidence checks, while prompt-to-report LLM workflows can obscure assumptions, interpretations, and computational support.

  • Method

    AI Soccer Analyst uses inspectable, revisable stages and intermediate artifacts so analysts can shape definitions and plans, review outputs, and trace claims to execution evidence.

  • Results

    33 of 48 analytical tasks met the operational completion criteria, and completed-task ratings for output quality, task achievement, reliability, and verifiability remained above the neutral midpoint after Holm correction.

  • Takeaways & Limitations

    Stage-aware human–AI collaboration can support inspectable, revisable, and verifiable soccer analyses while retaining domain-expert involvement in consequential decisions.

  • Takeaways & Limitations

    The evidence comes from focused samples of five formative participants and 16 evaluation participants, with ratings covering only 33 completed tasks.

Abstract

from arXiv · show

Sports data analysts translate domain questions into insights by combining computation with sport-specific domain expertise. Large language models ease programming, but prompt-to-report workflows may obscure decisions and evidence. We present AI Soccer Analyst, a mixed-initiative system with revisable stages: Data Understanding, Problem Definition, Structured Planning, Execution, Evidence-Grounded Reporting, and Interaction and Refinement. A formative study with five analysts first informed design goals for automation, verifiability, human control, and accessibility. Subsequently, a task-based evaluation with 16 participants combined system logs, retained artifacts, ratings, and open responses; 33 of 48 tasks met the operational completion criteria. Exploratory tests supported favorable participant perceptions of completed-task output quality, task achievement, reliability, and verifiability after Holm correction. Interaction records showed domain knowledge emerging through clarification, planning, and refinement. These findings position stage-aware human-AI collaboration as a practical approach for producing inspectable, revisable, and verifiable analyses while retaining domain-expert involvement in consequential decisions.

1 Introduction

Soccer analysis requires translating domain questions into data-supported computations and judgments, while LLM prompt-to-report workflows can obscure assumptions and evidence. AI Soccer Analyst addresses this with inspectable, revisable stages and reports empirical evidence from analyst studies and task evaluation.

  • Motivation: Soccer analysts must determine what data supports, implement computations, inspect intermediate results, and communicate evidence because analysis combines computation with situated domain judgment.Sport settings connect analytical results to contextual and tactical decisions with practical consequences.
  • Motivation: LLM-generated analyses can sound convincing while hiding question interpretation, unavailable data assumptions, incomplete calculations, and the basis for tactical conclusions.Users of LLM-assisted data analysis may also struggle to provide sufficient context and verify generated results.
  • System contribution: AI Soccer Analyst externalizes analytical decisions and artifacts, requests domain input when needed, routes revisions to responsible stages, and connects reported claims to execution evidence.Its workflow includes Data Understanding, Problem Definition, Structured Planning, Execution, Evidence-Grounded Reporting, and Interaction and Refinement.
  • Evaluation findings: 33 of 48 analytical tasks met the operational completion criteria, while completed-task ratings for output quality, task achievement, reliability, and verifiability remained above the neutral midpoint after Holm correction.Domain knowledge emerged through clarification, planning, and refinement; incomplete tasks concentrated around understanding dataset support.
  • Study contributions: The formative study identified requirements for efficient automation, domain control, verifiability, accessibility, and practical communication that informed the system design.The evaluation then examined usefulness, operational incompletion, domain-knowledge entry points, and interaction and verification support.

2 Related Work

Related work frames soccer analysis as context-dependent collaboration requiring expert control, while LLM-assisted tools need inspectable intermediate steps and evidence connections. AI Soccer Analyst combines stage-aware intervention with provenance-oriented verification to address this gap.

  • Sports analytics in practice: Soccer concepts can have multiple valid operational definitions, so analysts must decide which interpretation and computation fit the question and available data.Different definitions of an attacking play can produce different measures and conclusions.
  • Mixed-initiative workflows: Mixed-initiative interaction lets users and AI initiate actions with control shifting as the task evolves, matching computational automation with analyst domain expertise.The appropriate balance between automation and human control may change within a single task.
  • Mixed-initiative workflows: Stage-aware intervention places confirmation at analytical-intent and planning checkpoints, while requesting input elsewhere only for unresolved decisions that could materially change the analysis.This approach aims to avoid both unchecked decisions and poorly timed interruptions.
  • Verifiability and provenance: Verification requires access to the analysis process because convincing answers and final explanations alone do not establish correctness.Inspecting assumptions, task steps, intermediate outputs, and evidence connections helps users judge and redirect the analysis.
  • Verifiability and provenance: Analytical verifiability records how human decisions relate to automated work so users can follow reported evidence and revise the process when the connection is wrong.This extends provenance from recording history to supporting correction of the analytical chain.
  • Research gap: Existing research links useful sports analysis to practical context and meaningful expert control, but these concerns have rarely been combined in one workflow.The paper examines whether stage-aware soccer analysis can preserve control and verification while retaining LLM-assisted efficiency.

3 Formative Study

The formative study used surveys and selective interviews with five soccer analysts to explore requirements for responsible AI assistance. Its themes shaped design goals around efficiency, verification, human control, communication, security, and accessibility.

  • 3.1 Method: The formative study examined current practices, workflow difficulties, and expectations about efficiency, verifiability, trust, and AI assistance.Closed-ended responses were descriptive, while qualitative analysis combined deductive and inductive coding.
  • 3.1 Method: Five of seven contacted soccer analysts completed the survey, and two of those five completed selective follow-up interviews.Participants worked in collegiate or professional contexts.
  • 3.1 Method: The analysis was exploratory, conducted by a single coder, and informed design goals rather than independently establishing them.The small sample limited closed-ended analysis to descriptive contextualization.
  • 3.2 Thematic Findings: Participants identified five themes: workflow efficiency, verifiability and trust, automation boundaries, practical communication, and secure and accessible use.These themes describe where AI assistance may fit and the conditions required for responsible use.
  • 3.2 Thematic Findings: Manual data preparation, report production, and verification created substantial effort, with slow collection constraining both available time and analytical scope.One participant reported spending more than a day and a half aggregating phases of play and event-location coordinates.
  • 3.2 Thematic Findings: Trust depended on process inspection, source-level checking, domain validation, interpretability, alignment with soccer knowledge, and communicated uncertainty.Participants distinguished meaningful checking from qualitative validation constrained by data limitations.
  • 3.2 Thematic Findings: Participants supported automation for routine work but expected consequential analytical decisions to remain under human control without eroding expertise or learning opportunities.Concerns included loss of analytical thinking and excessive verification burden.
  • 3.2 Thematic Findings: Useful outputs needed prioritization, fit with staff workflows, focused revision, security across the data lifecycle, and accessibility for nontechnical stakeholders.Participants cited report-design effort, excessive output volume, information leakage, and the needs of coaches and players.

4 Design Goals

The design goals translate formative findings into a workflow that automates supported analytical work while preserving inspection, revision, human control, accessibility, and constrained data handling.

  • Design goals: DG1 targets efficient, data-grounded automation across preparation, computation, verification, reporting, and revision within supported workflow stages.External data-access constraints may limit achievable benefits.
  • Design goals: DG2 preserves verifiability and human analytical control through cross-stage inspection, revision, traceable claims, and oversight of unresolved consequential decisions.The goal is appropriate reliance rather than automation without review.
  • Design goals: DG3 supports accessible and actionable use through clear reporting, accessible interaction, prioritization, and self-service use across stakeholders with different technical expertise.Outputs should be understandable and usable without continual specialist support.
  • Cross-cutting requirement: A cross-cutting requirement reduces unnecessary data exposure and limits the effects of untrusted operations across workflow stages.This requirement applies across the design goals rather than to one interaction objective.

5 System Overview

AI Soccer Analyst combines a web interface, coordinating backend, local LLM, isolated coder worker, and persistent artifacts to support inspectable soccer-data analysis. Local deployment and execution isolation support controlled data handling and bounded computation.

  • Architecture: The system coordinates a web interface, backend agents, local LLM, isolated coder worker, and persistent analysis artifacts.Artifacts preserve problem definitions, plans, execution records, outputs, review decisions, reports, and interaction history.
  • Web Interface: Each stage produces inspectable outputs that users can revisit through a persistent analysis workspace.The interface connects user requests and datasets with progress, artifacts, and interaction history.
  • Model Deployment: Local deployment keeps analytical data, prompts, and workflow context within the controlled environment across stages.The same model checkpoint and inference configuration are used consistently across workflow stages and evaluation conditions.
  • Execution Isolation: The isolated coder worker runs generated scripts in a separate container with limits on backend access, time, concurrency, and output size.This separation reduces exposure to unintended file operations, excessive resource use, and sensitive backend configuration.

6 Analytical Workflow

The analytical workflow moves through six revisable stages from dataset understanding to evidence-grounded reporting and refinement. Automation handles profiling, planning support, coding, review, and artifact routing while users confirm analytical intent and domain validity.

  • Workflow Structure: The six-stage workflow progresses from Data Understanding through Problem Definition, Structured Planning, Execution, Evidence-Grounded Reporting, and Interaction and Refinement.Execution Review is an internal substage, and confirmed refinements are routed to affected stages.
  • User Control: Users judge dataset suitability, confirm the analytical definition and plan, and remain responsible for domain validity.The system profiles data and requests clarification, but users decide whether the data and definitions support their intended analysis.
  • Execution and Review: The CoderAgent automatically generates and executes scripts, while the ReviewAgent independently checks plan adherence and computational consistency.Review acceptance advances reporting; revision feedback returns the workflow to execution within a configured round limit.
  • Evidence-Grounded Reporting: Evidence-Grounded Reporting links report claims to reviewed outputs and execution records through identifiers.The report uses the confirmed problem definition and cleaned code as an explanatory view of the computation.
  • Interaction and Refinement: Refinement requests are routed to the earliest responsible stage so affected downstream artifacts can be regenerated consistently.Before refinement, users confirm that the proposed target and scope match the intended change.

7 User Experiment

The user experiment evaluated AI Soccer Analyst across descriptive, comparative, and tactical soccer-analysis tasks using system records, artifacts, ratings, and open responses. The design separated increasing analytical and domain demands while defining completion as production of a final report.

  • Dataset: The evaluation used the Wyscout 2017 event dataset, whose recorded actions support filtering, aggregation, sequence analysis, and spatial summaries.Analyses remain limited to actions in the event log, excluding off-the-ball movement and formation changes.
  • Task Design: Participants attempted L1 descriptive retrieval, L2 comparative interpretation, and L3 tactical synthesis tasks with increasing analytical and soccer-domain demands.L1 required direct filtering and aggregation, whereas L2 and L3 required contextual comparison and integration of event patterns into tactical accounts.
  • Measures: Operational completion required the system to finish Evidence-Grounded Reporting and generate a final report.Tasks ending as failed, cancelled, or otherwise unfinished did not meet the completion criterion.
  • Measures: Participants rated completed tasks on output quality, task achievement, reliability, and verifiability, while overall assessment covered trust, validity judgment, process and evidence verification, efficiency, and interaction usefulness.Open responses addressed strengths, difficulties, untrustworthy outputs, and additional evidence needed for verification.
  • Participants and Procedure: 16 participants completed 48 analytical tasks, with 16 tasks at each difficulty level.Participants formulated their own soccer-analysis question at each level after standardized instructions and a workflow walkthrough.

8 Results

The evaluation found favorable perceptions of completed analyses, while incomplete tasks exposed data-resolution and execution constraints. Interaction records showed that participants contributed domain knowledge through planning, clarification, and refinement, motivating design implications for provider-specific data understanding, persistent collaboration, and active verification.

  • 8.2 Operationally Incomplete Task Analysis: 15 tasks were operationally incomplete, including 4 infeasible requests requiring unavailable tracking or off-ball data and 6 failures involving unresolved dataset fields, identifiers, tags, or records.The remaining incomplete tasks involved memory exhaustion, an incorrect report path, and a time-limit cancellation.
  • 8.3 Human–AI Collaboration Patterns: Participants introduced or revised explicit soccer-domain knowledge in 14 of 48 tasks, most often during Structured Planning, followed by post-result refinement and clarification.Planning converted participant expertise into executable definitions, while refinement applied soccer judgment after results became visible.
  • 8.4 Overall System Ratings: After Holm correction, reduced manual work and helpful interaction were the only significant overall-system ratings, while validity judgment and evidence verification showed less consistent support.Evidence verifiability was significant before correction but not afterward; validity judgment had mean 3.56 and median 4 with substantial variation.
  • 8.5 Open-Response Themes: Open responses emphasized interaction continuity and recovery, alongside analytical credibility and verification, with 11 participants mentioning each theme.These themes, together with data-resolution failures, motivated the subsequent design implications.
  • 8.6 Design Implications: The evaluation motivates provider-specific capability knowledge, persistent human–AI refinement, and active verification beyond tracing how outputs were produced.The proposed directions include feasibility checks before execution, reusable analyst-approved definitions, and domain-specific validation or independent recomputation.

9 Discussion

The discussion frames AI Soccer Analyst as a process-level approach that makes analytical decisions inspectable and revisable while integrating domain expertise throughout the workflow. Its benefits are qualified by weaker evidence for validity assessment and by the uncontrolled difficulty differences across task levels.

  • 9 Discussion: AI Soccer Analyst organizes analysis as an inspectable process in which human expertise can change the analysis at the stage where intervention is needed.Process artifacts support inspection and verification, while clarification, planning feedback, and refinement allow domain knowledge to enter as analysis develops.
  • 9.1 Revisability as Process-Level Transparency: Externalizing the problem definition, plan, execution record, and claim–evidence links exposes analytical decisions and routes corrections to the responsible stage.This shifts transparency from explaining a final answer to exposing how the answer was produced.
  • 9.1 Revisability as Process-Level Transparency: Revisability lets analysts change a questionable definition or plan and regenerate affected artifacts instead of accepting or rejecting the entire report.Process evidence indicates what to examine and where an intervention can be made.
  • 9.2 Human–AI Collaboration as Expertise Integration: The system performs repetitive computational work, while analysts define soccer concepts, select meaningful comparisons, and judge tactical credibility.This division keeps domain knowledge active beyond the initial request and aligns system automation with analyst judgment.
  • 9.2 Human–AI Collaboration as Expertise Integration: L3 tasks had lower completion and required more execution rounds, but differing task content and fixed order prevent treating this pattern as a controlled difficulty effect.The discussion therefore recommends stronger opportunities to revise definitions, plans, and interpretations for complex analysis.
  • 9.2 Human–AI Collaboration as Expertise Integration: Stage-aware collaboration makes analytical work revisable and open to domain judgment, extending its value beyond manual-work reduction.Analysts receive explicit opportunities to shape, inspect, and verify an analysis as it develops.

10 Limitations and Future Work

The evaluation provides focused evidence from limited participant and data settings, while LLM-generated analyses remain vulnerable to assumptions and shared blind spots. Future work should test broader transfer and stronger correctness checks.

  • 33 of 48 analytical tasks met the operational completion criteria, but ratings covered only completed tasks and omitted the 15 incomplete tasks.The task-based evaluation included 16 participants, while task-level ratings came from 15 participants with at least one completed task.
  • The formative and task-based studies included 5 and 16 participants, respectively, limiting coverage of soccer-analysis roles and organizational settings.The authors call for larger and more diverse samples including professional and academy analysts, coaches, and technical staff.
  • The implementation and evaluation used structured JSON event data with a Wyscout 2017 configuration, constraining evidence about transfer across providers, schemas, competitions, data types, and sports.Event logs support only concepts directly represented or derivable from recorded fields, and sample-based profiling may miss rare structures.
  • Claim–evidence links improve inspectability, but traceability alone does not establish computational or soccer-analytical correctness.The authors propose combining model-based review with deterministic checks, independent models, or expert assessment.

11 Conclusion

AI Soccer Analyst organizes soccer event-data analysis into inspectable, revisable stages rather than a one-step prompt-to-report workflow. The conclusion emphasizes traceable evidence and continued domain-expert involvement in consequential analytical decisions.

  • AI Soccer Analyst organizes LLM-assisted soccer analysis into inspectable stages from Data Understanding through Interaction and Refinement.
  • The workflow lets analysts revise definitions and plans before assumptions propagate, then trace report claims back to execution evidence.Revisions target the responsible analytical stage so affected outputs can be updated consistently.
  • The system combines computational automation with domain-expert opportunities to shape and verify analyses.The design highlights control before execution and inspection of claim–evidence connections afterward.
  • The evaluation materials measured task-level output quality, task achievement, reliability, and verifiability after L1, L2, and L3 tasks.

B Supplementary Formative Study Results

The supplementary results describe a small formative cohort, distributed workflow difficulties, strong expectations for transparent AI assistance, and the task-evaluation cohort and analysis boundaries.

  • B.1 Study Data and Participants: The formative study surveyed 5 participants and selectively interviewed 2 of them using semi-structured follow-up questions.Participants had sustained soccer-analysis experience, with 4 of 5 reporting at least 4 years.
  • B.2 Participant and Workflow Context: Participants combined manual, automated, and AI-assisted workflow modes, while human effort remained substantial in interpretation, checking, and communication.
  • B.2 Participant and Workflow Context: Workflow difficulties spanned data preparation, insight discovery, validation, communication, and time cost rather than concentrating in one predefined bottleneck.Report preparation received no selections, while follow-up responses supplied additional context.
  • B.2 Participant and Workflow Context: Time cost received the highest median current-workflow rating, while verification burden and process opacity varied across participants.
  • B.2 Participant and Workflow Context: Every expectation rating was 4 or 5, with strongest consensus around transparency, verifiability, reduced manual work, and greater speed.
  • B.3 Qualitative Coding and Codebook: Qualitative responses were coded with deductive categories covering trust and calibrated reliance, verifiability and evidence, and efficiency and workload.
  • C.1 Study Data and Participants: The task-based evaluation retained 16 survey responses and 48 linked analytical tasks, with one L1, L2, and L3 task per participant.
  • C.2 Cohort and Analysis Boundaries: Task-rating analyses covered 33 completed tasks—12 L1, 11 L2, and 10 L3—while overall-system ratings included all 16 participants.

C.3 Question Types, Operationally Incomplete Tasks, and System Behavior

The evaluation characterizes task demands, operational incompletion, system behavior, and participant ratings across analytical tasks. It also reports how open responses and prompt designs informed interpretation of the system’s performance and refinement needs.

  • Operationally Incomplete Tasks: 4 tasks were infeasible because they required unavailable tracking or off-ball information, while 6 failed because required dataset fields or records could not be resolved.
  • Completed-Task Ratings: All four participant-aggregated completed-task outcomes remained significant after Holm correction, with adjusted p-values from .0002 to .0044.
  • Overall System Ratings: Only reduced manual work and helpful interaction remained significant among overall-system ratings after Holm correction, with adjusted p-values of .0049 and .0184.
  • Open Responses: Open-response themes supported favorable assessments of efficiency, accessibility, and analytical support while motivating persistent refinement and active verification.
  • Prompt Templates: The appendix documents prompts for five LLM-mediated stages and two Execution substages, while deterministic profiling handles Data Understanding.

G.4 Generated Report

The generated report aggregates canonical soccer actions by team from Wyscout event files and exports the results as a CSV. Its evidence links expose the computational basis of claims, while the cleaned code remains explanatory rather than rerun provenance.

  • The complete per-team summary contains 142 teams, with FC Barcelona recording 23,260 passes and Real Madrid recording 22,081 passes and 631 shots.
  • 1,665,508 passes, 43,078 shots, and 879,083 duels were counted across the dataset.
  • The workflow discovered event labels, mapped them to Pass, Shot, and Duel, aggregated counts by team, joined team names, and exported a CSV.
  • The evidence view links report claims to cleaned-code sections that explain analytical logic but do not independently validate analytical interpretation.
  • The cleaned code was not rerun, so it serves as an explanatory view rather than the executable provenance record.
Loading 2609.11224v1…