Source-linked AI summary
MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production
Shih-Yu Lai, Wen-Fan Wang, Sai Ling, Shaune Jan, Bing-Yu Chen, Xiang Anthony Chen
TL;DR
Animation and VFX review must translate loosely specified intent into executable revisions, but rationale and evidence often disappear across senior–junior handoffs. The paper proposes and instantiates an intent–evidence–action workflow in MOONWALK, where AI coordinates evidence and tasks while practitioners retain aesthetic authority. In an in-studio comparison with chat-only review, MOONWALK produced stronger reported alignment, traceability, and checklist executability, though its evaluation did not isolate individual system mechanisms.
Problem
Animation and VFX review lacks a persistent shared representation connecting creative intent, evidential grounding, and executable action across senior–junior handoffs.
Method
MOONWALK instantiates an intent–evidence–action framework through shared records, evidence anchoring, structured comparison, and supervisor-authorized revision planning, with AI limited to coordination.
Results
MOONWALK was rated above neutral on 12 of 13 Likert items and led eight of nine comparative questions, including checklist executability, clarification reduction, and evidence-linked feedback.
Takeaways & Limitations
The findings support structured, evidence-linked review as an alternative to unstructured conversational AI while retaining aesthetic judgment and final prioritization with practitioners.
Takeaways & Limitations
The single-session exploratory evaluation compared MOONWALK with chat rather than non-AI structured baselines or isolated system components, and did not provide longitudinal production logs.
Abstract
from arXiv · showhide
Animation and VFX pre-production review requires teams to translate loosely specified creative intent--briefs, evolving specifications, heterogeneous references, and verbal decisions--into revisions that junior artists can execute without repeated clarification. In practice, criteria drift across iterations, review judgments lose their evidential basis, and the reasoning behind a request rarely survives the senior-junior handoff. We contribute a design framework for intent-evidence-action alignment: intent is articulated into a shared project record, judgments are anchored to grounded evidence, and authorized decisions are converted into clear revision tasks tied directly to reference notes. We instantiate this framework in MOONWALK, a professional pre-production review system comprising a shared intent record, reference/specification anchoring, structured work-in-progress comparison, and supervisor-authorized action planning. In this workflow, AI handles administrative coordination--flagging missing context and organizing notes--while artists retain full creative direction. An in-studio study with professional practitioners compares MOONWALK with a chat-only (chatbot) interface using matched production materials, while participants' existing workflows provide a retrospective ecological baseline. Results indicate stronger intent alignment, decision traceability, and checklist executability, while also showing that aesthetic authority and final prioritization must remain with practitioners. The evaluation establishes the value of the integrated structured workflow over unstructured conversational AI chatbot. Code: https://github.com/Akinesia112/Moonwalk/tree/english-version
1 INTRODUCTION
MOONWALK addresses the fragile intent–evidence–action chain in animation and VFX review by preserving creative intent, grounding judgments in inspectable evidence, and converting authorized decisions into executable revision actions. An in-studio evaluation with 19 practitioners found stronger alignment and actionability than chat-only review, while retaining aesthetic authority with practitioners.
- Problem: Pre-production review breaks down when creative intent, supporting evidence, and revision rationale are lost across senior–junior handoffs.Vague feedback can lead juniors to revise the wrong visual property or search for irrelevant references.
- Framework: The intent–evidence–action framework links active goals and constraints to supporting evidence and authorized revisions with completion conditions.These links persist across iterations so revisions can be traced to their rationale.
- System: MOONWALK operationalizes the framework through shared specifications, annotated references, structured comparisons, and evidence-linked action items.Its AI requests missing context, retrieves relevant evidence, checks stated requirements, and organizes authorized revisions without introducing aesthetic criteria.
- Human agency: AI supports coordination while artists retain creative direction and aesthetic judgment.The system does not independently judge lighting, color, composition, or style beyond the project record and explicit human judgment.
- Evaluation: MOONWALK was rated above neutral on 12 of 13 Likert items and led eight of nine comparative questions.Reported comparative outcomes included junior-executable checklists at 74%, reduced senior–junior clarification at 63%, and blind-spot identification and evidence-linked feedback at 79% each.
2 RELATED WORK
Related work emphasizes persistent criteria, inspectable traces, visual grounding, and human control in collaborative creative systems. MOONWALK applies these principles to the asymmetric senior–junior production handoff while complementing, rather than replacing, existing tracking platforms.
- Persistent collaboration: Collaborative creativity systems externalize goals, interpretations, and decision histories so collaborators can revisit them across iterations.Temporal support connects prior decisions with current activity instead of summarizing isolated sessions.
- Collaborative breakdowns: Prior HCI systems support shared criteria, visible meeting structure, continuity, disagreement, and inspectable feedback quality.These systems address adjacent collaborative breakdowns involving common ground, premature consensus, and accountability.
- Human control: AI assistance in creative work requires steering, inspection, rollback, and explicit human authorization because more assistance is not automatically better.MOONWALK positions AI as a coordination aid that prompts, retrieves, compares, and drafts.
- Production systems: Previsualization and production-tracking systems structure filmmaking work by supporting shared shot planning, capture, retrieval, versions, notes, review, and approval.MOONWALK targets the review rationale connecting a specific reference to a revision request rather than replacing these systems.
- Model boundaries: Multimodal models are more reliable for observational reference matching than for interpretive judgments requiring tacit domain standards.Human-in-the-loop systems therefore emphasize user-defined criteria, agreement inspection, and active auditing.
3 FORMATIVE STUDY
The formative study found that intent, evidence, criteria, and revision obligations often failed to survive animation and VFX handoffs. These findings motivated design goals for a persistent shared record, evidence traceability, actionable tasks, and visibility into both roles’ interpretations.
- 3 FORMATIVE STUDY: The formative study involved 12 practitioners across two studios and multiple animation and VFX production contexts.Participants included directors, a supervisor, senior and junior artists, and a project manager; sessions lasted 30–60 minutes.
- 3.1.1 Intent Did Not Survive the Handoff as a Shared Interpretation: Existing review relied on in-person discussions, tracking notes, chat, shared documents, calls, and annotated images, leaving rich visual context poorly recorded.Written feedback often omitted the reference region shown during a conversation, while disconnected channels obscured which instruction superseded another.
- 3.1.2 Review Judgments Were Delivered as Verdicts Rather Than Grounded Evidence: Review judgments were often delivered as verdicts such as “feels off” without identifying the motivating reference, specification, prior decision, or visual property.Junior artists cycled through plausible interpretations without a narrowing signal.
- 3.1.3 Criteria and Revision Obligations Drifted Across Iterations: Criteria drifted across iterations because persistent rationale did not connect successive versions, leaving absent references, undetected gaps, and undifferentiated comment lists.Participants reported that legitimate client revisions could become indistinguishable from contradictions.
- 3.2 Design Goals: The study identified a missing shared representation linking creative intent, review evidence, and executable action across both handoff directions.This breakdown directly motivated the three design goals.
- 3.2 Design Goals: The design goals require persistent intent articulation, evidence anchoring and traceability, and reference-grounded revision tasks.Tasks should state what to change, why it follows from evidence, priority, and the resolution condition.
- 3.2 Design Goals: Both roles need visibility into interpretations, rationale, uncertainty, and requests for additional evidence before action is finalized.MOONWALK operationalizes this through unified specifications, annotated reference hubs, interpretation records, and structured feedback checklists.
4 SYSTEM DESIGN & IMPLEMENTATION
MOONWALK carries intent, evidence, and authorized actions through one connected review record. Its implementation combines reference-grounded analysis with supervisor-controlled checklist creation, while acknowledging that some prompts could invite unanchored suggestions.
- Integrated workflow: MOONWALK connects the Spec/Brief, Reference Hub, comparison workspace, Review Canvas, and checklist as views of one review history.The workflow spans articulating intent, grounding evidence, and authorizing action.
- Articulate Intent: Supervisors annotate reference targets and requirements, while artists record interpretations, unresolved questions, and intentional deviations before submission.These notes remain attached to the WIP for subsequent review.
- Ground Evidence: The comparison workspace surfaces candidate differences against specifications, references, prior decisions, and artist interpretations for supervisor review.Region-level annotations identify the evidence motivating a comment.
- Authorize Action: Revision Consolidation turns accepted observations into prioritized checklist items retaining their supporting evidence, urgency, and completion condition.Supervisors can rewrite or merge items before returning the checklist to the junior artist.
- Technical implementation: Eleven analysis dimensions feed a three-model synthesis, with non-visual, low-score, and high-disagreement results escalated to human review.The models divide visual observation, specification/reference comparison, and synthesis responsibilities.
- Technical implementation: Broad evaluator prompts such as “what needs improvement” could invite suggestions beyond explicit project anchors, so the study does not establish that every observation was evidence-bounded.The supported claim is limited to a human-authorized structured review workflow.
5 SUMMATIVE STUDY
The summative study used a within-subject comparison of MOONWALK and Chat-Only with professional animation and VFX practitioners. It measured alignment, executable review output, and review awareness, traceability, and agency using questionnaires, comparative preferences, and interviews.
- Research questions: The study examined whether MOONWALK improves cross-role alignment, executable review output, and review awareness, traceability, and agency.These correspond to RQ1, RQ2, and RQ3.
- Participants: Nineteen practitioners participated across two studio contexts, including 10 senior-role and 9 junior-role participants.Senior participants had 4–16 years of experience, while junior participants had 0.5–3 years.
- Study design: Participants experienced MOONWALK and Chat-Only in two within-subject conditions using the same underlying AI and matched review materials.MOONWALK included structured reference and specification linkage, while Chat-Only exposed the conversational front end.
- Study design: Existing Workflow was collected after the task as a retrospective ecological comparison rather than a third controlled condition.The comparison captured participants’ studio practice without treating it as experimentally equivalent to the two interfaces.
- Measures and analysis: The evaluation combined a 13-item seven-point Likert questionnaire, nine three-way preference questions, and thematic interview analysis.The questionnaire covered collaboration, outcome/efficiency, and self-reflection.
- Measures and analysis: Likert ratings were tested against the neutral midpoint and against Chat-Only with paired Wilcoxon tests, while comparative preferences used chi-square goodness-of-fit tests.The three-way chi-square test assessed departure from uniformity rather than pairwise MOONWALK-versus-Existing-Workflow significance.
6 RESULTS & FINDINGS
MOONWALK produced positive results across alignment, executable review output, and review awareness measures. Participants nevertheless retained stronger preferences for existing workflows on some agency-related judgments, and the interface did not resolve every usability concern.
- Overall findings: MOONWALK was significantly above neutral on 12 of 13 Likert items and received the plurality on eight of nine comparative questions.Sense of control was the comparative-question exception.
- RQ1: Improving Collaboration & Intent Alignment: For articulating creative intent and review criteria, 63% chose MOONWALK versus 11% for Chat-Only and 26% for Existing Workflow.The three-way goodness-of-fit test was significant at p < .05.
- RQ1: Improving Collaboration & Intent Alignment: MOONWALK was rated significantly above neutral for proactive clarification, reference/specification-linked feedback, and handling conflicts between system suggestions and personal judgment.The corresponding significance levels were p < .01, p < .01, and p < .05.
- RQ2: Improving Outcomes & Efficiency: For junior-executable checklists, 74% chose MOONWALK, while 63% chose it for reducing senior–junior clarification.Chat-Only received 2 and 0 selections respectively; Existing Workflow received 3 and 7.
- RQ2: Improving Outcomes & Efficiency: Participants linked checklist executability to explicit reference and specification targets, estimating that juniors could independently resolve 70–80% of foundational errors.Examples included scale, color-temperature relations, and consistency with supplied references.
- RQ3: Review Awareness, Traceability, & Human Agency: Artifact–intent gap reflection had the highest mean across all 13 items at M = 5.74, with p < .01.Participants also tied the result to persistent specifications and broader coverage beyond the reviewer’s current focus.
- RQ3: Review Awareness, Traceability, & Human Agency: For blind-spot identification and evidence-linked feedback, 79% chose MOONWALK and none chose Chat-Only in either question.Both comparisons were significant at p < .01.
- RQ3: Review Awareness, Traceability, & Human Agency: Although traceability was significantly above neutral and 74% chose MOONWALK for tracking decision evolution, 47% chose Existing Workflow for sense of control versus 37% for MOONWALK.The sense-of-control comparison was not significant (p = .23), reinforcing the need for practitioner judgment.
7 DISCUSSION, LIMITATIONS, AND FUTURE WORK
The discussion frames MOONWALK as a shared intent–evidence–action record that supports role-specific coordination while preserving human creative authority. It also identifies mentorship, visual-native interaction, tool integration, evidential completeness, and study design as boundaries for deployment and future work.
- Discussion: Participants valued persistent specifications and references, rationale-linked judgments, and prioritized action records returned to artists.These properties support continuity of intent, evidence, and action across senior–junior handoffs.
- Collaboration limits: The prototype stored artists’ interpretations with submitted work, but the study did not test an explicit mutual confirmation step before review.Future work should isolate whether confirmation improves review outcomes.
- Role-based coordination: The shared record supports different role needs: seniors need missing-evidence and target identification, while juniors need explicit requirements, references, and bounded next actions.The same workflow can provide high-level discrepancy flags to supervisors and step-by-step guidance to junior artists.
- Agency and mentorship: Human authorization preserves practitioners’ final production decisions, but grounded outputs and artists’ ability to question them remain necessary for creative agency.The paper also warns that automation may reduce tacit mentorship interactions, so evidence checking should remain distinct from teaching.
- Limitations and future work: Text-heavy output, separate-platform friction, incomplete context, and exploratory single-session evaluation constrain production adoption and causal interpretation.The authors call for visual previews, infrastructure integration, low-friction context ingestion, structured non-AI baselines, counterbalanced studies, and longitudinal production metrics.
8 CONCLUSION
The paper presents MOONWALK as an intent–evidence–action framework and working review system for professional junior–supervisor workflows. Its evaluation supports structured review as an alternative to unstructured conversational AI while retaining aesthetic authority and final judgment with practitioners.
- Contribution: MOONWALK preserves articulated intent, inspectable evidence, and supervisor-approved actionable revision tasks across iterative junior–supervisor handoffs.The system is presented as a working pre-production review system.
- Findings: Practitioners valued persistent specifications and references, evidence-linked review records, and executor-ready checklists, alongside role-based differences and over-reliance risks.The evaluation covered formative and within-subject in-studio studies.
- Scope: The findings support MOONWALK as a structured alternative to unstructured conversational AI within the evaluated tasks and materials.The authors do not claim that AI is necessary beyond structured review support or isolate individual mechanisms.
Appendix A: Technical Implementation Details
The implementation stores shared project evidence, analyzes work-in-progress across multiple dimensions, and synthesizes model outputs into reference-grounded observations for human inspection.
- Project state: The project state stores typed brief fields, categorized and prioritized references, annotation notes, artist interpretations, and prior review decisions.Analysis, review, and consolidation read the same per-project evidence state.
- Analysis pipeline: The system analyzes work-in-progress across eleven dimensions using image-derived signals, lexical alignment, internal confidence, and agreement to organize candidate discrepancies.The dimensions include lighting, composition, color, style, perceptual quality, sketch/line quality, specification faithfulness, controllability, consistency, efficiency, and stability.
- Feature groups: Five image-signal groups cover composition, lighting and color, no-reference quality, artifact checks, and style or prompt alignment.These signals are combined with the active specification, reference notes, and the artist’s submitted interpretation.
- Model synthesis: A three-model pipeline generates observations, compares them with specification and reference context, and synthesizes analyses with artwork and reference metadata.Passing reference pixels and metadata enables outputs to cite visual regions rather than only filenames.
Appendix B: Analysis Dimensions Definitions
The supplied appendix material describes the analysis rubric alongside formative-study participants, procedures, interview topics, and post-task questionnaire items. Together, these passages connect technical dimensions with practitioner-centered evaluation of review coordination.
- Analysis dimensions: Seven analysis dimensions are artifact- and evidence-facing, while four others come from an earlier generation-evaluation rubric.The evidence-facing group includes visual properties, perceptual and artifact quality, and specification faithfulness.
- Formative study: The formative study involved practitioners across animation and VFX roles, studio contexts, and production types, using interviews, follow-ups, transcription, and thematic analysis.The supplied participant passages report directors, a supervisor, artists, and a project manager across advertising, character animation, virtual production, and film/TV VFX.
- Evaluation measures: The post-task items evaluated proactive context requests, reference-linked feedback, conflict resolution, translation of intent, junior-executable checklists, and follow-up effort.These items operationalize the workflow’s intended coordination benefits.
Section 3 | Improve Self-Reflection
The self-reflection evaluation examined whether the system helped practitioners clarify intent, identify blind spots, trace decisions, and retain control while structuring feedback rather than replacing judgment.
- Intent clarification: It examined whether the system helped reviewers compare current artifacts with original intent instead of relying on intuitive judgment alone.
- Blind-spot discovery: The study assessed whether system feedback exposed overlooked details and review blind spots.
- Decision traceability: It evaluated whether reviewers could trace each decision to its rationale and supporting references or specifications across iterative reviews.
- User control: The evaluation considered whether the interface reduced information overload and preserved users’ control over workflow direction and final decisions.
- Mediator role: The study also asked whether the system functioned as a coordinator or mediator rather than replacing practitioner judgment, alongside overall preference and satisfaction.
- Intent clarification: The evaluation asked whether the system clarified creative intent and review criteria before feedback.
Part 0 | Overall Experience
The overall-experience evaluation explored MOONWALK’s differences from prior workflows, its effects on collaboration and execution, and practitioners’ perceptions of its structured mediator role and implementation boundaries.
- Overall experience: Participants were asked to describe MOONWALK’s overall experience and its most fundamental difference from their previous workflow.
- Intent and collaboration: The evaluation examined whether structuring creative intent through the system differed from relying on memory or chat logs.
- Intent and collaboration: It assessed communication efficiency between senior and junior artists and the value of linking review items to specific references or specifications.
- Outcome and efficiency: The study asked whether junior artists could execute generated checklists without follow-up questions and whether the system reduced late-stage decision changes.
- Outcome and efficiency: It examined whether tracking decision evolution across iterative reviews addressed inconsistencies in review standards.
- Integration and implementation: Participants discussed challenges, missing features, future workflow integration, suitable team sizes, and the system’s multi-agent review implementation.
- Integration and implementation: The implementation boundary notes that some agent prompts request critical issues, negative impacts, target values, and verdicts without explicit abstention when evidence is missing.