Source-linked AI summary

The Brand War: A Gamified AI-Feedback System for Time-Limited EFL Writing

Jing-Yuan Huang, Vivien Lin, Yujong Park, Yi Miao, Yun-Hua Hsiao, Michael Pin-Chuan Lin, Daniel Chang, Seong Min Park, Marco Ho, Michael S. Hsiao, Jeeho Ryoo

arXiv:2608.28604v1cs.CY

TL;DR

Timed EFL writing remains under-supported, particularly in real-world time-constrained settings. The Brand War embeds iterative GPT-4.1 scoring in a competitive narrative-writing game; scores improved across revisions, and AI and human evaluations converged strongly.

  • Problem

    Timed EFL writing remains under-supported despite its prevalence in exams, assessments, and workplace tasks.

  • Method

    The Brand War is a web-based competitive game in which Taiwanese undergraduate EFL students write a 500-word brand story within 60 minutes using up to five GPT-4.1 scoring passes.

  • Results

    AI scores rose across revisions, multi-cycle completers scored significantly higher on final versus first reviews (p = .032), and last AI scores correlated strongly with human final scores (r = 0.722, p < .001).

  • Takeaways & Limitations

    Integrating AI evaluative scaffolding with competitive game mechanics appears pedagogically feasible for EFL writing, while students generally prioritized revision over competition.

  • Takeaways & Limitations

    The exploratory single-session, non-experimental study was small, and the system provided scores and dimensional breakdowns rather than explicit revision suggestions.

Abstract

from arXiv · show

Writing is cognitively demanding and anxiety-provoking for English as a Foreign Language (EFL) learners, especially under time pressure. This paper presents The Brand War, a web-based gamified writing application combining competitive game mechanics with iterative GPT-4.1-powered formative feedback for undergraduate EFL learners completing a timed narrative writing task. Students role-play as marketing interns competing for a job offer, using review passes to receive AI feedback, attack opponents, or shield their own passes while drafting a 500-word brand story. We conducted an exploratory single-session classroom study with 29 university EFL students in Taiwan to examine engagement patterns, whether iterative AI feedback improved writing performance across revisions, and how AI and human scores related to overall outcomes. Students wrote within 60 minutes, using up to five AI feedback passes before a final human-graded submission. Most (65.5%) used the AI feedback system, and within-student AI scores improved modestly across revisions (M = +3.7, SD = 7.4), with larger gains among students completing more cycles and significantly higher final- versus first-review scores among multi-cycle completers (p = .032). AI-assessed and human final scores showed strong convergent validity (r = 0.722, p < .001), and AI-feedback users scored descriptively, though not significantly, higher than non-users. Students maintained a high mean focus ratio (82.4%), and competitive mechanics were used sparingly, suggesting most prioritized writing over social interference even when available. Findings suggest embedding iterative AI scoring within a competitive game context is feasible and may scaffold writing improvement, with implications for EFL writing pedagogy and AI-mediated gamified learning design.

I. INTRODUCTION

EFL writing becomes especially demanding under time pressure, while gamification and AI feedback offer complementary but rarely integrated forms of support. The Brand War combines both in a timed, competitive writing environment and examines engagement, revision performance, and score relationships.

  • Time-constrained EFL writing compounds cognitive and emotional difficulty, while instructional support for these conditions remains underdeveloped.
  • Gamification can increase motivation and engagement, but most gamified writing systems lack substantive, adaptive writing feedback.
  • The Brand War places undergraduate EFL students in a 60-minute, 500-word brand-story competition with five GPT-4.1-powered AI feedback passes.
  • Students can use passes for iterative feedback, attack opponents’ passes, or shield their own passes while drafting.
  • An exploratory classroom study with 29 university EFL students in Taiwan investigates engagement, successive revision performance, and relationships between AI and human scores.

II. LITERATURE REVIEW

Research distinguishes gamification from full game-based learning and reports that game-related writing designs can support engagement, while highlighting limits and design trade-offs.

  • Gamification applies game design elements such as points, levels, and leaderboards to non-game activities rather than creating a complete learning game.
  • Studies report improved engagement and writing quality in time-limited tasks, although excessive gamification can reduce motivation after reward saturation.
  • A review of 22 studies found digital games were the most common format, with competition and storylines supporting engagement but substantive-feedback designs remaining scarce.
  • A competition-and-collaboration language game was accepted across proficiency levels, with weaker students gaining the most, suggesting competition can be tuned rather than treated as binary.

B. Game-Based Learning and Narrative Writing

Narrative game-based learning gives EFL writing a purposeful professional context, while AI extends interactive game support into formative scoring. The Brand War builds on these approaches through self-regulated and motivational design.

  • B. Game-Based Learning and Narrative Writing: Narrative framing makes writing purposeful and audience-directed, reducing abstraction associated with EFL writing anxiety and L2 fatigue.
  • B. Game-Based Learning and Narrative Writing: The Brand War extends AI from an interactional NPC role to a formative scoring agent within a professional narrative scenario.
  • B. Game-Based Learning and Narrative Writing: Prior work found AI feedback approached human quality for surface-level linguistic features, while human raters retained an advantage in nuanced holistic judgment.
  • B. Game-Based Learning and Narrative Writing: The design maps competition to relatedness, iterative scoring to competence, and pass choices to autonomy within Self-Determination Theory.
  • B. Game-Based Learning and Narrative Writing: Self-Regulated Learning frames the system as supporting cycles of planning, monitoring, and reflection during writing and revision.

E. Research Gap

The Brand War addresses a research gap by integrating competitive mechanics, narrative writing, and iterative AI formative scoring in a live EFL classroom. Its task uses a timed professional brand-story genre aligned with the game scenario.

  • E. Research Gap: Existing research examines gamification, narrative contexts, and AI writing feedback separately, but not their integration in one real-classroom system.
  • E. Research Gap: The study implements this integration through an exploratory classroom deployment with undergraduate EFL learners in Taiwan.
  • E. Research Gap: Students write a 500-word narrative brand story within 60 minutes for a mandatory undergraduate EFL writing course.
  • E. Research Gap: The task requires setting, theme, mood, character, and plot, linking instructional content to professional brand-storytelling demands.

B. Game Scenario and Storyline

The Brand War frames timed EFL writing as a competitive internship challenge, linking career relevance with optional social competition and strategic pass use.

  • B. Game Scenario and Storyline: Students role-play marketing internship candidates competing for a job offer by completing a brand story.The scenario places the task at Global Marketing, Inc. in Taipei 101.
  • B. Game Scenario and Storyline: The storyline grounds writing in a professional situation, creates competitive stakes, and connects with students’ career aspirations.The paper relates these features to relevance as a driver of autonomous motivation.
  • B. Game Scenario and Storyline: Students can spend passes on AI feedback or attack opponents, creating a trade-off between improving their draft and reducing another student’s revision capacity.

D. AI Feedback and Scoring System

The AI system provides rubric-aligned GPT-4.1 scoring during drafting, giving students diagnostic information for self-directed revision without generating explicit rewrite suggestions.

  • D. AI Feedback and Scoring System: GPT-4.1 scores drafts on Content, Mechanics, and Structure, each from 0–100, and displays their composite average.The rubric covers narrative content, language mechanics, and organization.
  • D. AI Feedback and Scoring System: The system functions as evaluative scaffolding by providing scores and dimensional breakdowns rather than explicit revision suggestions.This design preserves writing agency while supplying diagnostic signals for revision.
  • D. AI Feedback and Scoring System: Platform logs capture focus and blur events, submission timestamps, and attack or shield activations for engagement measurement.Derived metrics include focus ratio, focus sessions, task duration, and passes used.

A. Participants

The exploratory classroom study involved 29 Taiwanese university EFL students completing a structured, approximately 80-minute session with timed writing, optional AI review, and logged revision activity.

  • A. Participants: Participants were 29 university students in an elective EFL writing course at a Taiwanese university.Two students lacked final essays for human-score analyses, and one lacked focus-event logs.
  • A. Participants: The classroom session lasted approximately 80 minutes and was organized into five phases.
  • A. Participants: The writing and revision phase lasted approximately 60 minutes, with optional GPT-4.1 scoring or opponent attacks before final submission.
  • A. Participants: Review-level data comprised 43 submissions from 19 students, while student-level data covered all 29 participants.Review records included attempt and dimensional AI scores; student records included engagement, AI gains, mechanics, and human scores.

D. Analysis

The study used exploratory descriptive, nonparametric, correlational, and validity analyses to examine engagement, revision changes, group differences, and relationships between AI and human scores.

  • D. Analysis: Analyses summarized key variables and mean AI composite-score trends across attempts 1–5.
  • D. Analysis: Wilcoxon signed-rank tests compared first and final reviewed submissions among students completing at least two review cycles.
  • D. Analysis: Mann-Whitney U tests compared human final scores between AI-feedback users and non-users.The comparison involved 19 users and 8 non-users.
  • D. Analysis: Spearman correlations examined focus ratio with human scores and passes used with AI score gain, while Pearson correlation assessed AI–human convergent validity.
  • D. Analysis: 19 of 29 students (65.5%) used at least one AI review pass, while mean focus ratio was 0.824 and attacks were used by 17.2%.No student used a shield; one pass was the most common usage level.

B. RQ2. Writing Performance Across AI Feedback Attempts

AI-assessed writing scores generally increased across successive review attempts, with significant first-to-final gains among students completing at least two cycles. AI–human score convergence was strong, while engagement measures showed weaker relationships with writing outcomes.

  • Performance across attempts: M = 3.7 (SD = 7.4) was the mean AI score gain from first to last attempt among 19 AI-feedback users.Individual gains ranged from −5 to +22.
  • Performance across attempts: M ≈65.6 at Attempt 1 increased to ≈79–81.3 by Attempts 4–5, although later attempts included only three students each.The steepest gain occurred between Attempts 2 and 3.
  • Performance across attempts: Among 11 students completing at least two review cycles, final-review scores were significantly higher than first-review scores (W = 8.5, p = .032).The reported rank-biserial effect size was .74.
  • Relationship to outcomes: r = 0.722 (p < .001) indicated strong convergence between students’ last AI review scores and human final scores.Focus ratio was not significantly associated with human final score, with Spearman ρ = −0.147 (p = .464).
  • Relationship to outcomes: AI-feedback users scored descriptively higher than non-users on human final scores, but the difference was not significant (M = 76.9 vs. M = 70.4, p = .167).Passes used showed a positive but non-significant trend with AI score gain (ρ = 0.428, p = .068).

VI. DISCUSSION

The discussion interprets rising AI-assessed scores as evidence that evaluative scaffolding can support revision within the game. It also qualifies this interpretation because feedback was minimalist and later-attempt results were affected by attrition and self-selection.

  • AI Feedback as Evaluative Scaffolding: M = 65.6 at Attempt 1 rose to M = 81.3 at Attempt 5, which the authors interpret as productive responses to GPT-4.1 scoring.The discussion frames scoring as externalizing rubric expectations within a monitor–evaluate–adjust loop.
  • Evidence overview: Table II summarizes the study’s key statistics, including revision trends and relationships between AI scores, human scores, and engagement measures.The supplied table caption identifies the table but does not provide its individual entries.
  • AI Feedback as Evaluative Scaffolding: The system delivered scores and dimensional breakdowns rather than explicit revision suggestions, requiring students to infer what to change.This design was intended to preserve writing agency and reduce over-reliance on generated text.
  • Limitations: Later-attempt gains are constrained by declining participation from n = 19 at Attempt 1 to n = 3 at Attempt 5.The discussion attributes this pattern to self-selection by students who were more motivated or saw greater improvement potential.

B. Gamification, Engagement, and Self-Determination

The Brand War sustained high behavioral engagement while students generally favored constructive AI feedback over offensive or defensive competition. AI and human scores also showed strong convergence, supporting AI scoring as a formative aid while preserving human evaluation for summative judgment.

  • Engagement: M = 0.824 focus ratio indicates sustained behavioral engagement throughout the 60-minute writing task.The focus ratio was interpreted as evidence that the game context maintained engagement during the timed session.
  • Mechanic use: 19 students (65.5%) used AI feedback, compared with 5 students (17.2%) who attacked and none who shielded.Students therefore used constructive revision support more often than competitive interference mechanics.
  • Outcomes: AI-feedback users scored descriptively higher than non-users, but the group difference was not significant (p = .167).Users had M = 76.9 versus M = 70.4 for non-users, and the study was not powered to detect causal effects.
  • AI-human convergence: r = 0.722, p < .001 between AI last scores and human final scores supports convergent validity.The result supports AI scoring as a real-time formative tool while human scoring remains appropriate for summative evaluation.
  • Implications: The study concludes that iterative AI scoring within a competitive game context is pedagogically feasible for undergraduate EFL writing.The exploratory classroom study involved N = 29 students and found high focus alongside a preference for revision over competition.
Loading 2608.28604v1…