This study benchmarks STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) for improving IELTS writing performance through multi-round AI and teacher-supported revisions.
Method Article
This study benchmarks STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) for improving IELTS writing performance through multi-round AI and teacher-supported revisions.
Automated writing feedback systems are prevalent, yet most deliver static, fragmented comments that provide limited scaffolding for revision. This study evaluates the STEP-DWCF-R framework (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent), in which AI-generated feedback, moderated by a teacher, is delivered via a robotic AI agent across multiple iterative rounds within a one-week task cycle. In an eight-week quasi-experimental trial, 32 EFL learners were randomized to either traditional written corrective feedback (one round per task) or STEP-DWCF-R. Both groups completed IELTS Task 2 essays at baseline and post-test, which were anonymized, randomized, and scored by two independent raters (ICC = 0.86–0.93). Linear mixed-effects models demonstrated that the STEP-DWCF-R group exhibited significantly greater gains in overall band score (Δ = 1.03 vs. 0.31 bands) and across all four analytic dimensions, with the largest improvement observed in Coherence and Cohesion. Process data indicated that STEP-DWCF-R learners completed an average of 2.26 revision rounds per task, with error counts decreasing linearly across rounds. These findings suggest that the integrated STEP-DWCF-R framework, encompassing AI-generated feedback, teacher moderation, and iterative robotic AI agent delivery, is associated with greater IELTS writing improvement than traditional single-round feedback, pointing to practical applications for AI-enhanced dynamic feedback in EFL contexts.
Written corrective feedback (WCF) remains central to L2 writing pedagogy. Evidence shows that comprehensive WCF improves learners’ accuracy over time and, when aligned with classroom practice, can coexist with focused approaches that are feasible in authentic contexts1,2,3. Recent classroom studies show durable gains in accuracy and fluency from sustained, comprehensive WCF. Research on feedback scope cautions that teachers should target forms strategically rather than mark everything4,5,6,7,8,9. In parallel, automated writing evaluation (AWE) and automated written corrective feedback (AWCF) have grown rapidly. Findings are mixed: some studies show gains in task achievement and grammatical range with AWCF or Criterion-style programs, while others find no significant advantage over teacher-only feedback or note largely local, surface-level revisions. To reflect this heterogeneity, we deliberately triangulate across multiple AWCF studies rather than leaning on a single source10,11,12,13.
Learners’ beliefs about who is providing feedback also matter: quasi-experimental work on perceived source indicates that performance and trust can shift when identical feedback is framed as “teacher” versus “automated,” underscoring the need to account for perception effects in AWCF designs14,15. Extending this line, emerging studies suggest that educational robots, due to their embodied presence, can also act as feedback providers, potentially influencing learner engagement and trust in distinctive ways16.
A complementary line of research, dynamic written corrective feedback (DWCF), emphasizes frequent, manageable, comprehensive feedback cycles with coding and rapid revision. Evidence from multiple settings suggests DWCF reliably improves accuracy (with frequency effects on fluency), though results may vary by course goals and learner population17,18,19,20. Finally, early studies on generative-AI-mediated feedback report parity with teacher feedback on writing outcomes and potential benefits when paired with explicit metalinguistic guidance, motivating closer integration of human and AI feedback in L2 contexts21,22.
To address these challenges, we propose the STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) framework, a novel multi-stage DWCF feedback system designed to provide structured, iterative, and process-oriented writing support. Our system leverages the predictive and generative power of Large Language Models (LLMs), with delivery through a robotic AI agent interface and teacher moderation, to segment complex writing issues into manageable categories, deliver feedback through multiple rounds of interaction, and evaluate its effectiveness using both outcome-based and process-based metrics. By transforming feedback into a structured and interactive process, STEP-DWCF-R advances automated writing support from static error correction toward a more engaging, embodied, and learning-oriented system.
Our contributions center on three integrated advances. First, we propose a structured feedback representation schema that categorizes errors (e.g., grammar, coherence) so that robotic AI agent-delivered LLM feedback is both actionable and pedagogically interpretable. Second, we introduce an iterative multi-stage process that sequences feedback across language, structure, and reasoning across multiple rounds to reduce cognitive load and deepen engagement. Third, we extend evaluation beyond post-scores to include process-oriented metrics such as error reduction and feedback adoption, offering a richer view of learning impact. WCF frameworks represent the earliest attempts to systematize written corrective feedback within educational technology. Originating in Automated Writing Evaluation (AWE) and Automated Essay Scoring (AES), these systems emphasized error detection and proficiency scoring13,14. Their contribution lies in scalability: thousands of learners could receive feedback without the bottleneck of human grading. However, these frameworks typically provided static and fragmented comments, such as grammar checks, lexical suggestions, or global scores that learners struggled to transform into meaningful revisions23,24,25,26.
Over time, WCF research evolved to include more technology-mediated approaches. These newer systems sought to capture broader aspects of writing by leveraging computational methods13,21. More recently, the scope has expanded further with embodied technologies, where educational robots have begun to act as feedback providers in writing or tutoring contexts. Their physical presence and interactive affordances introduce new dynamics of trust, engagement, and learner motivation. Yet, despite these developments, the dominant orientation of WCF has remained outcome-focused8,27: producing scores or isolated comments in a one-shot manner. Pedagogically, this orientation is misaligned with formative assessment, where feedback should scaffold revision across drafts and support metacognitive engagement16,28,29,30,31. Thus, the conceptual boundary of WCF reveals an enduring gap: feedback is delivered, but not structured or sequenced to guide learning processes.
In response to these limitations, scholars have begun to explore more dynamic approaches to writing feedback. Early investigations highlighted the importance of iterative feedback cycles, where learners revise multiple times under scaffolded guidance17,32. Structured error categorization has also been proposed to improve the interpretability of feedback and align it with pedagogical objectives33,34. The notion of a DWCF framework consolidates these directions by embedding three design principles: explicit error schemas, staged and iterative delivery, and process-oriented evaluation metrics19,35. Rather than overwhelming students with a bulk of corrections at once, DWCF distributes attention across layers of writing—surface accuracy, structural coherence, and argumentative depth—through successive rounds17,33. Evaluation expands beyond final scores to include process indicators such as error reduction rates, feedback adoption, and revision depth10,19,25. Despite its promise, practical applications of DWCF remain sparse, and most studies stop short of operationalizing it in large-scale, technology-mediated contexts10,36,37. This gap highlights the need for frameworks that not only theorize DWCF but also demonstrate its feasibility in real-world educational settings. Figure 1 shows the broader theoretical framework of DWCF proposed in the literature. The figure provides a general rationale for how dynamic written corrective feedback may operate at multiple levels, including both measured outcomes (e.g., accuracy, coherence) and theorized but unmeasured constructs in the present study (e.g., writing anxiety reduction, fluency, and complexity). It is included for theoretical orientation rather than as a direct model of this trial. Elements corresponding to the outcomes and process metrics analyzed here are indicated in the figure; other pathways are illustrative and were not evaluated in this study.
The emergence of LLMs and multimodal systems offers a powerful means to operationalize DWCF at scale. LLMs, with their zero-shot and few-shot capabilities, can generate context-sensitive feedback without task-specific training21,22. Recent experiments suggest their potential for directive segmentation, staged refinement, and explanation generation, which align closely with the principles of DWCF13,37,38,39. Complementary research in human–AI collaboration has shown that iterative loops of engaging with, adapting, and selectively adopting AI feedback can improve both uptake and revision quality25,30. When these processes are embodied through robots, the interaction has the theoretical potential to become multimodal—combining gesture, voice, and presence—which may further enhance learner perception and motivation40.
Grounded in the Zone of Proximal Development (ZPD)41, the Input Hypothesis42, and the Skill Acquisition Theory43, we position STEP-DWCF-R as a controller that links learner support, input processing, and skill development. Concretely, rubric-linked itemization, timely consolidated lists, and multi-round teacher-moderated AI–, human–, and robot–cycles operate at an integration layer that mediates revision and yields the observable endpoints analyzed here (IELTS Overall and TR/CC/LR/GRA; rounds, uptake, persistent errors, time). Constructs such as writing anxiety reduction are theorized but not measured in this trial.
Access restricted. Please log in or start a trial to view this content.
This protocol was approved by the Ethics Committee of the School of Primary Education, Shangrao Preschool Education College. All participants were adults (≥18 years) and provided written informed consent prior to the study.
1. Study design
NOTE: Due to the constraints of the available EFL classroom (total enrollment N = 32), participants were randomized into two groups of 16 each. This sample size, while modest, was determined to be sufficient for a pilot investigation given the within-subject pre-post design and expected large effect sizes based on prior DWCF studies17,18,19,20.
2. Writing Tasks
3. Feedback Workflow
4. Outcome Measurement
5. Process Measurement
6. Data Analysis
7. Code Availability
All code, prompts, JSON schemas, and example implementation files required to reproduce the STEP-DWCF-R system are openly available at https://github.com/YifangGaoinPG/Dynamic-Writing-Correction-Feedback.
Access restricted. Please log in or start a trial to view this content.
Experiments and Analysis
All 32 learners provided baseline data. Post-test data were available for 16 learners in each group (WCF and STEP-DWCF-R; Figure 2). The study timeline and measures are presented in Figure 3, and the end-to-end feedback workflow is illustrated in Figure 4. Experimental setup and implementation parameters for our AI system are shown in Table 1. Ana...
Access restricted. Please log in or start a trial to view this content.
This study compared STEP-DWCF-R, a multi-component dynamic written corrective feedback framework with a robotic AI agent, with single-round WCF in an eight-week IELTS writing course. STEP-DWCF-R yielded larger pre–post gains in overall band and across TR, CC, LR, and GRA. Learners under STEP-DWCF-R also completed more revision rounds and showed faster error reduction, with models estimating a decline of about 4.1 errors per round. The greatest improvement was in coherence and cohesion (ΔΔ = 0.85), indicat...
Access restricted. Please log in or start a trial to view this content.
The authors have no competing interests.
This work was funded by a Universiti Sains Malaysia Bridging Grant, Project No: R501-LR-RND003-0000001342-0000.
Access restricted. Please log in or start a trial to view this content.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| Flask (Web Framework) | Pallets Projects | https://flask.palletsprojects.com | Version 2.1; used for robotic AI agent web interface |
| GPT-5 API | OpenAI | https://platform.openai.com | Model gpt-5, temperature=0.7; for structured JSON feedback generation |
| Jinja2 | Pallets Projects | https://jinja.palletsprojects.com | Server-rendered templates for web interface |
| Python | Python Software Foundation | https://www.python.org | Version 3.10; backend scripting |
| Windows 11 | Microsoft | https://www.microsoft.com/windows | Operating system for AI feedback system |
Request permission to reuse the text or figures of this JoVE article
Request Permission