Method Article

A Dynamic Written Corrective Feedback Framework Integrating AI Agent Delivery for Structured and Iterative Essay Support

DOI:

10.3791/71992

July 24th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study benchmarks STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) for improving IELTS writing performance through multi-round AI and teacher-supported revisions.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Automated writing feedback systems are prevalent, yet most deliver static, fragmented comments that provide limited scaffolding for revision. This study evaluates the STEP-DWCF-R framework (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent), in which AI-generated feedback, moderated by a teacher, is delivered via a robotic AI agent across multiple iterative rounds within a one-week task cycle. In an eight-week quasi-experimental trial, 32 EFL learners were randomized to either traditional written corrective feedback (one round per task) or STEP-DWCF-R. Both groups completed IELTS Task 2 essays at baseline and post-test, which were anonymized, randomized, and scored by two independent raters (ICC = 0.86–0.93). Linear mixed-effects models demonstrated that the STEP-DWCF-R group exhibited significantly greater gains in overall band score (Δ = 1.03 vs. 0.31 bands) and across all four analytic dimensions, with the largest improvement observed in Coherence and Cohesion. Process data indicated that STEP-DWCF-R learners completed an average of 2.26 revision rounds per task, with error counts decreasing linearly across rounds. These findings suggest that the integrated STEP-DWCF-R framework, encompassing AI-generated feedback, teacher moderation, and iterative robotic AI agent delivery, is associated with greater IELTS writing improvement than traditional single-round feedback, pointing to practical applications for AI-enhanced dynamic feedback in EFL contexts.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Written corrective feedback (WCF) remains central to L2 writing pedagogy. Evidence shows that comprehensive WCF improves learners’ accuracy over time and, when aligned with classroom practice, can coexist with focused approaches that are feasible in authentic contexts1,2,3. Recent classroom studies show durable gains in accuracy and fluency from sustained, comprehensive WCF. Research on feedback scope cautions that teachers should target forms strategically rather than mark everything4,5,6,7,8,9. In parallel, automated writing evaluation (AWE) and automated written corrective feedback (AWCF) have grown rapidly. Findings are mixed: some studies show gains in task achievement and grammatical range with AWCF or Criterion-style programs, while others find no significant advantage over teacher-only feedback or note largely local, surface-level revisions. To reflect this heterogeneity, we deliberately triangulate across multiple AWCF studies rather than leaning on a single source10,11,12,13.

Learners’ beliefs about who is providing feedback also matter: quasi-experimental work on perceived source indicates that performance and trust can shift when identical feedback is framed as “teacher” versus “automated,” underscoring the need to account for perception effects in AWCF designs14,15. Extending this line, emerging studies suggest that educational robots, due to their embodied presence, can also act as feedback providers, potentially influencing learner engagement and trust in distinctive ways16.

A complementary line of research, dynamic written corrective feedback (DWCF), emphasizes frequent, manageable, comprehensive feedback cycles with coding and rapid revision. Evidence from multiple settings suggests DWCF reliably improves accuracy (with frequency effects on fluency), though results may vary by course goals and learner population17,18,19,20. Finally, early studies on generative-AI-mediated feedback report parity with teacher feedback on writing outcomes and potential benefits when paired with explicit metalinguistic guidance, motivating closer integration of human and AI feedback in L2 contexts21,22.

To address these challenges, we propose the STEP-DWCF-R (Structured, Tiered, Evidence-driven Process for Dynamic Written Corrective Feedback with Robotic AI Agent) framework, a novel multi-stage DWCF feedback system designed to provide structured, iterative, and process-oriented writing support. Our system leverages the predictive and generative power of Large Language Models (LLMs), with delivery through a robotic AI agent interface and teacher moderation, to segment complex writing issues into manageable categories, deliver feedback through multiple rounds of interaction, and evaluate its effectiveness using both outcome-based and process-based metrics. By transforming feedback into a structured and interactive process, STEP-DWCF-R advances automated writing support from static error correction toward a more engaging, embodied, and learning-oriented system.

Our contributions center on three integrated advances. First, we propose a structured feedback representation schema that categorizes errors (e.g., grammar, coherence) so that robotic AI agent-delivered LLM feedback is both actionable and pedagogically interpretable. Second, we introduce an iterative multi-stage process that sequences feedback across language, structure, and reasoning across multiple rounds to reduce cognitive load and deepen engagement. Third, we extend evaluation beyond post-scores to include process-oriented metrics such as error reduction and feedback adoption, offering a richer view of learning impact. WCF frameworks represent the earliest attempts to systematize written corrective feedback within educational technology. Originating in Automated Writing Evaluation (AWE) and Automated Essay Scoring (AES), these systems emphasized error detection and proficiency scoring13,14. Their contribution lies in scalability: thousands of learners could receive feedback without the bottleneck of human grading. However, these frameworks typically provided static and fragmented comments, such as grammar checks, lexical suggestions, or global scores that learners struggled to transform into meaningful revisions23,24,25,26.

Over time, WCF research evolved to include more technology-mediated approaches. These newer systems sought to capture broader aspects of writing by leveraging computational methods13,21. More recently, the scope has expanded further with embodied technologies, where educational robots have begun to act as feedback providers in writing or tutoring contexts. Their physical presence and interactive affordances introduce new dynamics of trust, engagement, and learner motivation. Yet, despite these developments, the dominant orientation of WCF has remained outcome-focused8,27: producing scores or isolated comments in a one-shot manner. Pedagogically, this orientation is misaligned with formative assessment, where feedback should scaffold revision across drafts and support metacognitive engagement16,28,29,30,31. Thus, the conceptual boundary of WCF reveals an enduring gap: feedback is delivered, but not structured or sequenced to guide learning processes.

In response to these limitations, scholars have begun to explore more dynamic approaches to writing feedback. Early investigations highlighted the importance of iterative feedback cycles, where learners revise multiple times under scaffolded guidance17,32. Structured error categorization has also been proposed to improve the interpretability of feedback and align it with pedagogical objectives33,34. The notion of a DWCF framework consolidates these directions by embedding three design principles: explicit error schemas, staged and iterative delivery, and process-oriented evaluation metrics19,35. Rather than overwhelming students with a bulk of corrections at once, DWCF distributes attention across layers of writing—surface accuracy, structural coherence, and argumentative depth—through successive rounds17,33. Evaluation expands beyond final scores to include process indicators such as error reduction rates, feedback adoption, and revision depth10,19,25. Despite its promise, practical applications of DWCF remain sparse, and most studies stop short of operationalizing it in large-scale, technology-mediated contexts10,36,37. This gap highlights the need for frameworks that not only theorize DWCF but also demonstrate its feasibility in real-world educational settings. Figure 1 shows the broader theoretical framework of DWCF proposed in the literature. The figure provides a general rationale for how dynamic written corrective feedback may operate at multiple levels, including both measured outcomes (e.g., accuracy, coherence) and theorized but unmeasured constructs in the present study (e.g., writing anxiety reduction, fluency, and complexity). It is included for theoretical orientation rather than as a direct model of this trial. Elements corresponding to the outcomes and process metrics analyzed here are indicated in the figure; other pathways are illustrative and were not evaluated in this study.

The emergence of LLMs and multimodal systems offers a powerful means to operationalize DWCF at scale. LLMs, with their zero-shot and few-shot capabilities, can generate context-sensitive feedback without task-specific training21,22. Recent experiments suggest their potential for directive segmentation, staged refinement, and explanation generation, which align closely with the principles of DWCF13,37,38,39. Complementary research in human–AI collaboration has shown that iterative loops of engaging with, adapting, and selectively adopting AI feedback can improve both uptake and revision quality25,30. When these processes are embodied through robots, the interaction has the theoretical potential to become multimodal—combining gesture, voice, and presence—which may further enhance learner perception and motivation40.

Grounded in the Zone of Proximal Development (ZPD)41, the Input Hypothesis42, and the Skill Acquisition Theory43, we position STEP-DWCF-R as a controller that links learner support, input processing, and skill development. Concretely, rubric-linked itemization, timely consolidated lists, and multi-round teacher-moderated AI–, human–, and robot–cycles operate at an integration layer that mediates revision and yields the observable endpoints analyzed here (IELTS Overall and TR/CC/LR/GRA; rounds, uptake, persistent errors, time). Constructs such as writing anxiety reduction are theorized but not measured in this trial.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This protocol was approved by the Ethics Committee of the School of Primary Education, Shangrao Preschool Education College. All participants were adults (≥18 years) and provided written informed consent prior to the study.

1. Study design

NOTE: Due to the constraints of the available EFL classroom (total enrollment N = 32), participants were randomized into two groups of 16 each. This sample size, while modest, was determined to be sufficient for a pilot investigation given the within-subject pre-post design and expected large effect sizes based on prior DWCF studies17,18,19,20.

  1. Randomize participants into two groups (n = 16 each) using simple randomization. Generate the allocation sequence using a computer-generated random number list. Conceal allocation from the instructor until baseline assessment is completed.
  2. Ensure both groups are taught by the same instructor using identical materials and contact hours.
  3. Deliver feedback either through direct human correction (WCF) or through the STEP-DWCF-R workflow, where AI and teacher feedback are presented via a robotic AI agent.

2. Writing Tasks

  1. Administer a baseline IELTS Task 2 essay in Week 1 under exam conditions (≥250 words, 40 min).
  2. Conduct six weekly writing tasks (Weeks 2–7). For each task:
    1. In WCF, provide one round of human feedback on the essay.
    2. In STEP-DWCF-R, generate AI feedback, moderate it with a human teacher, and instruct the robotic AI agent to present the feedback interactively to the learner.
    3. In STEP-DWCF-R, repeat the STEP-DWCF-R feedback cycle for 2–3 rounds until the revision is complete.
  3. Administer a post-test IELTS Task 2 essay in Week 8 under the same exam conditions. Follow the same one-week calendar for each task in both groups. Deliver Round 1 feedback within 24 h of draft submission in STEP-DWCF-R. Require learners to revise and resubmit within 48 h. Follow the same schedule for subsequent rounds and complete all cycles within the weekly window.

3. Feedback Workflow

  1. Process each learner draft with the GPT-5 API (model = "gpt-5", temperature = 0.7, max_output_tokens = 2048) to generate itemized feedback across four fixed categories: Grammar, Vocabulary, Organization, and Reasoning. The exact system prompt and JSON output schema are provided in the open-source repository (https://github.com/YifangGaoinPG/Dynamic-Writing-Correction-Feedback/blob/main/evaluate.py, function build_prompt() and normalize_reasoning()).
  2. Moderate the AI-generated feedback according to the following criteria: remove false positives or inaccurate comments; add missed critical issues aligned with the IELTS rubric; ensure comments are actionable and encouraging; and maintain consistency with the four-category schema. Log moderation decisions manually.
  3. Save the moderated feedback in JSON format and upload it to the robotic AI agent interface (a lightweight Flask web application).
  4. Present feedback through the web interface. Display structured tables for each category (summary, issues, and revision tips). Record learner acknowledgments and clarification requests through system logs. Instruct learners to revise and resubmit their drafts. Repeat the cycle up to three times per task.
  5. Verification and troubleshooting.
    1. Verify successful AI feedback generation by confirming that the output is valid JSON containing all four required categories: Grammar, Vocabulary, Organization, and Reasoning. Confirm completion of teacher moderation by ensuring that false positives have been removed, missing issues have been added, and the moderated JSON file has been saved.
    2. Upload the moderated JSON file to the Flask-based robotic AI agent interface through the upload endpoint. Confirm successful upload through the interface confirmation message and database log entry. Record learner resubmissions automatically through form submission logs.
    3. Process non-compliant AI outputs (e.g., malformed JSON, missing categories, or inaccurate feedback) using the ensure_json() and normalize_reasoning() functions. Regenerate the API output with adjusted prompt parameters as needed. Manually correct the JSON during teacher moderation, if necessary, to ensure that all four categories are present before upload.
      NOTE: Examples of raw AI feedback, teacher-moderated versions, and the delivered interface output are available in the repository (data/ folder and notebooks/).

4. Outcome Measurement

  1. Define the primary outcome as the change in IELTS overall band score (Week 1 to Week 8).
  2. Define secondary outcomes as changes in Task Response, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy.
  3. Assign two independent raters with experience in IELTS Task 2 scoring to evaluate all essays. Remove names and group identifiers from all essays. Randomize essay order and blind raters to group assignment and time point. Use the same IELTS Task 2 prompt for both the Week 1 baseline and Week 8 post-test. Resolve scoring disagreements greater than 0.5 bands through discussion until consensus is reached. Assess inter-rater reliability using the average-measures intraclass correlation coefficient (ICC[2,2]).

5. Process Measurement

  1. Log the number of revision rounds per task.
  2. Record uptake of feedback (adopted, modified, declined) by learners.
  3. Track persistent errors between cycles.
  4. Record teacher feedback time, robot delivery time, and learner revision time.

6. Data Analysis

  1. Fit linear mixed-effects models with Group, Time, and Group × Time interaction.
  2. Apply generalized mixed-effects models for process measures.
  3. Report effect sizes, 95% confidence intervals, and adjusted p-values.
  4. Conduct robustness checks with ANCOVA, rater-specific models, and robust SEs.

7. Code Availability

All code, prompts, JSON schemas, and example implementation files required to reproduce the STEP-DWCF-R system are openly available at https://github.com/YifangGaoinPG/Dynamic-Writing-Correction-Feedback.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Experiments and Analysis

All 32 learners provided baseline data. Post-test data were available for 16 learners in each group (WCF and STEP-DWCF-R; Figure 2). The study timeline and measures are presented in Figure 3, and the end-to-end feedback workflow is illustrated in Figure 4. Experimental setup and implementation parameters for our AI system are shown in Table 1. Ana...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study compared STEP-DWCF-R, a multi-component dynamic written corrective feedback framework with a robotic AI agent, with single-round WCF in an eight-week IELTS writing course. STEP-DWCF-R yielded larger pre–post gains in overall band and across TR, CC, LR, and GRA. Learners under STEP-DWCF-R also completed more revision rounds and showed faster error reduction, with models estimating a decline of about 4.1 errors per round. The greatest improvement was in coherence and cohesion (ΔΔ = 0.85), indicat...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors have no competing interests.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This work was funded by a Universiti Sains Malaysia Bridging Grant, Project No: R501-LR-RND003-0000001342-0000.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Flask (Web Framework)Pallets Projectshttps://flask.palletsprojects.comVersion 2.1; used for robotic AI agent web interface
GPT-5 APIOpenAIhttps://platform.openai.comModel gpt-5, temperature=0.7; for structured JSON feedback generation
Jinja2Pallets Projectshttps://jinja.palletsprojects.comServer-rendered templates for web interface
PythonPython Software Foundationhttps://www.python.orgVersion 3.10; backend scripting
Windows 11Microsofthttps://www.microsoft.com/windowsOperating system for AI feedback system

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

AI Agent FeedbackIterative Essay RevisionIELTS Writing ImprovementEFL LearnersDynamic FeedbackTeacher ModerationRobotic AI AgentRevision RoundsLinear Mixed Effects

Related Articles