Research Article

Investigating The Effect Of Automated Writing Evaluation Through Deep Neural Network And Teachers' Written Evaluation On English Writing Performance

DOI:

10.3791/69995

May 8th, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Provision of feedback in the form of two NLP semantic-based features through the computer system and the teacher’s written feedback was employed to improve students' English writing fluency and correctness. Results showed positive results concerning the former.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Computer-assisted language learning technologies have made the greatest gains in language learning, especially in the context of artificial intelligence (AI), over the past few years. There are various technologies, such as automated writing evaluation (AWE hereafter) and automated essay scoring (AES), that are considered pillars of language learning, not only in computer disciplines but also in education. The evaluation can be done only in the form of holistic scores using AWE technology, which has been quite effective in enhancing language learning and, therefore, cannot provide detailed or in-depth feedback. To provide detailed writing feedback on two major elements of a language (i.e., Grammar and Fluency), a computer-aided implementation system entailing neural network models, and two semantic-based natural language processing (NLP hereafter) methods have been considered. To that end, 90 Pakistani university students who were English second-language learners (ESL) were randomly assigned to the control, instructor feedback, and experimental groups. The computer-assisted evaluation encompassed a neural network. The results of the comparison test between the AWE baseline model and the instructors who graded showed a correlation between the computer-aided feedback that used these neural networks. The implications of such results lie in ESL writing pedagogy, especially in situations where linguistic accuracy and fluency pose a challenge.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Feedback in writing addresses various attributes of language, such as organization, grammar, fluency, content, and mechanics, enabling learners to either rewrite or create a new piece of written text1,2,3. Presently, teachers of English writing have been a major beneficiary because of the fast changes and advancements of computer-assisted language learning (CALL hereafter) and computer-aided writing assessment in the shape of AWE or AES and the grading of written essays4. Additionally, computer-assisted assessment of English writing gives diagnostic feedback on linguistic elements such as content, grammar, vocabulary, etc., providing timely, individual, and objective responses to the written text of the students. The AWE system is also in a position to give clear written answers to the students in order to have them internalize the knowledge acquired through these written responses and increase the cognitive capacity6of the students.

The present study was based on a research4, which offered the concept of a multi-strategy, computer-based English writing learning with the English writing feedback theory in perspective and the process of the analytical scoring in mind. This included other NLP semantic-based models that incorporated deep learning into written feedback to offer elaborate written feedback scoring. It addresses the limitation of prior AWE technology, which only emphasizes the holistic but not the analytic, specific scoring. Consequently, it has more to do with accurate and immediate assessment with sub and total marks regarding a series of elements of linguistic indicators to improve English writing learning among students. Among various characteristics of linguistics (i.e., contents, complexities, correctness, fluencies), the present study will only examine fluency and the grammatical aspect of ESL learners in Pakistan. To this end, the current study suggests a strategy of NLP technology that entails the use of neural networks accompanied by teachers’ evaluation.

In the Pakistani language learning context, Pakistani students encounter issues in learning English writing, even though the students consider studying English as a compulsory subject starting at grade 1 and continuing up to grade 14. That is where numerous factors, such as wrong courses and the use of the old methods to explain the grammar, come into the picture7. At the university level, the number of students that teachers are required to teach is rather considerable since the students have an assortment of backgrounds that include various colleges and schools. This compels teachers to spend much of the time correcting students' writing mistakes, and this makes the workload way too high. On the other hand, students took the teachers' written feedback for granted and did not make efforts to recognize the mistakes and correct them8.

A study stated that fluency is mostly influenced by the fact that, because students are afraid of making an error, it becomes a barrier to writing fluently, and therefore, the students would rather avoid making mistakes more frequently than engaging with thoughts. There is occasional interference of the Urdu language with the English writing, in addition to other contextual factors. Furthermore, the Urdu language is the national language. As a result of these contextual factors, teachers find it difficult to provide precise, timely, and consistent feedback to every student whenever the students generate written output or edit the written manuscripts. Considering the theory of writing evaluation, this study will be able to follow the study that employs an automatic machine writing feedback strategy during the comparison of the traditional teacher evaluation and takes the analytic instead of the holistic scoring to ensure that students obtain instant feedback about the two significant linguistic dimensions (i.e. Grammar and Fluency)4.

Different industrial AWE and AES products have emerged, which include Pigai.org, Bingo English, My Access, E-rater, I-write, Criterion, etc. AES/AWE at its early stage of development is characterized mainly by the traditional machine learning system based on Bayes theorem10, linear regression11 etc. Currently, the extended development of deep learning, particularly the technology of neural network, has also remarkably benefited the field of writing feedback, such as recurrent neural networks or RNN12, long short-term memory13, Convolutional Neural Networks (CNN)14, BERT approach15. AWE also received benefits from the pre-training strategies employed at the start of NLP16. Many studies have taken into consideration various solutions focusing on neural networks, such as generating adversarial samples for the purpose of learning, learning through Multitask17, learning through Self-Supervision18, and Graph Algorithms19. Some have suggested different techniques to address the issue of the lack of training data for writing feedback20. There are also models of scoring that contribute to many neural network models21.

Grammatical error correction (GEC hereafter) is an independent way of NLP missions, which has the capacity to automatically detect and then correct the errors of composition. The task content of GEC is interrelated with that of the AWE aspects. AWE tasks are considered to take into account the diagnosis of grammatical errors, most of which are required to be highlighted in the correct form. However, the AWE studies have given less attention to the GEC, and only a few researches integrated the early GEC model into the AWE technology. Initially, research on GEC chooses N-gram most of the time22, and language model23 including BERT24 to identify and correct grammatical errors. However, the above solutions could not completely solve problems like disorder and component omission. The architecture of encoder-decoder25 has attained the same success in GEC by addressing the issues that other models could not address previously. It is important to note that advanced GEC architectures utilized multiple models to solve problems, including the Transformer-based approaches26.

Different studies confirmed the efficiency of AES as compared to human ratings when it comes to evaluating writing essays27,28. A study29 demonstrated the effect of automated evaluation on the English fluency of learners. The findings indicated that automatic evaluation had a positive effect on English writing fluency. In contrast, some other studies30disregard the high error detection rate by AES compared to human ratings. Recent studies highlight the growing effectiveness of AI-assisted tools for writing evaluation in language education. One such study31, made a comparison between automated grading with instructor’s feedback in the Chinese college students learning context.

The revolutionary role of AI-based tools in academic writing has been actively reported in the recent literature. Research concerning the application of AI-powered writing tools as supplements to the traditional EFL writing education highlights the paramount significance of equal opportunity and proactive teacher support in the successful implementation of technology32. The studies that investigated AWE, such as Pigai, indicated the strong correlation between the feedback under the support of computers and the enhancement of ESL writing, revision, and new pieces of writing33. Another study has highlighted the benefits of variational autoencoders (VAEs) that positively influence the aspect of training English writing in learners and get feedback on higher levels of deep learning, specializing in the various linguistic features34. Another research suggested that neural AWE systems would be effective to use with NLP-based GEC, which uses deep learning to offer an immediate evaluation35.

The improvement of multi-agent systems is an important step in the development of technologies that assist in writing. An example of being a multi-agent, an academic writing refinement assistant, the AcademyCraft, based on deep-learning, has demonstrated abilities to write and give detailed feedback, hence facilitating AI-based EFL/ESL writing support, upon complex analytical schemes36. In fact, technology-based language learning studies have also identified diverse emergent patterns like deep learning, automatic assessment, and semantic feedback with performance analysis that showed a gradual and consistent upsurge in writing accuracy, as well as learner engagement37.

Transformer-LSTM models have shown a certain propensity for automatic grading and feedback made on English writing. Such hybrid architectures mutually took the contextual understanding of transformers along with the sequential processing of LSTM systems, concluding to show an improved accuracy concerning grammar and coherence grading and provide more nuanced feedback generation38. Perception of EFL teachers on AI writing tools consistently reports measurable improvements in organization of content and writing quality, focusing on the crucial role of pedagogical help in triggering the usefulness of technological integration39. There are fewer studies that make a comparison between instructors’ feedback and computer-supported feedback, implying deep learning models to explore the effect on the English writing of ESL learners in terms of fluency and grammar.

Access restricted. Please log in or start a trial to view this content.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

All the software and tools used in the study are listed in the Table of Materials.

Grading criteria

The classification of evaluation adopted in the present study addresses two major domains, namely, fluency and correctness, which incorporate all the requirements for comprehensive evaluation of writing in English as a Foreign Language (EFL). These dimensions are intended to measure the key points of the writing quality that help in effective communication and demonstration of skill in language proficiency. The Fluency module measures naturalness, expressiveness, and coherence in writing by measuring the smoothness of ideas as well as how well words and sentence structure are used. The Correctness module targets precise grammatical accuracy, mechanical flawlessness, and compliance with standard English conventions and, hypothetically, measures and counts the language's erroneousness on several levels. (see Table 1)

Task definition

Following the conditions of automated writing assessment of students learning the English language, the present study suggests a new model of an EG computer-assisted English writing evaluation system. Compared to the previous studies, which have been limited by the scope of automated writing evaluation systems, the proposed system will help resolve those limitations through the introduction of innovative deep learning techniques to allow users a level of linguistic analysis, scope of modification, and system-wide feedback analysis based on the extensive processing of data. Operating on two of the most significant indicators of a composition, namely, fluency (how natural, coherent, and expressive in nature is the piece), or correctness (how correctly written the text, riddled with grammatical, mechanical and language errors is), with the separate components of the neural networks, the dual-model system has to perform a number of evaluations. Besides, it identifies mistakes on several linguistic levels, recommends correctional measures, and gives feedback in a list form to enable users to improve students’ language learning in a methodical way. In order to describe the process of calculating the computer-assisted automatic evaluation system, we suggest the following mathematical model in formulas (1)–(4) (see Table 2).

Final score equation, formula for weighted scoring method, Σ(wSf + wSc), educational use.     (1)

Σf=σ(BERT output); equation diagram in AI model analysis.    (2)

Static equilibrium formula \( S_c = 100e^{-\lambda \cdot \text{ErrorRatio}} \); equation analysis.   (3)

Sigmoid function equation σ(x)=1/(1+e^-x); diagram for logistic regression, neural networks.   (4)

where: Final Score = the composite writing evaluation score; w1 = weight assigned to fluency score; w2 = weight assigned to correctness score; Sf = BERT-based fluency score (normalized to [0, 1]); Sc = exponential decay correctness score; λ = decay constant controlling error penalty; σ(x) = logistic sigmoid activation function; ErrorRatio = ratio of errors to total tokens in the text. The fluency scores are normalized to [0, 1] by the logistic sigmoid function σ(x) whereas the exponential decay function is used in scoring the correctness to impose an appropriate penalty on error frequency.

System architecture and design

Its system architecture uses a two-model system with a parallel processing architecture, in which the analysis of input essays is performed in parallel via parallel evaluation pathways. The Fluency and Correctness modules, in turn, yield respective scores, and the final composite score is the weighted average of the two sub-scores. The correctness module also offers a suggestion on error detection and correction, besides scoring (see Figure 1).

The fluency module

Writing fluency can be defined as the nature or pattern of the use of a language with regard to correct choice of vocabulary, sentence structure variety, syntactic complexity, coherence, etc.4. Fluency assessment models capture ideas being conveyed without stuttering or being awkward, sounding approximate to the nativist type of communication trends. Traditional fluency measurement is based on scales, highlighting inter-rater reliability concerns. The proposed system overcomes this drawback as it includes BERT based neural network to induce semantic coherence and linguistic naturalness automatically through deep situational understanding. This architectural decision was made considering the strengths of BERT, which has been found to be effective regarding the identification of contextual relationships and semantic patterns that are necessary in fluency measurement. Input sequences are fed into the system and passed through 12 transformer layers with multi-head self-attention to identify long-range dependencies and contextual relationships (see Figure 2).

The mechanism of attention would be as follows:

Mathematical formula for attention mechanism in neural networks, showing softmax operation.   (5)

Let Q, K, and V be the matrices that represent query, key, and value, respectively, and d and k be the dimensions of the key vectors. The CLS token embedding is the summary one, which records the overall semantic coherence of the whole input sequence.

The calculation of the fluency score is as follows:

Sigmoid function formula, Sf=sigmoid(WreghCLS+breg), mathematical expression.  (6)

In which hCLS is the BERT [CLS] token embedding, and Wreg, are breg the parameters of the regression head, and σ is the sigmoid activation function that normalizes the output to [0, 1].

The correctness module

Correctness assessment focuses on grammatical accuracy, mechanical exactness, and adherence to standard English conventions. This dimension addresses the technical aspects of writing, including syntax, morphology, punctuation, and spelling accuracy. The correctness component not only evaluates the number of linguistic errors but also examines the impact on the overall quality of the text. Unlike fluency assessment, which emphasizes semantic coherence and natural flow, the correctness evaluation module requires precise error identification and systematic correction. The module distinguishes between different error types, assesses the severity, and provides targeted corrections to support learning improvement. The correctness loss uses the FLAN-T5 model, which has been fine-tuned for grammatical error spotting and correction. This choice is informed by the text-to-text transformability of T5 that allows the system to not only fault find but also render relevant corrections within one system. The system is an encoder-decoder form, and the text is fed in the form of this transformation: (see Figure 3)

T5 text correction formula, T5_Decoder(T5_Encoder), coding method.   (7)

The encoder element generates context-sensitive representations that not only capture the grammatical tendencies but also potential areas of failure:

Neural network encoding equation: H_enc = T5Encoder(T_input) formula.   (8)

With the help of generating appropriate grammatical variations to falsely identified mistakes, the decoder generates corrected sequences:

Mathematical expression, T_corrected = T5Decoder(H_enc), neural network model, text decoding.   (9)

This kind of two-way processing enables the system to have the awareness of the contexts of errors, and how the contexts should be rectified, so the hints are semantically coherent, yet focus on any correction in grammar.

Quantification of error and scoring algorithm

The score of accuracy compared the original writing to the correct version of writing, considering the linguistic errors and the severity based on the position within the text and the interference with the meaning. This method captures the distinction between types of errors in comparison to the impact on the textual understanding of the text. The relationship between frequency of errors and poor quality of text was stressed in a logarithmic manner in the grading system. A fundamental step in token comparison of the original text and corrected text is two degrees of comparison that are founded on the bit-wise technique of string alignment. The second stage involves the identification and classification of types of errors based on some known linguistic taxonomies that differentiate between grammatical errors (subject-verb agreement, tense consistency, and syntactic structure reliability), mechanical (capitalization, punctuation, and spelling), and usage (word choice, idiomatic expression, and register appropriateness) errors. The third stage is the calculation of the error ratio based on the total length of the text. The fourth step exploits the exponential decay scoring algorithm that transforms the ratios of errors into directly interpretable correctness scores to approximate the non- additive and non- linear dependence between frequency of mistakes and text quality.

The mathematical treatment of the error quantification process begins with the calculation of the ratio of errors installed, which provides a normalized error installation rate, which in turn permits a fair comparison to be made of texts of different lengths:

Error ratio formula, Error_Ratio=Number_of_Corrections/Total_Tokens, mathematical equation.   (10)

Decay constant formula \(S_c = 100e^{-2.5 \cdot \text{Error Ratio}}\); mathematical equation.   (11)

It was established calibration parameter of 2.5 on an empirical basis by fully validating this with human expert judgments, so that the scoring function can offer results in line with human understanding in relation to the decrease in text quality as the frequency of errors increases. Setting this parameter value in this way ensures that any text with a low error rate (Error_Ratio < 0.1) has a high score of correctness as it closely follows the rules of standard English, any text with a moderate error rate (0.1 < Error_Ratio ≤ 0.3) has an intermediate score of correctness indicating that there are some linguistic concerns which are nonetheless not outright disastrous, and any text with a high error rate (Error_Ratio > 0.3) has gotten a low score of correctness because there are critical linguistic concerns which significantly affect the possibility of comprehension.

The correctness scoring algorithm for the correctness evaluation systematically takes into account all errors located and fixed by the T5-based system as penalty items in the calculation of the final error ratio. The penalty credit mechanism used for grammatical errors is represented by the exponential decrease of the with error proportion growth towards high values, which means that the quality of the text is very poor, and becomes 100 when the error ratio equals 0, which means there are absolutely no errors detected. It is a perfect usage of English standard conventions. Such a mathematical relation ensures that the code of correctness offers significant distinction within the whole spectrum of linguistic proficiency levels and is responsive to small successive advances, which indicate real acquisitions.

Datasets: training and evaluation corpus

Training data of sufficient quality and scope is paramount to the performance of the dual-model system. The current study utilized a well-designed dataset of student-written texts to develop and assess the model, which was produced by employing different writing tasks. The sample comprised 45 males and 45 female Pakistani university students with an average age of 19 years, studying ‘Functional English’ as a compulsory course. Each student wrote eight essays for eight weeks, including pre-test, intervention essays, and post-test occasions. Students were allowed to produce 400–450 words for each writing task. The statistics of the data sets provide the following characteristics:

70,500 words

Mean sentence length: 15.7 words (SD = 3.2)

The range of vocabulary (Type-Token Ratio): 0.64 (SD 0.08) Grammatical error density: 2.3 in 100 words (SD = 1.1)

The justification and data quality assurance of the selected data

The dataset that was chosen was driven by a few important considerations that allowed for the richness and relevance to research work on automated writing evaluation. To begin with, the Economics student population was selected due to the nature, similar to EFL learners in the Pakistani higher education, thus enabling presentation of more authentic writing samples that reflect real-world situations. Second, the establishment of the maximum and minimum potential word range, 400–450 words, was informed by pilot studies data that this range is linguistically complex enough to be reliably analyzed by a machine, and it remained at the same time of manageable size in terms of consistent human assessments. Each of the compositions was evaluated by three trained English teachers (inter-rater reliability κ = 0.847, fluency dimension, and 0.823, correctness dimension) on standardized 10-point fluency and correctness rating scales. Training of the human evaluators was rigorous and depended on developed scoring rubrics with respect to IELTS and TOEFL writing assessments. The data consists of a human-annotated dataset, which can be used in training and validating models to be trained on human-annotated data.

Procedures of data preprocessing and validation

The dataset was pre-processed systematically and involved text normalization, tokenizing the text with the BERT WordPiece tokenizer, and making quality validation checks. Those who have submitted incomplete work and have produced plagiarized text were not included to determine data integrity. The last data was randomly divided into training (70%, n = 189), validation (15%, n = 41), and test sets (15%, n = 40) through stratified sampling to make it balanced across proficiency levels and the type of tasks. Linguistic analysis of a large reference corpus containing professionally edited academic texts, articles from newspapers, and published essays, making approximately 50 million words. This corpus provides language patterns needed to perform the fluency tests and assists in the procedures of error detection by comparing these with the standard usage. In the reference corpus, the steps before preprocessing and indexing are tokenization, part-of-speech tagging, syntax parsing, and semantic labeling to facilitate efficient access during real-time assessment.

Fluency model training strategy

A sequential optimization system is proposed to train the data to achieve fluency. The training used state-of-the-art features, including adaptive learning rate scheduling, progressive batch size, and advanced regularization methods that prevent overfitting as training proceeds, thereby maintaining the model's effectiveness in identifying linguistic patterns indicative of language fluency. The learning rate is 5 × 10-4 with a linear decay schedule, configured to ensure convergence remains stable and does not oscillate around the optimal parameter values. This learning rate is chosen through grid search in [1 x 10-6, 1 x 10-4] with 5 x 10-5 being the most appropriate tradeoff between the convergence speed and the stability. The real training will be restricted to only 3 epochs, a configuration that has been optimized on the basis of empirical benefits by the validation loss tracking to prevent overfitting without preventing sufficient parameter updates to render learning effective. Early halting of mechanisms with a patience of 2 epochs and a minimum delta of 0.001 has been done to prevent degradation.

Optimization is carried out by AdamW optimizer, with simplified choices like β₁ = 0.9 (momentum, which controls the decay of first moment estimates exponentially), β₂ = 0.999 (second-moment estimation, which controls the decay of second moment estimates exponentially), and ε = 1 × 10⁻8 (numerical stability term). Weight decay regularization (λ = 0.01) prevents parameter magnitudes from growing excessively large, with this value selected through validation performance analysis across the range [0.001, 0.1]. Complex monitoring is able to monitor various metrics across training to guarantee the best performance of a model and avoid overfitting. Included in the monitoring framework are:

Loss tracking: Training and validation loss are tracked within every epoch, and the early stopping happens automatically when validation loss no longer improves during 2 consecutive epochs.

Performance metrics: The validation data are used in calculating Pearson correlation coefficients between predicted and actual scores every 100 training steps.

Gradient sensing: Gradient norms are monitored to detect vanishing and exploding gradients, and at a gradient norm of 1.0, gradient clipping is applied.

Learning rate scheduling: Validation-based learning rate decrease (factor = 0.5) in case of more than 1 epoch plateau is detected.

Regularization-Related: Dropout rates (0.1 for BERT’s, 0.3 for regression head) were found by systematic ablation. Several regularization methods based on preventing overfitting are used during the training process: (1) dropout regularization using optimized dropout rates per type of layer in the network, (2) weight decay as explained above, (3) early stopping, determined by the validation performance, and (4) data augmentation based on paraphrasing 15 percent of the training examples using back-translation methods. Checkpoints of the best-performing model were saved after every epoch, and the highest-scoring models were selected based on validation correlation scores. The loss function used in the fluency assessment part is the Mean Squared Error (MSE), which is the main objective function of the regression task of predicting the continuous fluency scores. The loss formulation mathematically is given as:

Fluency loss function equation, L_fluency, analysis of predicted vs. true values, mathematical formula.  (12)

N is the total training sample size, y_pred, i being the predicted fluency score of the i-th sample, and y_true, i being the corresponding ground truth score as was given by human expert evaluators. This loss is very useful in discouraging prediction-actual error quadratically so that bigger errors attract a larger penalty, and is differentiable as well. The learning rate scheduling mechanism is highly advanced with a two-phase process, and it starts with a linear warmup stage, followed by a systematic decay process in order to create optimal convergence features. This is done in the warmup stage, whereby the learning rate starts at zero and gradually changes to the maximum rate, after which the model parameters are optimized relative to the training dataset. The mathematical expression of the warm-up phase is:

Learning rate schedule equation, lr(t)=lrmax·min(t/warmup_steps,1.0), mathematical formula.   (13)

After the warmup phase, the learning rate is decayed in a systematic linear manner to avoid overshooting of the optimum parameter values and results in steady convergence. Mathematically, the decay phase reads as:

Learning rate schedule equation, lr(t)=lr_max*(1-(t-warmup_steps)/(total_steps-warmup_steps)).   (14)

In which t is the current training step, lr max is the maximum learning rate achieved during the warmup, warmup is the number of steps assigned to the warmup phase, and total is the total number of training steps over all epochs. The mentioned scheduling procedure ensures aggressive learning of the model at the lower levels and stability on the convergence levels.

Correctness model training strategy

The correctness model was based on the transfer learning strategy using the pre-trained FLAN-T5 model to take advantage of all the available grammar knowledge as required. This was meant to look into the general language understanding as well as the specific information regarding the mistakes committed by the ESL learners. The transfer learning method used calls upon the pszemraj/flan-t5-large-grammar-synthesis checkpoint as the base model, with several noteworthy advantages in terms of correctness. As the model is already trained and possesses a high language understanding capacity that has been trained in various fields of text, it will adapt to many writing styles and writing topics. Particularly, grammar-specific fine-tuning has been conducted on large correction datasets spanning multiple types of grammatical errors as would be encountered in second language writing. Moreover, the checkpoint comprises powerful error-detecting systems that are tested against the different types of errors (i.e., morphological errors, syntactic, and semantic errors), creating a perfect standard of comparison to the innate correction testing parameters of the study.

The domain adaptation process reflects the systematic parameter modifications that do not alter the general capabilities of language understanding of the pre-trained model but introduce the specific features of the writing assessment task solution. This process of adaptation has the mathematical derivation as:

Domain adaptation formula θ_final=θ_pretrained+Δθ_domain; machine learning model update.   (15)

Here θpretrained represents the parameters of the pre-trained model encoding general language understanding and grammatical knowledge, and Δθdomain represents the specific parameter updates learned from the target writing assessment task. The domain-specific updates are computed by gradient-based optimization on the target dataset, ensuring the model develops task-specific sensitivity to error types common in EFL writing without disrupting its broader linguistic capabilities.

Experiments with comparisons

To prove the functionality of the EG, the Pearson correlation coefficient is proposed as the key value of measurement, which determines the linear connection between automated scores and the expertise of humans. The correlation coefficient is found by the formula:

Correlation coefficient formula, r, equation, statistical analysis, data correlation calculation.    (16)

Where xi represents the automated scores, represents the human ratings, and static equilibrium calculation, equation ΣFx=0, ΣFy=0, structural diagram and Static equilibrium; ΣFx=0, ΣFy=0 diagram with force vectors; mechanics analysis tool are the respective means. The correlation coefficient ranges from −1 to 1, with values near 1 indicating a strong positive relationship between automated and human assessment.

Mean absolute error formula, MAE=1/N Σ|y_pred,y_true|; statistical analysis tool.    (17)

Root Mean Square Error formula, RMSE equation for statistical error analysis, predictive accuracy.    (18)

R-squared formula: R² = 1 - SSres/SStot; statistical analysis, data fitting equation.    (19)

Baseline model comparisons

To evaluate the performance of the EG and demonstrate its superiority to existing methods, in-depth comparisons are made with established AWE and traditional single-metric approaches. These baseline comparisons are essential for validating the added value of the EG architecture and demonstrating that the integration of fluency and correctness assessment provides measurable improvements over approaches that focus on individual dimensions in isolation. The selection of baseline models covers four categories of AWE, each reflecting a distinct orientation toward how writing should be assessed. The first one has a single BERT-based model to evaluate fluency. These models use transformer architectures to assess the textual patterns of smoothness and coherence. The second category, embedded in traditional linguistic methodologies, focuses on rule-based systems for grammatical error detection. Hence, the focus remained on correctness, overlooking other aspects related to writing quality, such as cohesion and content. The third category is feature-based AWE systems, which includes indicators of linguistic complexity—like variational sentence length, diversification of lexis, and syntactic complexity to generate overall quality scores. The fourth one takes hybrid systems into account by combining various surface-level linguistic features, often using ensemble methods or weighted averages. These models do not reach the level of integrated sophistication achieved by EG neural architectures. Baseline model performance is evaluated with three major dimensions. The first one measures alignment of system-generated grades with trained human experts’ (English instructors) ratings using standardized rubrics. The second determines generalizability, identifying if performance remains consistent across three essays with varying topics and complexity. The third dimension relies on predictive reliability, assessing the accuracy and stability of scoring through statistical tools such as mean absolute error, root mean square error, and confidence interval analysis. Collectively, these dimensions establish a robust framework for comparison and highlight the superiority of EG systems, particularly in achieving greater precision and stability than traditional approaches.

Experimental and control groups

This study comprised 90 undergraduate Economics students who were studying a functional English course in two intact classes. The data was taken at three occasions: pre-test, intervention, and post-test to track progress in writing accuracy and fluency over time. This human subject study was approved by the Departmental Advisory Board, COMSATS University Islamabad, Lahore Campus, Pakistan. Participants were exposed to two conditions: traditional human feedback and computer-supported feedback. The control group consisted of 45 students, who received the traditional written feedback from a trained English instructor. Participants received written feedback about students’ essays in the form of holistic marks, identification of errors, and recommendations in the shape of comments by an instructor on how to improve the work. The experimental group, in its turn, included 45 students who were instructed the same way in the classroom, yet students’ writing was supplemented with the help of rapid automated feedback provided by the EG, which provided them with the diagnostic information, both in terms of fluency and correctness, within about 30 seconds of the submission.

This study was carried out over a duration of 8 weeks, where a pre-test (Week 1) was carried out to determine the initial writing performance of the test subjects before the students were given an intervention. Computer-based feedback with NLP semantic markers in the experimental group and teachers’ evaluation in the control group were given to the participants for six weeks. Ultimately, the post-test (in Week 8) was carried out to understand the difference between pre- and post-test measures of accuracy and fluency scores. To obtain similar results in comparability, respondents were requested to undertake similar tasks in writing at every assessment period, thereby minimizing potential confounding variables of the task difficulty and familiarity of content in the writing. To ensure that the writing prompts are consistent with the level of difficulty, as well as setting content validity, the writing prompts were chosen in a manner that is congruent. Essay topics were chosen on the basis that the students cover various content areas so as to prevent practice effect and boredom on the part of the participants from having to do a similar task over a long duration of time.

During the pre-test phase, the participants were requested to compose an essay of 400–450 words and write about the academic and career goals. This descriptive assignment was based on personal experience and future projections, which implies the need to organize the writing, provide clear examples, and make a logical statement of personal goals and motivation. The intervention writing task was based on different topics, such as writing about the daily routine of life, hobbies, traveling, the first day at the university, memorable days of life, and best friend moments. The post-test prompt was connected to the subject matter of the ‘Effect of technology on communication,’ and the essay length required was 400–450 words.

Access restricted. Please log in or start a trial to view this content.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Validity and reliability of the statistical approach

The system was tested on various complementary statistical methods that all give comprehensive evidence of the effects of the EG on the development of writing. This is a multifaceted analytical method that acknowledges that educational interventions are subject to complex statistical analysis that cannot be reduced to mere comparisons of groups but rather to the analysis of change patterns, the magnitude of effects, and the ...

Access restricted. Please log in or start a trial to view this content.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The experimental results demonstrate several significant achievements. First, the dual-model system achieved a strong correlation with human expert judgments (r = 0.923), substantially outperforming single-model approaches and contemporary AWE systems. Second, the controlled intervention study revealed significant improvements in ESL students' writing performance, with large effect sizes (d = 1.28) exceeding those reported in recent literature. Third, the system demonstrated superior performance across multiple evalu...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

There are no conflicts of interest among all authors.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

We acknowledge the support of COMSATS University Islamabad, Lahore campus, the faculty of English, in helping to collect data from the students.

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
SOFTWARE / LIBRARIES USED IN THE STUDY
Deep Learning FrameworkPyTorch1.13.1Deep learning framework for neural network implementation
Pre-trained ModelsTransformers (Hugging Face)4.21.3Pre-trained model loading, BERT and T5 model implementations
Language ModelBERT (Bidirectional Encoder Representations from Transformers)bert-base-uncasedFluency assessment, semantic coherence evaluation, contextual understanding
Language ModelFLAN-T5 (Fine-tuned Language Net T5)pszemraj/flan-t5-large-grammar-synthesisGrammatical error correction, text-to-text transformation, error detection
Numerical ComputingNumPy1.23.5Numerical computations, array operations for mathematical functions
Data AnalysisPandas1.5.2Data manipulation, statistical analysis and preprocessing
Machine LearningScikit-learn1.1.3Statistical metrics calculation, correlation analysis, evaluation metrics
VisualizationMatplotlib3.6.2Data visualization, training progress plots, performance charts
VisualizationSeaborn0.12.1Statistical data visualization, correlation heatmaps
NLP ToolkitNLTK (Natural Language Toolkit)3.7Text preprocessing, tokenization, linguistic analysis
NLP ProcessingSpaCy3.4.4Advanced NLP preprocessing, POS tagging, syntactic analysis
TokenizationTokenizers0.13.2Fast tokenization for BERT WordPiece and T5 tokenization
STATISTICAL AND ANALYSIS SOFTWARE
Statistical SoftwareSPSS Statistical SoftwareVersion 26.0Statistical analysis, independent samples t-tests, ANOVA, correlation analysis
Training DatasetKaggle ASAP Dataset2012 versionTraining corpus reference (90% of AWE models use this dataset)
PYTHON OPTIMIZATION AND TRAINING LIBRARIES
Optimizertorch.optim.AdamWPyTorch 1.13.1Model optimization with weight decay regularization (λ=0.01), β1=0.9, β2=0.999, ε=1×10-8
Loss / Activationtorch.nn.functionalPyTorch 1.13.1Sigmoid activation, loss functions, dropout regularization
Data Loadingtorch.utils.data.DataLoaderPyTorch 1.13.1Progressive batch sizing (4→16), data loading and batching
Tokenizationtransformers.AutoTokenizerTransformers 4.21.3BERT WordPiece tokenization, text preprocessing
Model Loadingtransformers.AutoModelTransformers 4.21.3Pre-trained model loading for BERT and T5 architectures
Gradient Controltorch.nn.utils.clip_grad_norm_PyTorch 1.13.1Gradient clipping (threshold: 1.0) for training stability
LR Schedulingtorch.optim.lr_schedulerPyTorch 1.13.1Learning rate scheduling with linear decay and warmup
Metricssklearn.metricsScikit-learn 1.1.3Pearson correlation, MAE, RMSE calculation for evaluation
PYTHON NEURAL NETWORK IMPLEMENTATION
Base Classtorch.nn.ModulePyTorch 1.13.1Base class for custom neural network architectures (Fluency & Correctness modules)
Linear Layerstorch.nn.LinearPyTorch 1.13.1Linear regression head implementation (768→512→256→128→1 neurons)
Regularizationtorch.nn.DropoutPyTorch 1.13.1Dropout regularization (0.1 for BERT, 0.3 for regression head)
Activationtorch.nn.functional.sigmoidPyTorch 1.13.1Sigmoid activation for fluency score normalization [0,1]
Activationtorch.nn.functional.reluPyTorch 1.13.1ReLU activation in regression layers for non-linear transformations
Transformertransformers.BertModelTransformers 4.21.312 transformer layers with multi-head self-attention for fluency assessment
Seq2Seqtransformers.T5For
ConditionalGeneration
Transformers 4.21.3Encoder-decoder architecture for grammatical error correction
Loss Functiontorch.nn.MSELossPyTorch 1.13.1Mean Squared Error loss function for fluency regression training
PYTHON EVALUATION AND METRICS LIBRARIES
Correlationscipy.stats.pearsonrSciPy 1.9.3Pearson correlation coefficient for model-human agreement
Error Metricsklearn.metrics.mean_absolute_errorScikit-learn 1.1.3Mean Absolute Error (MAE) calculation for prediction accuracy
Error Metricsklearn.metrics.mean_squared_errorScikit-learn 1.1.3Root Mean Square Error (RMSE) for error magnitude assessment
Statistical Testscipy.stats.ttest_indSciPy 1.9.3Independent samples t-test for statistical significance testing
Correlation Matrixnumpy.corrcoefNumPy 1.23.5Correlation matrix calculation for inter-rater reliability
Data Correlationpandas.DataFrame.corrPandas 1.5.2Data correlation analysis and statistical reporting
Visualizationmatplotlib.pyplotMatplotlib 3.6.2Performance visualization, correlation plots, training progress charts
Heatmapseaborn.heatmapSeaborn 0.12.1Correlation matrix visualization and statistical data presentation
PYTHON TEXT PROCESSING AND NLP LIBRARIES
Text Processingre (Regular Expressions)Python 3.9+ built-inText pattern matching, data cleaning and preprocessing
Config HandlingjsonPython 3.9+ built-inConfiguration file handling, model parameter storage
SerializationpicklePython 3.9+ built-inModel serialization and checkpoint saving/loading
Progress Trackingtqdm4.64.1Progress bars for training loops and data processing
LoggingloggingPython 3.9+ built-inTraining progress logging, error tracking and debugging
ReproducibilityrandomPython 3.9+ built-inRandom seed setting for reproducible experiments
File SystemosPython 3.9+ built-inFile system operations, model path management
CLI ParsingargparsePython 3.9+ built-inCommand-line argument parsing for training configurations
COMMERCIAL AWE / AES PLATFORMS (Referenced)
Commercial PlatformPigai.orgCommercial AWE platformChinese EFL writing assessment, automated feedback systems
Commercial PlatformBingo EnglishCommercial AWE systemEnglish language learning support, writing skill development
Educational PlatformMy AccessEducational writing platformStudent writing assessment and feedback delivery
Scoring EngineE-raterETS automated scoring engineStandardized test essay scoring, holistic evaluation
Assessment ToolI-writeWriting assessment toolAcademic writing evaluation and diagnostic feedback
Practice ServiceCriterionETS writing practice serviceWriting skill development and automated feedback
AI Language ModelChatGPTOpenAI language modelAI-assisted writing feedback and evaluation (referenced in literature)
DEEP LEARNING ARCHITECTURES (Referenced in Literature)
ArchitectureRecurrent Neural Networks (RNN)Cai, 2019Sequential text processing — writing feedback systems and automated evaluation
ArchitectureLong Short-term Memory (LSTM)Jin et al., 2018Long-range dependency modeling — automated essay scoring and sequence analysis
ArchitectureConvolutional Neural Networks (CNN)Dong et al., 2017Local pattern recognition — text classification and feature extraction for scoring
ArchitectureTransformer-LSTM HybridXuan, 2025Combined contextual and sequential processing — automatic scoring and feedback generation
ArchitectureVariational Autoencoders (VAEs)Kumar et al., 2024Generative modeling — enhanced writing skill development and personalized feedback
ArchitectureMulti-Agent SystemsThompson et al., 2024Collaborative AI processing — advanced writing assistance (AcademiCraft platform)
MechanismAttention MechanismsDong et al., 2017Self-attention and multi-head attention — improved AES performance and contextual understanding
TRADITIONAL MACHINE LEARNING TECHNIQUES (Referenced)
Statistical MethodBayes' TheoremRudner & Liang, 2002Traditional ML approach — early AES/AWE development and probabilistic scoring
Statistical MethodLinear RegressionPhandi et al., 2015Statistical modeling — feature-based writing assessment and score prediction
Learning MethodRank Preference LearningChen & He, 2013Ranking-based methodology — comparative essay scoring and preference modeling
Language ModelN-gram ModelsXie et al., 2015Statistical language modeling — grammatical error correction and pattern recognition
ArchitectureEncoder-Decoder ArchitectureGe et al., 2018Sequence-to-sequence modeling — grammatical error correction and text transformation
EVALUATION STANDARDS AND FRAMEWORKS
Assessment StandardIELTS Writing AssessmentScoring rubric adaptationStandardized evaluation criteria for fluency and correctness
Assessment StandardTOEFL Writing AssessmentAcademic writing standardsProficiency evaluation and standardized scoring
Statistical MeasureCohen's d Effect Size0.2=small, 0.5=medium, 0.8=largePractical significance assessment for intervention effectiveness
Reliability MeasureInter-rater Reliability (κ)Fluency: 0.847 / Correctness: 0.823Kappa coefficient — human evaluator consistency and agreement validation

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Bitchener, J. Evidence in support of written corrective feedback. J Second Lang Writ. 17 (2), 102-118 (2008).
  2. Nusrat, A., Ashraf, F., Narcy-Combes, M. F. Effect of direct and indirect teacher feedback on accuracy of English writing: A quasi-experimental study among Pakistani undergraduate students. 3L Lang Linguist Lit. 25 (4), 84-98 (2019).
  3. Nusrat, A., Khan, S., Kashif, F., Fawad, R. Effect of teachers’ asynchronous e-feedback and synchronous oral feedback on English language learners’ writing accuracy. Front Educ. 7, 783684 (2022).
  4. Chen, B., Bao, L., Zhang, R., Zhang, J., Liu, F., Wang, S., et al. A multi-strategy computer-assisted EFL writing learning system with deep learning incorporated and its effects on learning: A writing feedback perspective. J Educ Comput Res. 61 (8), 60-102 (2024).
  5. Li, J., Link, S., Hegelheimer, V. Rethinking the role of automated writing evaluation (AWE) feedback in ESL writing instruction. J Second Lang Writ. 27, 1-18 (2015).
  6. Ellis, R. Explicit form-focused instruction and second language acquisition. The handbook of educational linguistics. , 437-455 (2008).
  7. Shamim, F. Trends, issues and challenges in English language education in Pakistan. Asia Pac J Educ. 28 (3), 235-249 (2008).
  8. Haider, G. An insight into difficulties faced by Pakistani student writers: Implications for teaching of writing. J Educ Soc Res. 2 (3), 17-27 (2012).
  9. Nawaz, M., Hussain, S. A., Bughio, F. A. Exploring the preferred corrective feedback and practiced corrective feedback among Pakistani ESL secondary school students and teachers in writing class: Matches and mismatches. Int J Lang Lit Transl. 6 (1), 31-45 (2023).
  10. Rudner, L. M., Liang, T. Automated essay scoring using Bayes theorem. J Technol Learn Assess. 1 (2), 1-22 (2002).
  11. Phandi, P., Chai, K. M. A., Ng, H. T. Flexible domain adaptation for automated essay scoring using correlated linear regression. , 431-439 (2015).
  12. Cai, C. Automatic essay scoring with recurrent neural network. , 1-7 (2019).
  13. Jin, C., He, B., Hui, K., Sun, L. TDNN: A two-stage deep neural network for prompt-independent automated essay scoring. 1, 1088-1097 (2018).
  14. Dong, F., Zhang, Y., Yang, J. Attention-based recurrent convolutional neural network for automatic essay scoring. CoNLL. , 153-162 (2017).
  15. Sharma, A., Kabra, A., Kapoor, R. Feature enhanced capsule networks for robust automatic essay scoring. , 365-380 (2021).
  16. Mim, F. S., Inoue, N., Reisert, P., Ouchi, H., Inui, K. Corruption is not all bad: Incorporating discourse structure into pre-training via corruption for essay scoring. IEEE ACM Trans Audio Speech Lang Process. 29, 2202-2215 (2021).
  17. Cummins, R., Rei, M. Neural multi-task learning in automated assessment. arXiv. , 1-9 (2018).
  18. Cao, Y., Jin, H., Wan, X., Yu, Z. Domain-adaptive neural automated essay scoring. , 1011-1020 (2020).
  19. Jiang, Z., Liu, M., Yin, Y., Yu, H., Cheng, Z., Gu, Q., et al. Learning from graph propagation via ordinal distillation for one-shot automated essay scoring. , 2347-2356 (2021).
  20. Li, X., Chen, M., Nie, J. Y. SEDNN: Shared and enhanced deep neural network model for cross-prompt automated essay scoring. Knowl Based Syst. 210, 106491 (2020).
  21. Beseiso, M., Alzubi, O. A., Rashaideh, H. A novel automated essay scoring approach for reliable higher educational assessments. J Comput High Educ. 33 (3), 727-746 (2021).
  22. Xie, W., Huang, P., Zhang, X., Hong, K., Huang, Q., Chen, B., et al. Chinese spelling check system based on an n-gram model. , 128-136 (2015).
  23. Brockett, C., Dolan, W. B., Gamon, M. Correcting ESL errors using phrasal SMT techniques. , 249 (2006).
  24. Zhang, S., Huang, H., Liu, J., Li, H. Spelling error correction with soft-masked BERT. , 882-890 (2020).
  25. Ge, T., Wei, F., Zhou, M. Reaching human-level performance in automatic grammatical error correction: An empirical study. arXiv. , 1-15 (2018).
  26. Zhao, W., Wang, L., Shen, K., Jia, R., Liu, J. Improving grammatical error correction via pre-training a copy-augmented architecture with unlabeled data. 1, 156-165 (2019).
  27. Shermis, M. D., Koch, C. M., Page, E. B., Keith, T. Z., Harrington, S. Trait ratings for automated essay grading. Educ Psychol Meas. 62 (1), 5-18 (2002).
  28. Mizumoto, A., Shintani, N., Sasaki, M., Teng, M. F. Testing the viability of ChatGPT as a companion in L2 writing accuracy assessment. Res Methods Appl Linguist. 3 (2), 100116 (2024).
  29. Ajabshir, Z. F., Ebadi, S. The effects of automatic writing evaluation and teacher-focused feedback on CALF measures and overall quality of L2 writing across different genres. Asian Pac J Second Foreign Lang Educ. 8 (1), 26 (2023).
  30. Almusharraf, N., Alotaibi, H. An error-analysis study from an EFL writing context: Human and automated essay scoring approaches. Tech Knowl Learn. 28 (3), 1015-1031 (2023).
  31. Chen, H., Pan, J. Computer or human: A comparative study of automated evaluation scoring and instructors' feedback on Chinese college students' English writing. Asian Pac J Second Foreign Lang Educ. 7 (1), 34 (2022).
  32. Liu, X., Zhang, Y., Chen, L. The transformative impact of AI-powered tools on academic writing: Perspectives of EFL university students. Lang Learn Technol. 28 (2), 78-95 (2024).
  33. Zhang, Q., Li, F. From process to product: Writing engagement and performance of EFL learners under computer-generated feedback instruction. System. 115, 102998 (2023).
  34. Kumar, S., Patel, R., Williams, T. Automated deep learning approaches in variational autoencoders (VAEs) for enhancing English writing skills. J Educ Technol Res. 45 (3), 234-251 (2024).
  35. Martinez, A., Rodriguez, P., Thompson, K. Neural automated writing evaluation with corrective feedback. Comput Linguist. 50 (1), 123-145 (2024).
  36. Thompson, J., Anderson, C., Davis, R. AcademiCraft: Transforming writing assistance for English for academic purposes with multi-agent system innovations. Artif Intell Educ. 12 (4), 445-462 (2024).
  37. Rodriguez, M., Kim, S., Brown, D. Technology-enhanced language learning in English language education: Performance analysis, core publications, and emerging trends. Appl Linguist Rev. 15 (2), 201-225 (2024).
  38. Xuan, Y. Transformer-LSTM models for automatic scoring and feedback in English writing assessment. IEEE Access. 13, 82084-82096 (2025).
  39. Anderson, R., Smith, K., Johnson, M. The impact of AI writing tools on the content and organization of students' writing: EFL teachers' perspective. Comput Educ. 195, 104712 (2023).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Computer Assisted Language LearningNeural Network ModelsNatural Language ProcessingAutomated Essay ScoringESL Writing PedagogyGrammar FeedbackWriting Fluency

Related Articles