Research Article

Readability and Quality of Large Language Model Patient Education for Trigeminal Neuralgia: A Cross-Sectional Study

DOI:

10.3791/70833

August 21st, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Here, we present a protocol to systematically evaluate the readability, educational suitability, and quality of large language model–generated patient information on trigeminal neuralgia to benchmark models and guide patient-facing communication.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Large language models (LLMs) are increasingly used for patient health education, yet the readability and educational quality of LLM-generated information on trigeminal neuralgia (TN) have been insufficiently evaluated. This cross-sectional benchmarking study compared TN educational content generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Twenty frequently asked TN questions covering basic disease knowledge, etiology/risk factors, diagnosis, treatment, and prevention/rehabilitation were presented to each model using standardized prompts. Readability was assessed using seven established indices, educational suitability using the Patient Education Materials Assessment Tool for Understandability and Actionability (PEMAT), and overall information quality using the Global Quality Score (GQS). Two clinical experts independently evaluated all responses, with disagreements resolved by a senior adjudicator. Statistical analyses compared model performance, thematic differences, and correlations among the evaluation metrics. Significant differences were observed among the models for readability, PEMAT, and GQS scores. GPT-5 generated the most linguistically complex responses but achieved the highest ratings for educational suitability and information quality. In contrast, Wenxin Yiyan produced the most readable text but generally scored lower on PEMAT and GQS. Content category influenced readability, with prevention/rehabilitation and etiology/risk-factor topics being more difficult to read, whereas PEMAT and GQS remained relatively consistent across themes. Readability indices showed strong internal consistency and weak-to-moderate positive correlations with PEMAT and GQS, while PEMAT and GQS demonstrated a moderate positive correlation. These findings suggest that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability. Because patient comprehension, satisfaction, trust, health outcomes, factual accuracy, and clinical safety were not evaluated, these results should be interpreted as an expert-rated benchmarking analysis rather than evidence of clinical readiness.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Trigeminal neuralgia (TN) is a neuropathic pain syndrome characterized by recurrent episodes of unilateral, brief, and extremely severe facial pain commonly described as electric shock-like or stabbing. According to the International Classification of Headache Disorders, 3rd edition (ICHD-3), pain is typically localized to the distribution of one or more branches of the trigeminal nerve, is often intolerable in intensity, and can be triggered by innocuous stimuli such as light touch, facial washing, tooth brushing, speaking, or chewing. Attacks are marked by sudden onset and abrupt cessation, with durations ranging from fractions of a second to several minutes1,2. Although TN is not usually life-threatening, its severe and unpredictable pain substantially disrupts essential daily activities, including eating, speech, sleep, and emotional regulation, resulting in a considerable psychosocial burden and impaired quality of life. Evidence from systematic reviews indicates that individuals with TN face a significantly elevated risk of depression, anxiety, and sleep disturbances compared to the general population3,4. In large cohort studies, approximately 35–55% of patients with TN exhibit clinically meaningful symptoms of anxiety or depression, and pain catastrophizing is highly prevalent, suggesting a bidirectional interaction between chronic severe pain, insomnia, and affective disorders5,6. Epidemiological estimates indicate that the prevalence of TN in the general population ranges from approximately 0.03% to 0.3%, with a higher occurrence among women and older adults and notable variations across geographic regions, ethnic groups, and age strata7,8. In highly populated countries, such as China and India, rapid population aging has been accompanied by a sustained increase in the number of individuals affected by TN, thereby imposing a growing burden on healthcare systems and underscoring the need for more effective, scalable, and data-informed approaches to disease assessment and management8,9.

TN can be classified into classical, secondary, and idiopathic subtypes, with increasing evidence demonstrating substantial differences in the etiological mechanisms and therapeutic strategies among these categories1,10. Current studies suggest that neurovascular conflict-related structural changes at the trigeminal nerve root entry zone, peripheral nerve hyperexcitability, and altered central pain modulation jointly contribute to the pathogenesis of the disease11,12. Clinical guidelines recommend carbamazepine and oxcarbazepine as first-line treatments, whereas surgical or interventional approaches are considered when pharmacological therapy is ineffective or poorly tolerated10. Despite these therapeutic advances, TN is characterized by a relapsing–remitting course, and long-term pain recurrence remains common, making permanent remission difficult to achieve13. Consequently, long-term follow-up and structured chronic disease management are essential for monitoring the risk of recurrence, treatment response, and complications. Longitudinal cohort studies indicate that under standardized management strategies (including pharmacological treatment, surgical intervention when indicated, and comprehensive follow-up), approximately half of patients receiving conservative management can achieve a ≥50% reduction in overall pain burden within two years, without inevitable disease progression14. These findings underscore the importance of sustained patient engagement and effective health education for long-term disease management. Evidence suggests that patients with a better understanding of disease etiology, pathophysiology, and treatment options are more likely to actively participate in and adhere to therapeutic regimens, leading to improved outcomes15. However, traditional health education approaches are often limited in reach and scalability. With the widespread adoption of the internet, particularly online question-and-answer platforms and social media, patients increasingly rely on digital sources to obtain medical information16,17. Nevertheless, substantial variability in individual health literacy and the ability to access, process, and understand basic health information for informed decision-making pose major barriers to effective patient education18. Moreover, the uneven quality of online health information, including inaccurate or misleading content, may negatively influence patients’ medical decisions and psychological well-being19.

Artificial intelligence (AI) technologies, particularly the rapid development of large language models (LLMs) in recent years, have created new opportunities for medical science communication and health education20. Trained on large-scale textual corpora, these models can generate structured and logically coherent responses through natural language interaction21. Representative international systems, such as ChatGPT, have demonstrated promising potential for medical question answering, patient consultation, and medical education22,23. Several functionally similar LLMs have emerged in China, with continuously expanding user coverage and application scenarios24. Despite these advances, the application of LLMs in medical science popularization still faces substantial challenges, including the reliability of training data, update frequency, the scientific accuracy of generated responses, and a lack of clinical validation25. In addition, differences among models in language style, readability, logical structure, and citation accuracy may directly influence patients’ comprehension and trust in health-related information26. To date, systematic evaluations of LLM performance in TN-related health education have been limited.

Therefore, this study aimed to systematically evaluate and compare the performance of four widely used Chinese LLMs (Doubao, DeepSeek, Wenxin Yiyan, and Tongyi Qianwen) and a representative international LLM (ChatGPT) in answering patient education questions related to TN. Using a cross-sectional design, we constructed a TN question set based on clinical guidelines and common patient concerns, and assessed the responses generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Readability was evaluated using multiple metrics, and understandability and actionability were assessed using the Patient Education Materials Assessment Tool (PEMAT). Overall information quality was evaluated using the Global Quality Score (GQS). In addition, we analyzed how different health education topics influence these evaluation metrics and their interrelationships, aiming to provide evidence-based guidance for selecting and optimizing LLMs in clinical health communication.

Protocol

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study analyzed only AI-generated text and did not involve human participants, animals, biological specimens, medical records, identifiable personal information, or intervention in clinical care. Therefore, ethics review and formal informed consent were not applicable.

Study design

This cross-sectional study evaluated the readability, quality, and educational suitability of responses generated by five LLMs in response to frequently asked questions about TN.

Identification of commonly asked TN-related questions

On November 25, 2025, two clinical experts with extensive experience in pain medicine independently reviewed the current clinical guidelines for trigeminal neuralgia (TN), including the International Classification of Headache Disorders, 3rd edition (ICHD-3), the European Academy of Neurology guideline, and the American Academy of Neurology/European Federation of Neurological Societies guideline1,10,11. Following a consensus discussion, they compiled a list of 20 TN-related questions commonly encountered in clinical practice. The number of questions was determined a priori to provide balanced coverage of the five predefined thematic domains, with four questions selected for each domain, while keeping the total number of model-generated responses within a range feasible for manual expert rating and verification. To improve reproducibility, the selection process followed a predefined domain-based framework: the experts first identified candidate questions within each domain, independently reviewed them for clinical relevance, patient-oriented wording, and avoidance of redundancy, and then reached consensus on the final four questions per domain. The final list of questions is provided in Table 1 and Supplementary File 1. The five predefined thematic domains were basic disease knowledge, etiology and risk factors, diagnosis, treatment, and prevention and rehabilitation. Because the question set was not derived from patient search logs, online search queries, or formal patient interviews, the potential for selection bias is acknowledged as a study limitation.

AI models and data collection procedure

On November 25, 2025, the finalized set of 20 TN-related questions was submitted to five publicly accessible large language model (LLM) platforms: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5/ChatGPT. All models were accessed through their standard public web-based chat interfaces; application programming interface (API) access was not used. For each platform, the default publicly available model available on the date of data collection was used without modification of any generation parameters. Because the public web interfaces did not display backend model snapshot identifiers, hidden system prompts, temperature, top-p, maximum output tokens, or other generation parameters, these items were recorded as "not displayed in the public web interface".

The prompt and response languages were English for all models. Each standardized prompt consisted solely of the verbatim English question without any additional model-specific instructions. Each question was submitted individually in a new, independent chat session with no previous conversation history, uploaded files, plug-ins, customized instructions, or follow-up prompts. The complete list of standardized prompts and the corresponding raw model outputs is provided in Supplementary File 1. Model metadata and data collection settings are summarized in Supplementary Table 1.

Readability assessment

The readability of the LLM-generated responses was evaluated using multiple established readability formulas available through an online readability assessment platform (http://readabilityformulas.com/). Given the absence of a universally accepted gold standard for readability assessment and consistent with previous studies, a comprehensive set of widely used readability indices was employed23,27. Seven indices were selected to capture complementary aspects of readability, including sentence length, word length, syllable burden, character-based complexity, reading ease, and estimated grade level, thereby reducing reliance on any single formula. Specifically, the Coleman-Liau Index (CL), Linsear Write Formula (LW), Automated Readability Index (ARI), Simple Measure of Gobbledygook (SMOG), Gunning Fog Index (GFOG), Flesch Reading Ease Score (FRES), and Flesch-Kincaid Grade Level (FKGL) were calculated. These indices estimate reading difficulty or grade level and provide quantitative measures of text readability (Table 2).

Evaluation of educational suitability and information quality

Educational suitability was assessed using the Patient Education Materials Assessment Tool (PEMAT), which comprises two domains: understandability and actionability28. The instrument includes 24 items, with 16 evaluating understandability and eight evaluating actionability. Each item was scored dichotomously (0 = criterion not met; 1 = criterion met), yielding a total score ranging from 0 to 24, with higher scores indicating greater suitability for patient education. Overall information quality was evaluated using the Global Quality Score (GQS)27, a widely used five-point Likert scale for assessing the quality of medical information (Table 3).

On November 25, 2025, two clinical experts with more than 3 years of experience in TN management evaluated all AI-generated responses using both assessment instruments. Before scoring, all AI-generated responses were anonymized with coded response identifiers, and reviewers were blinded to the LLM platform during PEMAT and GQS assessments. Disagreements were resolved by a third senior expert, who adjudicated the final scores. Inter-rater reliability was assessed using Cohen's kappa coefficient, with values >0.75 indicating excellent agreement, 0.40–0.75 indicating acceptable agreement, and <0.40 indicating poor agreement. Cohen's kappa values for both the PEMAT and GQS exceeded 0.75, demonstrating excellent inter-rater reliability.

Statistical analysis

All statistical analyses were performed using R software (version 4.4.3). Data normality was assessed using the Shapiro-Wilk test and Q-Q plots. Normally distributed variables were summarized as mean ± standard deviation and compared using one-way analysis of variance (ANOVA), followed by Tukey's post hoc test when appropriate. Non-normally distributed variables were summarized as medians and interquartile ranges and compared using the Kruskal-Wallis test, followed by Dunn's post hoc test with a Bonferroni correction for pairwise comparisons. Correlations among readability indices, PEMAT scores, and GQS values were evaluated using Spearman's rank correlation coefficient, with the Benjamini-Hochberg correction applied for multiple-testing. A two-sided adjusted P value < 0.05 was considered statistically significant.

Results

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This study systematically evaluated the effects of large language models (LLMs) and different categories of health education content on the readability and quality of generated texts. First, we compared the performance of five LLMs—Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5—with respect to patient education suitability, overall information quality, and multiple readability metrics, including the PEMAT score, GQS, and established readability indices (ARI, FRES, GFOG, FKGL, CL, SMOG, and LW). Second, we examined the same outcome measures across five thematic domains of health education content: basic disease knowledge, etiology and risk factors, diagnosis, treatment, and prevention and rehabilitation. By integrating these two analytical perspectives, we further explored how model type and content category jointly influence text quality and readability, thereby identifying systematic patterns in LLM-generated patient education materials.

Readability analysis across different large language models

Seven established readability indices were used to systematically compare the linguistic complexity of trigeminal neuralgia-related texts generated by five large language models: Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5. The overall results are summarized in Table 4, and the distributions of the individual readability metrics are shown in Figure 1A–G. Statistically significant differences were observed among the five models for all readability indices, with P < 0.001 for each.

GPT-5 exhibited the highest median values for ARI, GFOG, FKGL, CL, and SMOG and the lowest FRES, indicating that it generated texts with longer sentence structures, more complex vocabulary, and the greatest overall reading burden (Figure 1A–G). In contrast, Wenxin Yiyan demonstrated the lowest linguistic complexity. Its median ARI, GFOG, FKGL, CL, and SMOG values were the lowest among the five models, whereas its FRES was the highest.

Doubao, DeepSeek, and Tongyi Qianwen demonstrated intermediate readability levels between GPT-5 and Wenxin Yiyan. Doubao and DeepSeek showed comparable performance across grade-level indices, including ARI, GFOG, FKGL, and CL, and both had significantly higher values than Wenxin Yiyan, with all P < 0.05. These findings indicate a similar overall reading burden for Doubao and DeepSeek, with greater linguistic complexity than Wenxin Yiyan.

Tongyi Qianwen generally yielded values between those of Doubao/DeepSeek and GPT-5 across most readability indices. Notably, its CL score was closer to GPT-5's, suggesting slightly greater linguistic complexity than Doubao and DeepSeek, yet remaining less complex than GPT-5. The LW index exhibited a distribution pattern distinct from that of the other readability metrics. Wenxin Yiyan achieved the highest LW score (60.00), followed by Tongyi Qianwen, DeepSeek, and Doubao, whereas GPT-5 had the lowest LW score (46.50). This discrepancy reflects differences in how individual readability formulas weight sentence length and word difficulty (Figure 1G).

Overall, substantial differences in linguistic complexity were observed among the five large language models. GPT-5 generated the most linguistically complex text, Wenxin Yiyan produced the most readable content, and the remaining models exhibited intermediate levels of readability, although structural differences were evident.

Comparison of patient education suitability and overall quality across large language models

Significant differences in patient education suitability, as assessed using the PEMAT, were observed among the five language models (Figure 1H; Table 4), with all P < 0.001. GPT-5 achieved the highest expert-rated PEMAT score (9.40 ± 1.90), indicating greater educational suitability within the present benchmarking framework. However, this finding should not be interpreted as evidence that GPT-5-generated materials improve real patient comprehension, satisfaction, trust, health behaviors, or clinical outcomes.

DeepSeek achieved an intermediate PEMAT score (6.80 ± 1.06), with a relatively compact score distribution, suggesting consistent performance in generating patient education materials. Nevertheless, its overall educational suitability remained significantly lower than that of GPT-5 (P < 0.001). Doubao, Tongyi Qianwen, and Wenxin Yiyan yielded lower PEMAT scores (5.35 ± 1.04, 5.65 ± 1.09, and 5.35 ± 0.93, respectively). No significant differences were observed among these three models, and all performed significantly worse than DeepSeek and GPT-5 (all P < 0.01). These findings indicate comparable yet limited educational suitability across these models, particularly regarding instructional clarity, linguistic organization, and ease of understanding.

A similar pattern was observed for overall information quality, as assessed using the GQS (Figure 1I). Statistically significant differences were identified among the five models, with all P < 0.001. GPT-5 achieved the highest GQS score (4.00 ± 0.65). Its scores were concentrated within the 4–5 range, with limited variability, indicating consistently high performance in content coherence, structural completeness, and practical usefulness.

DeepSeek and Doubao followed, with GQS scores of 3.35 ± 0.59 and 3.40 ± 0.50, respectively. Their score distributions were centered between 3 and 4, suggesting moderate information quality and reasonable reliability. In contrast, Tongyi Qianwen and Wenxin Yiyan achieved the lowest GQS scores (2.50 ± 0.51 and 2.20 ± 0.52, respectively). Their distributions were skewed toward the lower end of the scale, indicating limitations in information accuracy, structural rigor, and content depth that may reduce their usefulness for patient education.

Overall, GPT-5 achieved the highest expert-rated PEMAT and GQS scores in this dataset, whereas DeepSeek and Doubao demonstrated intermediate performance. Tongyi Qianwen and Wenxin Yiyan received lower GQS ratings. These findings represent expert-rated benchmarking results and should not be interpreted as evidence of clinical readiness or patient-level effectiveness.

Comparison of readability, patient education suitability, and overall quality across health education content categories

Analysis by content category revealed significant differences in readability across the five health education themes (Table 5). For the FRES, statistically significant differences were observed among the thematic categories (P = 0.004). Higher FRES values indicate easier readability. Basic disease knowledge texts were the most readable (median = 28.50), followed by diagnostic content (median = 16.00), etiology and risk factor content (median = 12.50), and treatment-related content (median = 11.50). In contrast, prevention and rehabilitation texts exhibited the lowest readability (median FRES = 1.00), indicating the greatest reading difficulty.

Similar patterns were observed across the other readability indices. GFOG and FKGL differed significantly among the thematic categories (P = 0.002 and P = 0.036, respectively), with both indices indicating higher reading grade levels for etiology and risk factors and prevention and rehabilitation content. The SMOG index likewise suggested that etiology and risk factor texts contained a greater proportion of multisyllabic words (P = 0.039). The largest differences were observed for CL (P < 0.001). Prevention and rehabilitation texts had the longest sentences (median = 21.01), substantially increasing reading difficulty, whereas basic disease knowledge texts had the shortest sentences (median = 14.88), corresponding to the lowest reading burden. No significant differences were observed among content categories for the ARI or LW.

No statistically significant differences in patient education suitability were identified among the five content categories based on PEMAT scores (P = 0.252). Mean PEMAT scores ranged from 5.90 to 7.20, indicating that despite marked differences in linguistic complexity, educational suitability remained relatively consistent across themes. Similarly, overall information quality, as assessed using the GQS, did not differ significantly among content categories (P = 0.822). Mean GQS values ranged from 2.95 to 3.25, indicating that overall information quality was relatively stable across content themes.

Overall, substantial differences in readability were observed across health education themes, with etiology and risk factor content and prevention and rehabilitation content exhibiting the greatest reading difficulty, whereas basic disease knowledge content was the most readable. In contrast, the suitability of patient education and the overall quality of information remained relatively consistent across content categories.

Correlation analysis among readability, patient education suitability, and overall quality

The correlation matrix in Figure 2 shows consistent, robust associations among the readability indices, indicating strong internal consistency. Most readability metrics, including ARI, GFOG, FKGL, CL, and SMOG, exhibited moderate-to-strong positive correlations with one another, with correlation coefficients ranging from 0.64 to 0.97 (all P < 0.001). These findings further support the complementary nature of these indices in assessing linguistic complexity. In contrast, FRES showed significant negative correlations with multiple readability metrics, including ARI, GFOG, FKGL, CL, and SMOG, with absolute correlation coefficients greater than 0.70 (all P < 0.001). This pattern reflects the inverse scoring direction of FRES, where higher scores indicate greater readability. The LW index also demonstrated significant negative correlations with other readability indices, including ARI, FKGL, GFOG, and SMOG, with absolute correlation coefficients ranging from 0.75 to 0.90 (all P < 0.001). Conversely, LW was positively correlated with FRES (correlation coefficient = 0.69, P < 0.001).

Clear associations were also observed between readability indices and overall information quality, as measured by the GQS. GQS demonstrated weak-to-moderate positive correlations with most readability metrics, including ARI, GFOG, FKGL, CL, and SMOG, with correlation coefficients ranging from 0.326 to 0.391 (all P < 0.001). These findings suggest that greater linguistic complexity was associated with higher expert-rated information quality. Conversely, GQS was moderately negatively correlated with FRES (correlation coefficient = −0.408, P < 0.001), indicating that texts requiring a higher reading level generally received higher GQS ratings. GQS also showed a moderate negative correlation with LW (correlation coefficient = −0.421, P < 0.001).

The associations between readability and patient education suitability, as assessed using the PEMAT, were generally weaker but directionally consistent. PEMAT scores exhibited weak positive correlations with ARI, GFOG, FKGL, CL, and SMOG, with correlation coefficients ranging from 0.305 to 0.434 (all P < 0.01). In contrast, PEMAT scores were weakly negatively correlated with FRES and LW, with absolute correlation coefficients of 0.432 and 0.320, respectively (both P < 0.01). These findings suggest that, within this dataset, greater linguistic complexity was associated with modestly higher expert-rated educational suitability, although the strength of these associations was limited.

Importantly, a moderate positive correlation was observed between the PEMAT and GQS scores (correlation coefficient = 0.645, P < 0.001). This was the strongest association observed between the evaluation domains, indicating that responses with higher educational suitability also tended to receive higher overall information quality ratings.

DATA AVAILABILITY:

The complete dataset supporting this study is provided as supplementary material. Supplementary File 1 contains the full list of standardized questions, prompts, and the 100 archived LLM-generated responses. Supplementary Table 2 contains the expert PEMAT and GQS scoring sheets, including item-level PEMAT ratings where available, total PEMAT scores, GQS scores, rater-comparison information, and adjudication records when applicable. The readability scores, figure source data, and R analysis code are provided as Supplementary File 2. Because LLM outputs may change over time due to model updates, interface changes, and platform-level modifications, the archived outputs, rather than regenerated responses, should be used for verification and secondary analysis.

figure-results-1
Figure 1: Comparison of readability, patient education suitability, and overall information quality among five large language models. (A–G) This panel presents the readability indices: Automated Readability Index (ARI), Flesch Reading Ease Score (FRES), Gunning Fog Index (GFOG), Flesch-Kincaid Grade Level (FKGL), Coleman-Liau Index (CL), Simple Measure of Gobbledygook (SMOG), and Linsear Write Formula (LW). (H) This panel shows the total PEMAT scores, and panel (I) shows the Global Quality Scores (GQS). In each box plot, the center line represents the median, the box indicates the interquartile range (IQR), the whiskers extend to 1.5 × IQR, and points beyond the whiskers represent outliers. Abbreviations: ARI = Automated Readability Index; FRES = Flesch Reading Ease Score; GFOG = Gunning Fog Index; FKGL = Flesch Kincaid Grade Level; CL = Coleman Liau Index; SMOG = Simple Measure of Gobbledygook; LW = Linsear Write; PEMAT = the Patient Education Materials Assessment Tool for Understandability and Actionability, printable format; Global Quality Score = GQS. Please click here to view a larger version of this figure.

figure-results-2
Figure 2: Spearman correlation matrix among readability indices, patient education suitability, and overall information quality across all model-generated responses. The matrix reports Spearman correlation coefficients for pairwise associations among ARI, FRES, GFOG, FKGL, CL, SMOG, LW, PEMAT, and GQS. Abbreviations: ARI = Automated Readability Index; FRES = Flesch Reading Ease Score; GFOG = Gunning Fog Index; FKGL = Flesch-Kincaid Grade Level; CL = Coleman-Liau Index; SMOG = Simple Measure of Gobbledygook; LW = Linsear Write; PEMAT = Patient Education Materials Assessment Tool; GQS = Global Quality Score. Please click here to view a larger version of this figure.

Table 1: List of 20 questions related to trigeminal neuralgia. Please click here to download this file.

Table 2: Readability tools, formulas, and descriptions. Please click here to download this file.

Table 3: The Global Quality Score (GQS) quality criteria. Please click here to download this file.

Table 4: Analysis results of different large language models for the most common questions related to trigeminal neuralgia. Please click here to download this file.

Table 5: Analysis results across different dimensions for the most common questions related to trigeminal neuralgia. Please click here to download this file.

Supplementary Table 1: Readability scores of responses generated by the five large language models across all evaluated readability indices. Please click here to download this file.

Supplementary Table 2: Expert ratings of patient education suitability (PEMAT) and overall information quality (GQS) for responses generated by the five large language models. Please click here to download this file.

Supplementary File 1: Standardized trigeminal neuralgia questions and complete responses generated by the five large language models. Please click here to download this file.

Supplementary File 2: R scripts and source data used for statistical analyses and figure generation. Please click here to download this file.

Discussion

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This cross-sectional benchmarking study systematically compared the readability, educational suitability, and overall information quality of trigeminal neuralgia (TN) patient education materials generated by five publicly available large language models (LLMs). Four principal findings emerged. First, readability differed substantially across the models, with GPT-5 generating the most linguistically complex responses and Wenxin Yiyan producing the most readable content. Second, readability and expert-rated information quality were distinct evaluation dimensions, with GPT-5 achieving the highest PEMAT and GQS scores despite its lower readability. Third, content theme influenced readability, with prevention, rehabilitation, and etiology and risk factor topics generally requiring higher reading levels, whereas PEMAT and GQS remained relatively stable across content categories. Finally, readability indices demonstrated strong internal consistency, and correlation analyses revealed weak-to-moderate positive associations between linguistic complexity and both PEMAT and GQS, as well as a moderate positive correlation between PEMAT and GQS. Collectively, these findings provide a benchmark for evaluating LLM-generated TN patient education materials while highlighting the need to balance medical completeness with linguistic accessibility. Importantly, these results represent expert-rated benchmarking outcomes and should not be interpreted as evidence of patient comprehension, factual accuracy, clinical safety, or readiness for routine clinical use.

The findings of this study are consistent with previous evaluations of AI-generated medical information, which have reported substantial variability among LLMs and a frequent disconnect between readability and information quality26,27,29. Previous studies across multiple disease domains have shown that LLM-generated patient education materials often exceed recommended reading levels and that models producing the most readable text do not necessarily provide the most comprehensive or highest-quality information26,30. The findings in TN education support this observation. GPT-5 generated the most linguistically complex responses yet achieved the highest PEMAT and GQS scores, whereas Wenxin Yiyan produced more readable text but lower expert-rated educational suitability and overall quality. Moreover, readability varied across thematic domains, with prevention and rehabilitation and etiology and risk factor topics generally exhibiting greater linguistic complexity, suggesting that both model characteristics and topic complexity influence reading difficulty31.

The apparent trade-off between readability and information quality is consistent with the distinct constructs measured by the evaluation instruments. Readability formulas primarily assess surface-level linguistic features, such as sentence length and lexical complexity, whereas PEMAT and GQS evaluate higher-order attributes, including information organization, completeness, clarity of explanations, and actionability32. For TN education, high-quality responses frequently include discussions of red-flag symptoms, differential diagnosis, pharmacological management and adverse effects, interventional and surgical options, and individualized rehabilitation strategies33. Such clinically comprehensive content inevitably increases terminological density and syntactic complexity, resulting in higher readability scores and improved expert-rated educational quality. Importantly, these findings should not be interpreted as suggesting that more complex language is preferable. Rather, they indicate that current LLMs often improve informational completeness at the expense of linguistic accessibility. Future development of patient education systems should therefore focus on maintaining medical accuracy while simplifying language through plain-language revision, structured presentation of key points, explanation of technical terminology, and clear, actionable guidance appropriate for patients with varying levels of health literacy11,34,35.

Correlation analysis further supports the distinction between readability and educational quality. Most grade-level-based readability indices demonstrated weak-to-moderate positive associations with both PEMAT and GQS, whereas FRES showed the expected inverse relationship because higher FRES scores indicate easier readability. These findings do not imply that greater linguistic complexity is desirable; rather, they suggest that current LLMs often achieve greater informational completeness at the expense of accessibility34. Accordingly, the development of AI-assisted patient education should prioritize medical accuracy and completeness while incorporating plain-language editing, structured presentation of key points, clear explanations of technical terminology, and actionable recommendations tailored to patients with varying levels of health literacy11,35.

An additional methodological finding was the distinct behavior of the LW index compared with the other readability measures. Unlike ARI, FKGL, GFOG, CL, and SMOG, LW followed a pattern more similar to FRES, reflecting differences in how readability formulas operationalize text complexity36. Because LW was originally developed for technical manuals and emphasizes sampled sentence structure and word classification rather than overall sentence or syllable characteristics, it may respond differently to AI-generated medical texts that combine concise statements with longer explanatory passages and uneven distributions of technical terminology37,38. These findings reinforce the importance of a multi-metric assessment framework, as no single readability index adequately captures all dimensions of text complexity. Consistent with previous studies of AI-generated health information, overall readability should therefore be interpreted using convergent evidence across multiple complementary indices rather than any individual measure39.

Content category also influenced readability. Prevention, rehabilitation, and etiology and risk factor topics generally produced texts with higher reading difficulty than basic disease knowledge, likely because these topics require conditional recommendations, mechanistic explanations, and more specialized terminology40,41. In contrast, basic disease knowledge is primarily definition-based and can be conveyed in simpler language42. Despite these differences, PEMAT and GQS remained relatively consistent across thematic domains, suggesting that expert-rated educational suitability and overall information quality were maintained across different types of patient questions35,43.

These findings have several implications for clinical practice and future LLM development. As patients increasingly use LLMs to obtain information before or after medical consultations, AI-generated content should complement, rather than replace, clinician counseling, particularly regarding medication safety, indications for interventional or surgical treatment, and recognition of red-flag symptoms requiring urgent evaluation44. Healthcare organizations may benefit from providing expert-reviewed, plain-language educational resources while using LLMs as drafting or decision-support tools rather than autonomous patient education systems. Future models should incorporate readability-aware generation strategies, transparent reporting of model versions and generation dates, explicit communication of uncertainty, and governance procedures including clinician review, factual accuracy verification, hallucination screening, medication safety checks, red-flag guidance, and patient comprehension testing before clinical implementation26,44,45.

This study has several limitations. Model outputs were collected from publicly accessible web interfaces on a single date, and subsequent model updates, interface changes, hidden system prompts, or default generation settings may affect reproducibility. The 20-question set was developed by clinical experts rather than derived from patient search logs or formal patient interviews, a practice that may introduce selection bias. Furthermore, the study relied on expert-rated PEMAT and GQS assessments and did not evaluate patient comprehension, trust, satisfaction, health behaviors, or clinical outcomes. Readability formulas, PEMAT, and GQS also do not directly assess factual accuracy, hallucination risk, citation accuracy, completeness of contraindications, or clinical safety. Consequently, these findings should be interpreted as an expert-rated cross-sectional benchmarking analysis rather than evidence that any LLM-generated material is ready for routine clinical use. Future studies should include archived model outputs, detailed model metadata, multilingual evaluations, larger question sets, formal safety audits, factual accuracy assessments, and patient-centered outcome measures before LLM-generated educational materials are considered for clinical implementation.

In this expert-rated cross-sectional benchmarking study, publicly available large language models differed substantially in the readability, educational suitability, and overall information quality of trigeminal neuralgia patient education materials. GPT-5 achieved the highest expert-rated PEMAT and GQS scores but also generated the most linguistically complex responses, highlighting that readability and expert-rated educational quality represent related yet distinct dimensions of patient education. Although some models produced more readable content, this did not necessarily translate into greater educational suitability or overall information quality. Because this study did not evaluate factual accuracy through a dedicated safety audit, hallucination risk, citation accuracy, patient comprehension, trust, satisfaction, health behaviors, or clinical outcomes, the findings should be interpreted as an expert-rated benchmarking analysis rather than evidence of clinical readiness. Future studies should integrate expert review, formal safety and factual accuracy assessments, and patient-centered evaluations to support the safe and effective use of LLM-generated materials in patient education.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The authors declare that they have no competing interests.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. The authors thank the clinical colleagues who provided feedback on the trigeminal neuralgia question set and assisted in refining the evaluation framework. We also thank all contributors who supported manuscript preparation and revision.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
DeepSeek chat platformDeepSeekhttps://chat.deepseek.com/Publicly accessible large language model interface used to generate responses
Doubao chat platformDoubaohttps://www.doubao.com/chat/Publicly accessible large language model interface used to generate responses
GPT 5 OpenAIhttps://chat.openai.com/Publicly accessible large language model interface used to generate responses
R software, version 4.4.3R Foundation for Statistical ComputingNot applicableStatistical analysis and visualization
Readability Formulas online readability calculatorReadability FormulasNot applicableOnline tool used to compute readability indices
Tongyi Qianwen chat platformTongyi Qianwenhttps://www.tongyi.com/Publicly accessible large language model interface used to generate responses
Wenxin Yiyan chat platformWenxin Yiyanhttps://yiyan.baidu.com/Publicly accessible large language model interface used to generate responses

References

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,
  1. Headache Classification Committee of the International Headache Society (IHS). The International Classification of Headache Disorders, 3rd edition. Cephalalgia. 2018;38(1):1-211.
  2. Ashina S, et al. Trigeminal neuralgia. Nat Rev Dis Primers. 2024;10(1):39.
  3. Donertas-Ayaz B, Caudle RM. Locus coeruleus-noradrenergic modulation of trigeminal pain: Implications for trigeminal neuralgia and psychiatric comorbidities. Neurobiol Pain. 2023;13:100124.
  4. Martinelli R, et al. Psychological assessment in patients affected by trigeminal neuralgia: A systematic review. Neurosurg Rev. 2025;48(1):414.
  5. Melek LN, et al. Comparison of the neuropathic pain symptoms and psychosocial impacts of trigeminal neuralgia and painful posttraumatic trigeminal neuropathy. J Oral Facial Pain Headache. 2019;33(1):77-88.
  6. Zakrzewska JM, et al. Evaluating the impact of trigeminal neuralgia. Pain. 2017;158(6):1166-1174.
  7. Gambeta E, Chichorro JG, Zamponi GW. Trigeminal neuralgia: An overview from pathophysiology to pharmacological treatments. Mol Pain. 2020;16:1744806920901890.
  8. De Toledo IP, et al. Prevalence of trigeminal neuralgia: A systematic review. J Am Dent Assoc. 2016;147(7):570-576.e2.
  9. Svedung Wettervik T, et al. Incidence of trigeminal neuralgia: A population-based study in central Sweden. Eur J Pain. 2023;27(5):580-587.
  10. Bendtsen L, et al. European Academy of Neurology guideline on trigeminal neuralgia. Eur J Neurol. 2019;26(6):831-849.
  11. Cruccu G, et al. Trigeminal neuralgia: New classification and diagnostic grading for practice and research. Neurology. 2016;87(2):220-228.
  12. Araya EI, et al. Trigeminal neuralgia: Basic and clinical aspects. Curr Neuropharmacol. 2020;18(2):109-119.
  13. Chong MS, Bahra A, Zakrzewska JM. Guidelines for the management of trigeminal neuralgia. Cleve Clin J Med. 2023;90(6):355-362.
  14. Heinskou TB, et al. Favourable prognosis of trigeminal neuralgia when enrolled in a multidisciplinary management program: A two-year prospective real-life study. J Headache Pain. 2019;20(1):23.
  15. Liu S, et al. Comparative evaluation of ChatGPT and Gemini in brain-computer interface patient education: A multidimensional analysis of reliability, accuracy, comprehensibility, and readability. Int J Med Inform. 2026;206:106164.
  16. Forgie EME, et al. Social media and the transformation of the physician-patient relationship: Viewpoint. J Med Internet Res. 2021;23(12):e25230.
  17. Gantenbein L, Navarini AA, Maul LV, Brandt O, Mueller SM. Internet and social media use in dermatology patients: Search behavior and impact on the patient-physician relationship. Dermatol Ther. 2020;33(6):e14098.
  18. Youssef Y, et al. Social media and internet use among orthopedic patients in Germany: A multicenter survey. Front Digit Health. 2025;7:1486296.
  19. De Martino I, et al. Social media for patients: Benefits and drawbacks. Curr Rev Musculoskelet Med. 2017;10(1):141-145.
  20. Johnson KB, et al. Precision medicine, AI, and the future of personalized health care. Clin Transl Sci. 2021;14(1):86-93.
  21. Hirani R, et al. Artificial intelligence and healthcare: A journey through history, present innovations, and future possibilities. Life (Basel). 2024;14(5):557.
  22. Karatas G, Kirik F, Karatas ME, Cakir A, Ozdemir H. Comparative performance of ChatGPT o3-mini-high and DeepSeek-R1 in ophthalmology: An evaluation of diagnostic reasoning and case-based problem solving. J Fr Ophtalmol. 2025;49(1):104712.
  23. Hanci V, et al. Assessment of readability, reliability, and quality of ChatGPT, Bard, Gemini, Copilot, and Perplexity responses on palliative care. Medicine (Baltimore). 2024;103(33):e39305.
  24. Liu C, et al. Potential to perpetuate social biases in health care by Chinese large language models: A model evaluation study. Int J Equity Health. 2025;24(1):206.
  25. Haupt CE, Marks M. AI-generated medical advice—GPT and beyond. JAMA. 2023;329(16):1349-1350.
  26. Aydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: A scoping review of applications in medicine. Front Med (Lausanne). 2024;11:1477898.
  27. Kara M, et al. Evaluating the readability, quality, and reliability of responses generated by ChatGPT, Gemini, and Perplexity on the most commonly asked questions about ankylosing spondylitis. PLoS One. 2025;20(6):e0326351.
  28. Perez Rivera LR, et al. Evaluating the quality and reliability of large language models for plastic surgery patient education: A comparative analysis of ChatGPT and OpenEvidence. Aesthet Surg J. 2025. doi:10.1093/asj/sjaf249.
  29. Wilhelm TI, Roos J, Kaczmarczyk R. Large language models for therapy recommendations across three clinical specialties: Comparative study. J Med Internet Res. 2023;25:e49324.
  30. Will J, et al. Enhancing the readability of online patient education materials using large language models: Cross-sectional study. J Med Internet Res. 2025;27:e69955.
  31. Tukur Jido J, Al-Wizni A, Aung SL. Readability of AI-generated patient information leaflets on Alzheimer's disease, vascular dementia, and delirium. Cureus. 2025;17(6):e85463.
  32. Luo Z, Lin C, Kim TH, Shin YS, Ahn ST. Quality and readability analysis of artificial intelligence-generated medical information related to prostate cancer: A cross-sectional study of ChatGPT and DeepSeek. World J Mens Health. 2025. doi:10.5534/wjmh.250144.
  33. Joseph P, Silva NA, Nanda A, Gupta G. Evaluating the readability of online patient education materials for trigeminal neuralgia. World Neurosurg. 2020;144:e934-e938.
  34. Dong C, et al. Comparative evaluation of large language models in delivering guideline-compliant recommendations for topical NSAID use in musculoskeletal pain: A multidimensional analysis. Clin Rheumatol. 2025;44(11):4703-4710.
  35. Ozduran E, Buyukcoban S. Evaluating the readability, quality, and reliability of online patient education materials on post-COVID pain. PeerJ. 2022;10:e13686.
  36. Wang LW, Miller MJ, Schmitt MR, Wen FK. Assessing readability formula differences with written health information materials: Application, results, and recommendations. Res Social Adm Pharm. 2013;9(5):503-516.
  37. Singh SP, et al. Comprehension profile of patient education materials in endocrine care. Kans J Med. 2022;15(2):247-252.
  38. Marder RS, et al. ChatGPT-3.5 and ChatGPT-4.0 do not reliably create readable patient education materials for common orthopaedic upper- and lower-extremity conditions. Arthrosc Sports Med Rehabil. 2025;7(1):101027.
  39. Nasra M, et al. Can artificial intelligence improve patient educational material readability? A systematic review and narrative synthesis. Intern Med J. 2025;55(1):20-34.
  40. Bhatt C, et al. Evaluating readability, understandability, and actionability of online printable patient education materials for cholesterol management: A systematic review. J Am Heart Assoc. 2024;13(8):e030140.
  41. Avra TD, Le M, Hernandez S, Thure K, Ulloa JG. Readability assessment of online peripheral artery disease education materials. J Vasc Surg. 2022;76(6):1728-1732.
  42. Rooney MK, et al. Readability of patient education materials from high-impact medical journals: A 20-year analysis. J Patient Exp. 2021;8:2374373521998847.
  43. Ahmadzadeh K, et al. Patient education information material assessment criteria: A scoping review. Health Info Libr J. 2023;40(1):3-28.
  44. Duan L, Yao Z, Li X, Wu Y, Sheng D. Comparing large language models and human doctors in symptom-driven online medical consultations: A case study on trigeminal neuralgia. Digit Health. 2025;11:20552076251388140.
  45. Werner C, Harden J, Lawton J. Pathways to a diagnosis of trigeminal neuralgia: A qualitative study of patients' experiences. BMC Prim Care. 2025;26(1):65.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Large Language ModelsReadability AssessmentEducational QualityPEMAT EvaluationGlobal Quality ScoreHealth Information QualityModel BenchmarkingPatient Comprehension

Related Articles