Evaluating the Response of AI-Based Large Language Models to Common Patient Concerns About Endodontic Root Canal Treatment: A Comparative Performance Analysis

评估基于人工智能的大型语言模型对患者常见根管治疗问题的响应：一项比较性能分析

阅读：1

期刊：	Journal of Clinical Medicine	影响因子：	2.900
时间：	2025	起止号：	2025 Oct 22;14(21)
doi：	10.3390/jcm14217482

Abstract

Objectives: The aim of this study was to compare the responses of large language models (LLMs)-DeepSeek V3, GPT 5, and Gemini 2.5 Flash-to patients' frequently asked questions (FAQs) regarding root canal treatment in terms of accuracy and comprehensiveness, and to assess the potential roles of these models in patient education and health literacy. Methods: A total of 37 open-ended FAQs, compiled from American Association of Endodontists (AAE) patient education materials and online resources, were presented to three LLMs. Responses were evaluated by expert clinicians on a 5-point Likert scale for accuracy and comprehensiveness. Inter-rater and test-retest reliability were assessed using intraclass correlation coefficients (ICCs). Differences among models were analyzed with the Kruskal-Wallis H test, followed by pairwise Mann-Whitney U tests with effect sizes (Cliff's delta, δ). A p-value < 0.05 was considered statistically significant. Results: Inter-rater agreement was excellent, with ICCs of 0.92 for accuracy and 0.91 for comprehensiveness. Test-retest reliability also demonstrated high consistency (ICCs of 0.90 for accuracy and 0.89 for comprehensiveness). DeepSeek V3 achieved the highest scores, with a mean accuracy of 4.81 ± 0.39 and a mean comprehensiveness of 4.78 ± 0.41, demonstrating statistically superior performance compared to GPT 5 (accuracy 4.0 ± 0.0; comprehensiveness 4.05 ± 0.4; p < 0.05, δ = 0.81 for accuracy, δ = 0.69 for comprehensiveness) and Gemini 2.5 Flash (accuracy 3.83 ± 0.68; comprehensiveness 3.81 ± 0.7; p < 0.05, δ = 0.71 for accuracy, δ = 0.70 for comprehensiveness). No significant difference was observed between GPT 5 and Gemini 2.5 Flash for either accuracy (p = 0.109, δ = 0.16) or comprehensiveness (p = 0.058, δ = 0.21). Conclusions: LLMs, such as DeepSeek V3, which can provide satisfactory responses to FAQs may serve as valuable supportive tools in patient education and health literacy; however, expert clinician oversight remains essential in clinical decision-making and treatment planning. When used appropriately, LLMs can enhance patient awareness and support satisfaction throughout the root canal treatment.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用；引用内容仅为补充信息，不代表本站立场。

2、若认为本页面引用内容涉及侵权，请及时与本站联系，我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容，需注明“来源：[生知库]”并获得授权；使用引用内容的，需自行联系原作者获得许可。

4、投稿及合作请联系：info@biocloudy.com。

肿瘤免疫

炎症

T细胞

线粒体

凋亡

转录调控

巨噬细胞

自噬

传染病

氧化应激

肠道菌群

磷酸化

血管生成

囊泡

3D/类器官

单细胞

中性粒细胞

外泌体

DNA甲基化

miRNA

药物研究

铁死亡

细胞衰老

乙酰化

缺氧低氧

泛素化

树突状细胞

炎性小体

组蛋白修饰

肿瘤微环境

lncRNA

代谢重编程

焦亡

m6A/m5C/m7G

内质网应激

空间多组学

细胞基因治疗

治疗耐药

相分离

Treg

上皮间质转化

免疫代谢

染色质重塑

脂质过氧化

蛋白质稳态

脂代谢

细胞极性

铁代谢

氨基酸代谢

碱基编辑

cGAS-STING

肠脑轴

蛋白降解

乳酸化

翻译调控

circRNA

piRNA

肿瘤异质性

NK 细胞

氧化脂质

MDSC

NETosis

低氧缺氧

溶酶体功能

琥珀酰化

细胞干性

CAR-NK

冷应激

RNA 编辑

Tfh

巴豆酰化

器官芯片

表观遗传记忆

铜死亡

器官纤维化

线粒体未折叠蛋白反应

空间代谢组

程序性坏死

自噬流

MAIT 细胞

肠肝轴

丙酰化