Abstract
BACKGROUND: H. pylori infection is a worldwide health issue, fueling rising demand for medical counseling. LLMs have the potential to serve in medical counseling. However, their performance remains unclear. OBJECTIVE: This study aimed to evaluate the effectiveness of LLMs in providing H. pylori related medical counseling in a Chinese context. METHODS: 20 H. pylori-related questions were collected, covering four domains: definition and symptoms, diagnosis, treatment, and prevention. Each question was asked thrice in Chinese to each LLM. We assessed the responses across five dimensions (accuracy, relevance, completeness, clarity, and reliability). RESULTS: 1. In the first batch of tests, the overall performance distribution was 33.3% good, 66.1% medium, and 0.6% poor, respectively. No significant differences were observed among the three LLMs (p=0.158). Good performance was observed with 47.8% in accuracy, 53.9% in relevance, 68.3% in completeness, 36.7% in clarity, and 36.1% in reliability. No significant differences were observed in accuracy, relevance, completeness, or clarity. Reliability differed significantly (p<0.001), with Ernie Bot achieving the best performance. 2. The second test batch yielded performance rates of 70.6% good, 29.4% medium, and 0% poor, with a significant difference among the three LLMs (p=0.018). Doubao attained the best performance, surpassing other models in relevance and clarity. 3. The newly assessed AI batch showed markedly superior overall performance to the counterpart evaluated more than a year prior. CONCLUSION: This study is the first to evaluate the effectiveness of various LLMs in H. pylori-related medical counseling in a real-world setting. The study showed that while LLMs generally performed acceptably in terms of accuracy, relevance, and completeness, their clarity and reliability were less satisfactory. Ernie Bot, developed by Chinese company, outperformed ChatGPT in certain aspects of medical counseling in Chinese. With the guidance of professionals, LLMs can serve as potential aids for medical counseling.