Comparative Performance of Gemini 3 Pro and GPT-5 Family Models on Ophthalmology Board-Style Questions

Gemini 3 Pro 和 GPT-5 系列模型在眼科考试题型上的性能比较

阅读：1

作者：Shean,Ryan S,Mallapu,Jayanth Kumar,Shah,Tathya,Rasheed,Haroon Adam,Younessi,David N,Tham,Yih Chung,Nguyen,Van,Bolo,Kyle,Xu,Benjamin Y

期刊：	Ophthalmology Science	影响因子：	4.600
时间：	2026	起止号：	2026 May;6(5):101145
doi：	10.1016/j.xops.2026.101145

Abstract

OBJECTIVE: To compare the performance of state-of-the-art Gemini and GPT models on ophthalmology board-style questions and examine variation by subspecialty, cognitive complexity, and question type. DESIGN: A cross-sectional evaluation of 12 distinct large language model (LLM) configurations using a standardized ophthalmology question set. SUBJECTS: Five hundred multiple-choice questions (250 from the American Academy of Ophthalmology's Basic and Clinical Science Course [BCSC]; 250 StatPearls). METHODS: Twelve configurations of the following LLMs: Gemini 3 Pro, Gemini 2.5 Pro, GPT-5.1 Pro, GPT-5 Pro, GPT-5.2, GPT-5.1, and GPT-5, interpreted the questions using standardized prompting procedures. Questions were categorized by subspecialty, multimodal content (image vs. text-only), and cognitive complexity (first, second, or third order). Accuracy, paired discordance (McNemar tests), and one-way analysis of variance with Tukey correction were used to compare performance. Human benchmarking used BCSC percent-correct data. MAIN OUTCOME MEASURES: Overall accuracy, subspecialty accuracy, image vs. nonimage accuracy, cognitive-complexity accuracy, and paired model-level discordance. RESULTS: Model accuracy ranged from 81.4% to 94.0%. Gemini 3 Pro High Reasoning achieved the highest accuracy (94.0%), followed by Gemini 3 Pro Low Reasoning (92.4%). GPT-5.1 Pro led the GPT family (90.4%), whereas GPT-5.2 Base Model performed lowest (81.4%). Analysis of variance showed significant heterogeneity (P < 0.001), but most Tukey-corrected pairwise differences were nonsignificant. McNemar tests demonstrated significantly more correct paired responses for Gemini 3 Pro High Reasoning than for GPT-5.2 and all GPT-5/5.1 variants. Models performed markedly better on BCSC (mean 94.4%) than StatPearls (81.9%); human BCSC mean accuracy was 64.5%. Image-based items produced a 10- to 22-point accuracy decrement across all systems. Accuracy declined with increasing cognitive complexity, with the clearest separation on third-order management questions. CONCLUSIONS: Gemini 3 Pro had the best general-purpose LLM performance on ophthalmology board-style questions, providing near-perfect accuracy, while outperforming all GPT-5 family variants across domains and complexity levels. Significant deficits on image-based and third-order questions highlight persistent multimodal limitations and the need for ongoing benchmarking using challenging, clinically grounded datasets. FINANCIAL DISCLOSURES: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用；引用内容仅为补充信息，不代表本站立场。

2、若认为本页面引用内容涉及侵权，请及时与本站联系，我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容，需注明“来源：[生知库]”并获得授权；使用引用内容的，需自行联系原作者获得许可。

4、投稿及合作请联系：info@biocloudy.com。

肿瘤免疫

炎症

T细胞

凋亡

线粒体

转录调控

巨噬细胞

自噬

传染病

氧化应激

血管生成

肠道菌群

磷酸化

囊泡

3D/类器官

单细胞

中性粒细胞

外泌体

药物研究

DNA甲基化

细胞衰老

miRNA

铁死亡

缺氧低氧

乙酰化

泛素化

组蛋白修饰

炎性小体

树突状细胞

代谢重编程

肿瘤微环境

焦亡

lncRNA

m6A/m5C/m7G

空间多组学

细胞基因治疗

内质网应激

相分离

治疗耐药

Treg

免疫代谢

上皮间质转化

染色质重塑

脂质过氧化

蛋白质稳态

铁代谢

脂代谢

cGAS-STING

肠脑轴

细胞极性

氨基酸代谢

碱基编辑

乳酸化

蛋白降解

circRNA

翻译调控

肿瘤异质性

piRNA

低氧缺氧

NK 细胞

MDSC

氧化脂质

溶酶体功能

NETosis

RNA 编辑

细胞干性

琥珀酰化

CAR-NK

冷应激

Tfh

器官芯片

巴豆酰化

表观遗传记忆

空间代谢组

器官纤维化

铜死亡

线粒体未折叠蛋白反应

程序性坏死

自噬流

肠肝轴

MAIT 细胞

丙酰化