清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整的填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

Performance of ChatGPT on the Chinese Postgraduate Examination for Clinical Medicine: Survey Study

理解力 背景(考古学) 集合(抽象数据类型) 医疗保健 政府(语言学) 代理(哲学) 质量(理念) 医学教育 医学 心理学 计算机科学 语言学 经济 古生物学 程序设计语言 哲学 认识论 生物 经济增长
作者
Peng Yu,Changchang Fang,Xiaolin Liu,Wanying Fu,Jitao Ling,Zhiwei Yan,Yuan Jiang,Zhengyu Cao,Maoxiong Wu,Zhiteng Chen,Wengen Zhu,Yuling Zhang,Ayiguli Abudukeremu,Yue Wang,Xiao Liu,J. Wang
出处
期刊:JMIR medical education [JMIR Publications Inc.]
卷期号:10: e48514-e48514 被引量:5
标识
DOI:10.2196/48514
摘要

Background ChatGPT, an artificial intelligence (AI) based on large-scale language models, has sparked interest in the field of health care. Nonetheless, the capabilities of AI in text comprehension and generation are constrained by the quality and volume of available training data for a specific language, and the performance of AI across different languages requires further investigation. While AI harbors substantial potential in medicine, it is imperative to tackle challenges such as the formulation of clinical care standards; facilitating cultural transitions in medical education and practice; and managing ethical issues including data privacy, consent, and bias. Objective The study aimed to evaluate ChatGPT’s performance in processing Chinese Postgraduate Examination for Clinical Medicine questions, assess its clinical reasoning ability, investigate potential limitations with the Chinese language, and explore its potential as a valuable tool for medical professionals in the Chinese context. Methods A data set of Chinese Postgraduate Examination for Clinical Medicine questions was used to assess the effectiveness of ChatGPT’s (version 3.5) medical knowledge in the Chinese language, which has a data set of 165 medical questions that were divided into three categories: (1) common questions (n=90) assessing basic medical knowledge, (2) case analysis questions (n=45) focusing on clinical decision-making through patient case evaluations, and (3) multichoice questions (n=30) requiring the selection of multiple correct answers. First of all, we assessed whether ChatGPT could meet the stringent cutoff score defined by the government agency, which requires a performance within the top 20% of candidates. Additionally, in our evaluation of ChatGPT’s performance on both original and encoded medical questions, 3 primary indicators were used: accuracy, concordance (which validates the answer), and the frequency of insights. Results Our evaluation revealed that ChatGPT scored 153.5 out of 300 for original questions in Chinese, which signifies the minimum score set to ensure that at least 20% more candidates pass than the enrollment quota. However, ChatGPT had low accuracy in answering open-ended medical questions, with only 31.5% total accuracy. The accuracy for common questions, multichoice questions, and case analysis questions was 42%, 37%, and 17%, respectively. ChatGPT achieved a 90% concordance across all questions. Among correct responses, the concordance was 100%, significantly exceeding that of incorrect responses (n=57, 50%; P<.001). ChatGPT provided innovative insights for 80% (n=132) of all questions, with an average of 2.95 insights per accurate response. Conclusions Although ChatGPT surpassed the passing threshold for the Chinese Postgraduate Examination for Clinical Medicine, its performance in answering open-ended medical questions was suboptimal. Nonetheless, ChatGPT exhibited high internal concordance and the ability to generate multiple insights in the Chinese language. Future research should investigate the language-based discrepancies in ChatGPT’s performance within the health care context.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
更新
大幅提高文件上传限制,最高150M (2024-4-1)

科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
8秒前
含糊的茹妖完成签到 ,获得积分10
9秒前
微卫星不稳定完成签到 ,获得积分0
21秒前
Jenny完成签到,获得积分10
26秒前
会飞的鹦鹉完成签到 ,获得积分10
1分钟前
1分钟前
1分钟前
2分钟前
科研通AI2S应助帮帮我好吗采纳,获得10
2分钟前
彭于晏应助木木三采纳,获得10
2分钟前
小羊咩完成签到 ,获得积分10
2分钟前
席江海完成签到,获得积分10
2分钟前
2分钟前
2分钟前
科研通AI2S应助帮帮我好吗采纳,获得10
2分钟前
2分钟前
2分钟前
木木三发布了新的文献求助10
3分钟前
桐桐应助科研通管家采纳,获得20
3分钟前
英俊的铭应助帮帮我好吗采纳,获得10
3分钟前
wenbo完成签到,获得积分10
3分钟前
3分钟前
qiao发布了新的文献求助10
3分钟前
chenying完成签到 ,获得积分0
3分钟前
大咖完成签到 ,获得积分10
3分钟前
qiao完成签到,获得积分10
3分钟前
fantw完成签到,获得积分10
3分钟前
zhao完成签到,获得积分10
3分钟前
3分钟前
小二郎应助木木三采纳,获得10
3分钟前
4分钟前
Drwenlu发布了新的文献求助10
4分钟前
木木三发布了新的文献求助10
4分钟前
沙海沉戈完成签到,获得积分0
4分钟前
木木三完成签到,获得积分20
4分钟前
4分钟前
研友_Z119gZ完成签到 ,获得积分10
4分钟前
theo完成签到 ,获得积分10
4分钟前
Science完成签到,获得积分10
4分钟前
可爱的函函应助颖宝老公采纳,获得10
4分钟前
高分求助中
Sustainability in Tides Chemistry 2800
The Young builders of New china : the visit of the delegation of the WFDY to the Chinese People's Republic 1000
Rechtsphilosophie 1000
Bayesian Models of Cognition:Reverse Engineering the Mind 888
Defense against predation 800
Very-high-order BVD Schemes Using β-variable THINC Method 568
Chen Hansheng: China’s Last Romantic Revolutionary 500
热门求助领域 (近24小时)
化学 医学 生物 材料科学 工程类 有机化学 生物化学 物理 内科学 纳米技术 计算机科学 化学工程 复合材料 基因 遗传学 催化作用 物理化学 免疫学 量子力学 细胞生物学
热门帖子
关注 科研通微信公众号,转发送积分 3137034
求助须知:如何正确求助?哪些是违规求助? 2788014
关于积分的说明 7784270
捐赠科研通 2444088
什么是DOI,文献DOI怎么找? 1299724
科研通“疑难数据库(出版商)”最低求助积分说明 625522
版权声明 600999