基于大语言模型的临床试验质量问题文本的智能化分类研究
2.香港大学深圳医院临床试验中心,广东 深圳 518053;
3.桂林医科大学人工智能医学院,广西 桂林 541199;
4.深圳大学生命与海洋科学学院,广东 深圳 518060;
5.深圳理工大学药学院,广东 深圳 518107
收稿日期: 2025-12-22
修回日期: 2026-03-04
录用日期: 2026-03-18
网络出版日期: 2026-08-17
基金资助
国家药品监督管理局药品审评检查大湾区分中心监管科学研究课题(GBA-JGKX-2501);国家药品监督管理局信息中心课题(20240718001)
Intelligent Classification of Interpretability for Clinical Trial Quality Issue Text Based on Large Language Models
Received date: 2025-12-22
Revised date: 2026-03-04
Accepted date: 2026-03-18
Online published: 2026-08-17
目的:探索大语言模型识别与分类临床试验质量问题的能力。方法:构建含12个一级指标、67个二级指标的临床试验质量问题文本分类框架,整理1035条问题文本,经数据扩增至5733条。比较Qwen3、Bert、Gemma3在LoRA微调下的分类性能,引入Shap进行可解释分析。结果:Bert模型表现最优,数据增强后,一、二级指标准确率分别达90.3%和81.2%,较增强前提升11.8%和36.8%。尽管Bert模型整体表现更优,但在生物样本采集、实验室检查、合并用药等特定场景下,Qwen3与Gemma3准确率超过Bert 0.4%~13.4%。Shap分析显示,模型决策关键词与人类判断一致。结论:微调后的大语言模型可有效分类临床试验质量问题文本,为解决专业质量控制领域文本分类任务中存在的从业人员的素质与判断标准不一的问题提供可行方案,从而有效辅助人工质量控制流程。
徐磊
,
余满
,
卓宝珊
,
许燕萍
,
黄龙凯
,
罗嘉慧
,
孔艺
,
周文菁
.
基于大语言模型的临床试验质量问题文本的智能化分类研究
Objective: To investigate the capability of large language models (LLMs) in recognizing and classifying quality issue texts of clinical trials.Methods: A text classification framework comprising 12 primary and 67 secondary indicators was established. A total of 1035 quality issues texts were collected and augmented to 5733 samples using LLMs. Three models (Qwen3, Bert, Gemma3) were fine-tuned with LoRA and evaluated, with Shap employed for interpretability analysis.Results: Bert achieved the best performance, with post-augmentation accuracies of 90.3% for primary and 81.2% for secondary indicators-improvements of 11.8% and 36.8%, respectively. Although Bert outperformed overall, Qwen3 and Gemma3 exceeded Bert by 0.4%-13.4% in specific scenarios such as biospecimen collection, laboratory tests, and concomitant medication. Shap analysis confirmed alignment between model decisions and human reasoning.Conclusion: Fine-tuned LLMs can effectively classify clinical trials quality issue texts, offering a viable solution to address inconsistencies arising from varying practitioner expertise and judgment standards, thereby supporting manual quality control workflows.
[1] 国家药品监督管理局.国家药监局 国家卫生健康委关于发布药物临床试验质量管理规范的公告(2020年第57号)[EB/OL].(2020-04-23)[2025-11-19].https://www.nmpa.gov.cn/yaopin/ypggtg/20200426162401243.html?type=pc&m=.
[2] 谭琴,邱攀博,李高扬,等.药物临床试验机构质量管理现状调查分析[J].医药导报,2023,42(12):1884-1889.
[3] 尚美霞,阎小妍,李雪迎,等.采用多阅片者多病例设计评估AI辅助医疗产品临床试验的样本量估算和应用[J].中国卫生统计,2022,39(1):14-18.
[4] Seufferlein T, Ettrich T, Stein A. Predicting resistance to first-line FOLFOX plus bevacizumab in metastatic colorectal cancer: final results of the multicenter, international PERMAD trial[J]. J Clin Oncology, 2021,39(3_suppl):115.
[5] 国家药品监督管理局.国家药监局关于发布人工智能医用软件产品分类界定指导原则的通告(2021年第47号)[EB/OL].(2021-07-01)[2025-11-19].https://www.nmpa.gov.cn/ylqx/ylqxggtg/20210708111147171.html?type=pc&m=.
[6] 国家药品监督管理局医疗器械技术审评中心.国家药监局器审中心关于发布人工智能医疗器械注册审查指导原则的通告(2022年第8号)[EB/OL].(2022-03-07)[2025-11-19].https://www.cmde.org.cn//xwdt/shpgzgg/gztg/20220309090800158.html.
[7] Jiang Min, Zhao Shuhua, Mei Yun, et al. Real-time, risk-based clinical trial quality management in China: development of a digital monitoring platform[J].JMIR Med Inform,2025,13:e64114.
[8] 广东省药学会.药物临床试验 质量管理·广东共识(2025年版)[J].今日药学,2025,3(10):721-728.
[9] 张基,梁淑红.药物临床试验质量管理规范培训融入医院药学研究生培养体系的实践与思考[J].中国药物与临床,2025,25(3):160-163.
[10] 国家药品监督管理局食品药品审核查验中心.药品注册核查要点与判定原则(药物临床试验)(试行)[EB/OL].(2021-12-17)[2026-01-14].https://www.cfdi.org.cn/cfdi/resource/news/14200.html.
[11] Ding Bosheng, Qin Chengwei, Zhao Ruochen , et al. Data augmentation using large language models: data perspectives, learning paradigms and challenges[EB/OL].(2024-03-05)[2026-06-06].https://arxiv.org/abs/2403.02990.
[12] Yang An, Li Anfeng, Yang Baosong, et al. Qwen3 Technical Report[EB/OL].(2025-05-14)[2026-06-06].https://arxiv.org/abs/2505.09388.
[13] Koroteev MV. BERT: A Review of Applications in Natural Language Processing and Understanding[EB/OL].(2021-03-22)[2026-06-06].https://arxiv.org/abs/2103.11943.
[14] Aishwarya Kamath, Johan Ferret, Shreya Pathak.et al. Gemma 3 Technical Report[EB/OL].(2025-03-25)[2026-06-06].https://arxiv.org/abs/2503. 19786.
[15] Hu EJ, Shen Yelong, Wallis P, et al. LoRA: Low-rank adaptation of large language models[EB/OL].(2022-04-25)[2026-06-06].https://openreview.net/forum?id=nZeVKeeFYf9.
[16] Diederik P. Kingma, Jimmy Ba.Adam: A Method for Stochastic Optimization[EB/OL].(2014-12-22)[2026-06-26].https://arxiv.org/abs/1412. 6980.
[17] Mosca E, Szigeti F, Tragianni S, et al. SHAP-based explanation methods: a review for NLP interpretability[C]//Proceedings of the 29th international conference on computational linguistics, 2022:4593-4603.
/
| 〈 |
|
〉 |