• 中国核心期刊数据库收录期刊
  • 中文科技期刊数据库收录期刊
  • 中国期刊全文数据库收录期刊
  • 中国学术期刊综合评价数据库统计源期刊等

快速检索引用检索图表检索高级检索

专栏:药物临床试验

基于大语言模型的临床试验质量问题文本的智能化分类研究

  • 徐磊 ,
  • 余满 ,
  • 卓宝珊 ,
  • 许燕萍 ,
  • 黄龙凯 ,
  • 罗嘉慧 ,
  • 孔艺 ,
  • 周文菁
展开
  • 1.南方科技大学生物医学工程系,广东 深圳 518055;

    2.香港大学深圳医院临床试验中心,广东 深圳 518053;

    3.桂林医科大学人工智能医学院,广西 桂林 541199;

    4.深圳大学生命与海洋科学学院,广东 深圳 518060;

    5.深圳理工大学药学院,广东 深圳 518107

徐磊,男,在读硕士,研究方向:AI医疗大数据,人工智能辅助临床试验

收稿日期: 2025-12-22

  修回日期: 2026-03-04

  录用日期: 2026-03-18

  网络出版日期: 2026-08-17

基金资助

国家药品监督管理局药品审评检查大湾区分中心监管科学研究课题(GBA-JGKX-2501);国家药品监督管理局信息中心课题(20240718001)

Intelligent Classification of Interpretability for Clinical Trial Quality Issue Text Based on Large Language Models

Expand
  • 1.Department of Biomedical Engineering Southern University of Science and Technology Guangdong Shenzhen 518055, China
    2.The University of HongKong-Shenzhen Hospital Clinical Trials Center Guangdong Shenzhen 518053, China
    3.School of Artificial Intelligence in Medicine Guilin Medical University Guangxi Guilin 541199, China
    4.College of Life Sciences and Oceanography Shenzhen University Guangdong Shenzhen 518060, China
    5.Faculty of Pharmaceutical Sciences Shenzhen University of Advanced Technology Guangdong Shenzhen 518107, China

Received date: 2025-12-22

  Revised date: 2026-03-04

  Accepted date: 2026-03-18

  Online published: 2026-08-17

摘要

目的:探索大语言模型识别与分类临床试验质量问题的能力。方法:构建含12个一级指标、67个二级指标的临床试验质量问题文本分类框架,整理1035条问题文本,经数据扩增至5733条。比较Qwen3BertGemma3LoRA微调下的分类性能,引入Shap进行可解释分析。结果:Bert模型表现最优,数据增强后,一、二级指标准确率分别达90.3%81.2%,较增强前提升11.8%36.8%。尽管Bert模型整体表现更优,但在生物样本采集、实验室检查、合并用药等特定场景下,Qwen3Gemma3准确率超过Bert 0.4%~13.4%Shap分析显示,模型决策关键词与人类判断一致。结论:微调后的大语言模型可有效分类临床试验质量问题文本,为解决专业质量控制领域文本分类任务中存在的从业人员的素质与判断标准不一的问题提供可行方案,从而有效辅助人工质量控制流程。


本文引用格式

徐磊 , 余满 , 卓宝珊 , 许燕萍 , 黄龙凯 , 罗嘉慧 , 孔艺 , 周文菁 .

基于大语言模型的临床试验质量问题文本的智能化分类研究

[J]. 中国医药导刊, 2026 , 28(6) : 767 -767-774 . DOI: 10.1009-0959.2026.030014

Abstract

Objective: To investigate the capability of large language models LLMs in recognizing and classifying quality issue texts of clinical trials.Methods: A text classification framework comprising 12 primary and 67 secondary indicators was established. A total of 1035 quality issues texts were collected and augmented to 5733 samples using LLMs. Three models Qwen3 Bert Gemma3 were fine-tuned with LoRA and evaluated with Shap employed for interpretability analysis.Results: Bert achieved the best performance with post-augmentation accuracies of 90.3% for primary and 81.2% for secondary indicators-improvements of 11.8% and 36.8% respectively. Although Bert outperformed overall Qwen3 and Gemma3 exceeded Bert by 0.4%-13.4% in specific scenarios such as biospecimen collection laboratory tests and concomitant medication. Shap analysis confirmed alignment between model decisions and human reasoning.Conclusion: Fine-tuned LLMs can effectively classify clinical trials quality issue texts offering a viable solution to address inconsistencies arising from varying practitioner expertise and judgment standards thereby supporting manual quality control workflows.


参考文献

    [1 国家药品监督管理局.国家药监局 国家卫生健康委关于发布药物临床试验质量管理规范的公告(2020年第57号)[EB/OL.2020-04-23)[2025-11-19.https//www.nmpa.gov.cn/yaopin/ypggtg/20200426162401243.htmltype=pc&m=.

         2  谭琴,邱攀博,李高扬,等.药物临床试验机构质量管理现状调查分析[J.医药导报,20234212):1884-1889.

         3  尚美霞,阎小妍,李雪迎,等.采用多阅片者多病例设计评估AI辅助医疗产品临床试验的样本量估算和应用[J.中国卫生统计,2022391):14-18.

         4  Seufferlein T Ettrich T Stein A. Predicting resistance to first-line FOLFOX plus bevacizumab in metastatic colorectal cancer final results of the multicenter international PERMAD trialJ. J Clin Oncology 2021393_suppl):115.

         5  国家药品监督管理局.国家药监局关于发布人工智能医用软件产品分类界定指导原则的通告(2021年第47号)[EB/OL.2021-07-01)[2025-11-19.https//www.nmpa.gov.cn/ylqx/ylqxggtg/20210708111147171.htmltype=pc&m=.

         6  国家药品监督管理局医疗器械技术审评中心.国家药监局器审中心关于发布人工智能医疗器械注册审查指导原则的通告(2022年第8号)[EB/OL.2022-03-07)[2025-11-19.https//www.cmde.org.cn//xwdt/shpgzgg/gztg/20220309090800158.html.

         7  Jiang Min Zhao Shuhua Mei Yun et al. Real-time risk-based clinical trial quality management in China development of a digital monitoring platformJ.JMIR Med Inform202513e64114.

         8  广东省药学会.药物临床试验 质量管理·广东共识(2025年版)[J.今日药学,2025310):721-728.

         9  张基,梁淑红.药物临床试验质量管理规范培训融入医院药学研究生培养体系的实践与思考[J.中国药物与临床,2025253):160-163.

         10 国家药品监督管理局食品药品审核查验中心.药品注册核查要点与判定原则(药物临床试验)(试行)[EB/OL.2021-12-17)[2026-01-14.https//www.cfdi.org.cn/cfdi/resource/news/14200.html.

         11 Ding Bosheng Qin Chengwei Zhao Ruochen et al. Data augmentation using large language models data perspectives learning paradigms and challengesEB/OL.2024-03-05)[2026-06-06.https//arxiv.org/abs/2403.02990.

         12 Yang An Li Anfeng Yang Baosong et al. Qwen3 Technical ReportEB/OL.2025-05-14)[2026-06-06.https//arxiv.org/abs/2505.09388.

         13 Koroteev MV. BERT A Review of Applications in Natural Language Processing and UnderstandingEB/OL.2021-03-22)[2026-06-06.https//arxiv.org/abs/2103.11943.

         14 Aishwarya Kamath Johan Ferret Shreya Pathak.et al. Gemma 3 Technical ReportEB/OL.2025-03-25)[2026-06-06.https//arxiv.org/abs/2503. 19786.

         15 Hu EJ Shen Yelong Wallis P et al. LoRA Low-rank adaptation of large language modelsEB/OL.2022-04-25)[2026-06-06.https//openreview.net/forumid=nZeVKeeFYf9.

         16 Diederik P. Kingma Jimmy Ba.Adam A Method for Stochastic OptimizationEB/OL.2014-12-22)[2026-06-26.https//arxiv.org/abs/1412. 6980.

         17 Mosca E Szigeti F Tragianni S et al. SHAP-based explanation methods a review for NLP interpretabilityC//Proceedings of the 29th international conference on computational linguistics 20224593-4603.

文章导航

/