苏菲的工房

每日医学AI论文简报 v0.3

[日报]

日期: 2026-08-08 来源: PubMed + arXiv | 分析: DeepSeek-v4-flash(PMC HTML解析,含表格+图注) 过滤: 近3天滚动窗口 | 期刊门槛(IF≥3/未收录放行) | 去重+状态追踪(PMID/DOI/标题) | 双层相关度(硬规则+LLM显式标准) | 研究设计识别 | PMC全文优先(不足摘要级补齐) | 复合排序


💡 今日学习推荐 CS231n 卷积神经网络(Stanford) — 一、卷积神经网络(CNN)架构(进度0%) ⏭ 下一步:CS229 机器学习讲义(Andrew Ng)


1. An autonomous multimodal AI agent for evidence-grounded ophthalmic diagnosis. 🆕

Cell Rep Med (IF: 45.5) | 2026 Aug 5 | PubMed ⚠️ | DOI | ✅ 高置信 | 研究设计: Guideline 分析来源: 期刊全文(隧道下载)
领域: 眼科人工智能(多模态AI诊断)
方法: 开发并评估AgentEYE,一种自主多模态AI智能体,整合眼底照相和B超图像,路由至专用分析工具,检索指南/网络证据,并生成可追溯的诊断报告。内部基准测试302例,比较AgentEYE与仅LLM基线、无专用影像工具消融、无检索消融的性能;另由3名眼科医生对200例进行盲法评估,比较诊断正确性、完整性、安全性和引文依据;并进行外部数据集分析。
核心发现: AgentEYE在内部基准中诊断正确性和完整性优于仅LLM基线和无专用影像工具消融;与无检索消融性能相似,表明检索主要增强证据依据和引文可审计性。盲法评估确认AgentEYE在诊断正确性、完整性、安全性和引文依据方面优于仅LLM自引基线。外部分析显示性能呈分布依赖性。研究支持AgentEYE作为可追溯的决策支持原型,需前瞻性多中心验证。
相关度(医学AI研究者): 高。展示了多模态AI智能体在眼科诊断中整合影像工具与证据检索的可行性和优势,强调可审计性和安全性,为医学AI从单任务模型向自主智能体发展提供范例,并指出前瞻性验证的必要性。

2. Clinical Trial Source Document Verification Using a Large Language Model. 🆕

JACC Adv (IF: N/A) | 2026 Aug 5 | PubMed ⚠️ | DOI | ⚠️ 中置信 | 研究设计: 未识别 分析来源: PMC全文
领域: 临床研究监查 / 人工智能辅助源数据核验(多中心临床试验中的既往病史数据验证)
方法: 基于大型语言模型(LLM,OpenAI o4-mini)构建源文件核验(SDV)工作流,将电子病例报告表(eCRF)中的24项既往病史变量与原始医疗记录(PDF转文本,使用Apple Vision OCR)配对,由LLM判定每项eCRF记录是否得到病历支持,并与人工审阅者结果比较一致性。数据来自INVESTED试验的293例受试者(随机抽取约5%)。
核心发现: LLM源文件核验与参考审阅者的一致率达96%,接近人工审阅水平,提示可自动化SDV并降低成本、提前标记无依据数据。
相关度(医学AI研究者): 高,展示了LLM在临床试验监查与源数据验证中的实际应用潜力,可显著减少人力成本,为医学AI在临床研究运营中的落地提供直接证据。
一句话: 大语言模型可自动核验多中心试验中受试者既往病史数据,与人工审阅结果96%一致,有望大幅降低临床试验监查成本。

3. A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation. 🆕

J Med Internet Res (IF: 5.8) | 2026 Aug 7 | PubMed ⚠️ | DOI | ⚠️ 中置信 | 研究设计: 未识别 分析来源: PMC全文
领域: 医学影像学、放射学诊断、多模态大语言模型评估
方法: 构建双语放射学基准RadM-Bench,包含720例病例(360例英文公开教学病例和360例中文常规临床病例),覆盖9个放射亚专科。评估4个专有模型(如GPT-4o、Gemini等)和6个开源模型(如Qwen2.5-VL、InternVL3等),在4种输入条件下测试:仅临床病史、病史+放射科医生选择的2D关键图像、病史+2帧/秒容积数据、病史+10帧/秒容积数据。采用0-3分四级诊断质量评分,由2名放射科医生盲评,并使用跨语言对照实验区分语言与临床内容的影响。
核心发现: 所有10个模型在两个数据集上的平均诊断质量评分均低于1.5(0-3分制)。添加放射科医生选择的2D关键图像后的具体性能变化结果在提供的摘要文本中未完整展示。
相关度(医学AI研究者): 高,该基准直接评估了当前多模态大语言模型在放射学诊断实践中的真实能力,对模型选择、临床部署及后续优化具有重要参考意义。
一句话: 本研究构建了RadM-Bench双语放射学基准,系统评估10个多模态大语言模型,发现其平均诊断质量评分均低于1.5分(满分3分),提示模型在放射学诊断任务中性能仍然有限。

4. Theoretical Exploration of Error Thresholds for Clinical AI Decision Support in Nursing: Exploratory Simulation Study Grounded in Human-AI Reliance Data. 🆕

JMIR Nurs (IF: 5.8) | 2026 Aug 5 | PubMed ⚠️ | DOI | ⚠️ 中置信 | 研究设计: 未识别 分析来源: PMC全文
领域: 护理信息学、临床AI决策支持、人类-AI协作、错误阈值
方法: 基于3项独立随机实验(N=3502)的9个实证数据点,构建并校准线性依赖模型(自变量包括AI准确性A、临床经验E、任务复杂性C),采用加权最小二乘法拟合;使用2000次自举(bootstrap)量化校准不确定性;在27种因素组合下模拟预测错误率;参照已发表的大型语言模型(LLM)在复杂临床任务上的准确性基准,评估模型运行范围。
核心发现: 校准后β_A为0.201(自举95%CI 0.023-0.234;P(β_A>0)>0.99)。对于新手护士×高复杂度任务,将错误率控制在<10%所需最低AI准确率为0.89(95%CI 0.88-0.90),<20%所需为0.78(95%CI 0.75-0.79)。当通用LLM在复杂临床任务上的准确率约为0.5-0.7时,模型预测高风险条件下错误率约为26%-41%。作者指出,该模型基于非护理依赖数据校准,阈值属于模型输出而非护理来源的经验标准;未来需要在护理情境中进行行为验证。
相关度(医学AI研究者): 高。该研究为护理领域AI决策支持系统设定安全准确性阈值提供了决策理论框架,提示当前通用LLM在高复杂度护理场景中可能未达到足够低的错误率目标,对AI模型部署、风险评估和阈值设计具有参考价值。
一句话: 文章通过模拟模型提示:新手护士在高复杂度任务中使用AI辅助决策时,若要将错误率控制在10%以下,AI准确率需至少达到约0.89,而当前通用LLM在复杂临床任务中可能难以稳定达到该水平。

5. Development of SCOPE: Social and Clinical Opioid Use Disorder Predisposition Evaluation. 🆕

J Addict Med (IF: N/A) | 2026 Aug 5 | PubMed ⚠️ | DOI | ⚠️ 中置信 | 研究设计: 未识别 分析来源: 摘要

领域: 物质使用障碍(阿片类药物使用障碍,OUD)的风险预测与临床决策辅助

方法: 基于“All of Us Research Program”的参与者数据,采用多变量逻辑回归识别与OUD相关的社会健康决定因素(SDoH)、人口学和临床预测因子;以参与者完成SDoH调查后≥90天诊断为OUD作为结局变量;基于模型计算的优势比构建评分式风险计算器SCOPE;通过AUROC、敏感性和特异性评估模型性能,并使用10折交叉验证。

核心发现: 共纳入624,431名参与者,其中2475名(0.39%)在调查后>90天被诊断为OUD。SCOPE模型在10折交叉验证中的平均AUROC为0.85(SD 0.011),特异性78%(SD 0.16%),敏感性78%(SD 2.42%)。以评分8作为高危截断值时,敏感性为0.79(95% CI: 0.76-0.83),特异性为0.76(95% CI: 0.75-0.76),阳性似然比为3.35,阴性似然比为0.26。

相关度(医学AI研究者): 高。该研究提供了一个基于多维度数据(含SDoH)的OUD风险预测工具,展示了如何利用大规模真实世界数据构建临床决策支持模型,且量化了模型性能,为个性化风险识别和早期干预提供了可操作的方法学参考。

一句话: 基于大规模真实世界数据开发的SCOPE风险计算器,可有效预测阿片类药物使用障碍风险,AUROC达到0.85,具有较好的临床决策辅助潜力。


🔬 新增AI临床试验

近1日新增: 78 项 | 数据源: ClinicalTrials.gov

NCT03351764 Development of Non-Invasive Brain Stimulation Techniques

状态: COMPLETED | 期: NA | 赞助: | 招募: 83 条件: Normal Physiology

NCT05998616 Feasibility of Remote Exercise Training for Hispanics/Latinos With MS

状态: TERMINATED | 期: NA | 赞助: | 招募: 33 条件: Multiple Sclerosis

NCT06498674 AI-Assisted Ultrasound Review Before Thyroid Surgery

状态: COMPLETED | 期: NA | 赞助: | 招募: 515 条件: Thyroid Cancer; Thyroid Nodule

NCT02535702 Development Of Neuroimaging Methods To Assess The Neurobiology Of Addiction

状态: RECRUITING | 期: NA | 赞助: | 招募: 192 条件: Normal Physiology

NCT07162714 SCRT Followed by AK112 in pMMR/MSS Mid-low Rectal Cancer

状态: RECRUITING | 期: PHASE2 | 赞助: | 招募: 30 条件: Rectal Cancer; Radiation; AK112


📄 医学AI预印本速览(arXiv 近3天)

1. Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi et al. | 2026-08-06 | arXiv Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists’ workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fr…

2. RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer Xinye Wang, Junxiao Liu, Shujian Huang | 2026-08-06 | arXiv Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promising paradigm, providing dense token-level supervision on stu…

3. Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents Tao Wang, Qihao Yang, Rongjiao Liang et al. | 2026-08-06 | arXiv Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, hig…

4. Does FLAIR super-resolution erase or hallucinate small white-matter lesions? Zahra Khodakarami, Yue Li, Pulkit Khandelwal et al. | 2026-08-06 | arXiv White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascular pathology and neurodegeneration. FLAIR is usually acquired with thick slices in clinical settings, giving it poor through-plane resolution.…

5. On-Policy Self-Distillation without Any Supervision Yijiang Li, Bingyang Wang, Yijun Liang et al. | 2026-08-06 | arXiv On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and …

6. QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction Mutasim Fuad Sarker, Adiba Rahman Namira, Wafa Binte Alam et al. | 2026-08-06 | arXiv Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries derived from early admissio…

7. NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering Jonas Gann, Michael Gertz | 2026-08-06 | arXiv Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. However, its reasoning process remains largely opaque: intermediate reasoning steps are difficult to verify and cannot be reliabl…

8. Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints Omid Bazgir, Md Nasir, Jacob Hoffman et al. | 2026-08-06 | arXiv Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaki…