苏菲的工房

每日医学AI论文简报 v0.3

[日报]

日期: 2026-08-06 来源: PubMed + arXiv | 分析: DeepSeek-v4-flash(PMC HTML解析,含表格+图注) 过滤: 近3天滚动窗口 | 期刊门槛(IF≥3/未收录放行) | 双层相关度(硬规则+LLM显式标准) | PMC全文优先(不足摘要级补齐) | 复合排序


💡 今日学习推荐 CS231n 卷积神经网络(Stanford) — 一、卷积神经网络(CNN)架构(进度0%) ⏭ 下一步:CS229 机器学习讲义(Andrew Ng)


1. TrialTriage, a Semiautonomous Prescreening Workflow for Resolving Ambiguity in Phase I Oncology Trial Eligibility: Development and Proof-of-Concept Study Using Synthetic Cases.

JMIR Form Res (IF: 5.8) | 2026 Aug 5 | PubMed ⚠️ | DOI | ⚠️ 中置信 分析来源: PMC全文
领域: 肿瘤学临床试验患者预筛查、AI辅助电子健康记录数据提取与资格判定
方法: 基于n8n平台构建半自主预筛查工作流TrialTriage,结合大语言模型从自由文本临床叙述和研究者邮件回复中提取变量,并采用确定性规则引擎执行预设的7项标准方案。对缺失或不确定的资格信息触发结构化邮件询问研究者,捕获回复后重新分类。若48小时无回复则转人工审查。使用Claude Sonnet 4.6、Gemini 3.1和Grok 4生成的90个合成患者病例(30例/模型,均衡分布三类结果)进行验证,并由5名独立评审员对其中Claude数据集进行复核。
核心发现: TrialTriage在全部90个合成病例中的分类结果与作者确认的金标准100%一致(95% CI 96.0%-100.0%)。所有模糊病例均正确升级至研究者邮件询问。平均处理时间每30例数据集为2.3(SD 0.5)分钟,约每例3.5-5.5秒。5名评审员平均准确率96.7%(SD 3.3%),Fleiss κ=0.910,评审30例平均耗时9.8(SD 4.8)分钟。在6个首次通过模糊病例的亚组测试中,研究者回复后4例被重新确定性分类,2例仍因回复缺乏可操作信息而保持模糊。
相关度(医学AI研究者): 高度相关,展示了一种将大语言模型提取与规则引擎结合、并通过邮件交互解决资格判定歧义的半自主工作流,对医学AI在临床试验筛查场景中的实际集成和延迟减少有参考价值;但基于合成病例且标签定义与规则引擎设计一致,属于概念验证。
一句话: 该研究通过合成病例验证了TrialTriage半自主预筛查工作流,该系统能在资格数据模糊时自动邮件询问研究者并重新分类,实现100%分类一致性和快速处理,但需注意其概念验证性质及规则与标签同源可能带来的乐观偏倚。

2. Designing clinical AI for patient-centered support beyond the visit: the PACT framework for health systems.

Npj Health Syst (IF: N/A) | 2026 Aug 5 | PubMed ⚠️ | DOI | ⚠️ 中置信 分析来源: PMC全文
领域: 临床人工智能、健康系统、患者中心护理
方法: 观点/框架论文,基于现有临床AI发展现状和护理连续性缺口,提出PACT框架(Patient-centered, AI-enabled Continuity and Timely action),并阐述其操作要素与应用场景
核心发现: 临床AI在就诊内任务(预测、文档、消息生成)进展迅速,但就诊外的随访、交接、沟通断裂仍导致大量可预防失败;提出将临床AI重新定位为支持全流程(诊前、诊中、诊后)连续性、完成度、升级和公平性的健康系统功能,并通过PACT框架明确所有权、沟通渠道、确认规则、升级路径和结局测量等操作要素;以诊后协调为例展示如何整合工作流、监测和分级人工支持
相关度(医学AI研究者): 高,为设计超越单次就诊的临床AI系统提供了可操作框架,并强调从任务性能转向运维随访和系统级责任
一句话: 论文提出PACT框架,主张临床AI应作为健康系统功能来保障患者全程连续性护理和及时行动,而非仅优化就诊内任务。

3. Accuracy, Usefulness, and Impact Variability of ChatGPT-4 for COPD Medication Management: A Modified Delphi Study.

Chronic Obstr Pulm Dis (IF: N/A) | 2026 Aug 4 | PubMed ⚠️ | DOI | ⚠️ 中置信 分析来源: 摘要

领域: 慢性阻塞性肺疾病(COPD)药物治疗管理;大语言模型(LLM)在临床决策支持中的应用

方法: 在单次会话中,将5个COPD治疗问题同时输入三台电脑的ChatGPT-4.0,共生成15条回答。三位经过住院药师培训、委员会认证的临床药师采用三轮改良德尔菲法,对每条回答在准确性、有用性和影响变异性三个维度上进行3点量表(0–2分)评分。共识定义为三位评审者完全一致。

核心发现: 共45项“回答-维度”评分中,40项(88.9%)达成共识。准确性评分范围从差(0)到好(2),有用性从有些用(1)到非常有用(2),影响变异性从低(0)到高(2)。针对稳定性COPD药物治疗的一个问题,三条同时生成的回答均引用了已退休的临床实践指南,导致准确性评为差。针对治疗急性加重和处理复杂病例的两个问题,各有一条回答的有用性高于同问题的其他回答。

相关度(医学AI研究者): 该研究提示即使在同一会话中向ChatGPT-4.0输入相同COPD用药问题,其回答在准确性和潜在临床影响上存在实质性差异;强调LLM输出应作为临床判断的辅助而非替代,并需结合现行指南和本地监督进行部署。

一句话: 同一ChatGPT-4.0会话中对相同COPD用药问题的并行回答在准确性和临床影响方面存在显著不一致。


🔬 新增AI临床试验

近1日新增: 67 项 | 数据源: ClinicalTrials.gov

NCT07747233 Precision Study of CLAiR AI Software

状态: COMPLETED | 期: NA | 赞助: | 招募: 56 条件: Cardiovascular Risk

NCT05675410 A Study to Compare Standard Therapy to Treat Hodgkin Lymphoma to the Use of Two Drugs, Brentuximab Vedotin and Nivolumab

状态: RECRUITING | 期: PHASE3 | 赞助: | 招募: 1875 条件: Lugano Classification Limited Stage Hodgkin Lymphoma AJCC v8

NCT07475026 A Study of Neoadjuvant Tislelizumab Plus Lenvatinib in Resectable HCC at High Risk of Recurrence

状态: RECRUITING | 期: PHASE3 | 赞助: | 招募: 198 条件: HCC - Hepatocellular Carcinoma

NCT06730113 Youth-Onset Type 2 Diabetes and Heart Disease: The Young at Heart Prospective Cohort Study

状态: RECRUITING | 期: N/A | 赞助: | 招募: 930 条件: Obesity; Type 2 Diabetes

NCT07476092 Evaluation of the Effect of Digital-based Games on the Visual and Cognitive Performance of Young Children With Intellectual Disabilities

状态: RECRUITING | 期: NA | 赞助: | 招募: 60 条件: Intellectual Disability


📄 医学AI预印本速览(arXiv 近3天)

1. ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Yang Yang, Qinyu Zhao, Mouxiang Chen et al. | 2026-08-04 | arXiv Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixed computat…

2. SocietyBench: Forecasting Counterfactual Social-World Evolution Zhenran Wang, Zhonghan Bian, Jinsong Li et al. | 2026-08-04 | arXiv Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task – fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social eve…

3. WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament Zhenran Wang, Zhonghan Bian, Jinsong Li et al. | 2026-08-04 | arXiv Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the…

4. Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Mohsen Hariri, Weicong Chen, Nahal Shahini et al. | 2026-08-04 | arXiv Large language models can solve substantially harder reasoning problems with more inference-time compute. The term “test-time scaling,” however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample completed candidates and aggregate t…

5. Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation Seyed Kahaki, Shijie Li, Weijie Chen et al. | 2026-08-04 | arXiv Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications. This work investigates and addresses limitati…

6. Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss? Hailong Jiang, Feng Yu, Emran Hossain et al. | 2026-08-04 | arXiv Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether large language models (LLMs) can recover such semantics from heterogeneous C/C++ context and realize them as validated, contract-…

7. ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Jinhe Bi, Chennan Zhou, Zengjie Jin et al. | 2026-08-04 | arXiv On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-…

8. HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification Salah Eddine Bekhouche, Abdessalam Bouchekif, Hichem Telli et al. | 2026-08-04 | arXiv Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hallucination resources often assign a binary label to an entire response, indicating whether it is hallucinated or non-hallucinate…