苏菲的工房

每日医学AI论文简报 v0.3

[日报]

日期: 2026-08-14 来源: PubMed + arXiv | 分析: DeepSeek-v4-flash(PMC HTML解析,含表格+图注) 过滤: 近3天滚动窗口 | 期刊门槛(IF≥3/未收录放行) | 去重+状态追踪(PMID/DOI/标题) | 双层相关度(硬规则+LLM显式标准) | 研究设计识别 | PMC全文优先(不足摘要级补齐) | 复合排序


💡 今日学习推荐 CS231n 卷积神经网络(Stanford) — 一、卷积神经网络(CNN)架构(进度0%) ⏭ 下一步:CS229 机器学习讲义(Andrew Ng)


1. Preparing for the Augmented Intelligences: A Roadmap for Physicians in the Era of Artificial Intelligence and Robotics. 🆕

J Am Geriatr Soc (IF: N/A) | 2026 Aug 14 | PubMed ⚠️ | DOI | ⚠️ 中置信 | 研究设计: 未识别 分析来源: 摘要

领域: 老年医学、人工智能与机器人临床应用、医师角色转型

方法: 专家评论/框架性论述,围绕信息收集、数据分析、治疗实施三大临床领域提出医师角色转变路径,并借鉴Beers Criteria提出临床AI工具认证建议。

核心发现: 医师角色并非被取代,而是转变为信息收集的监督者、数据分析的批判性评估者与伦理仲裁者、治疗实施的协调者。AI应作为增强智能,其输出需保持可查询、有来源、从属于负责医师,同时需警惕训练数据偏见、数字年龄歧视、大语言模型幻觉和自动化偏差。

相关度(医学AI研究者): 高。该文为老年复杂共病场景下的AI与机器人整合提供了清晰的医师角色框架,并指出关键风险与治理机制(如AI认证标准),对AI工具的设计、评估和临床落地具有直接参考价值。

一句话: 老年医学医师应通过“增强智能”原则,从执行者转向监督者、评估者和协调者,以引领AI与机器人时代的临床转型。

2. Preventing Participant Leakage in Reissued Clinical AI Benchmarks: A Biomedical Informatics Evaluation Protocol and DAIC-WOZ/E-DAIC Audit. 🆕

Methods Inf Med (IF: N/A) | 2026 Aug 14 | PubMed ⚠️ | DOI | ⚠️ 中置信 | 研究设计: 未识别 分析来源: 摘要
领域: 生物医学信息学/临床人工智能基准测试,具体为抑郁症筛查(PHQ-8)相关的语音/文本数据基准(DAIC-WOZ / E-DAIC)
方法: 对发布谱系与参与者身份进行审计,使用持久标识符、内容哈希、PHQ-8标签核对、折间转移交叉表;对比存在泄漏与参与者不重叠的E-DAIC到DAIC-WOZ协议,在相同保留参与者上使用配对bootstrap推断;构建无泄漏参考基线。
核心发现: 所有189名DAIC-WOZ参与者均以字节完全相同的录音和相同PHQ-8总分出现在E-DAIC中;官方分区中104/189参与者折叠归属发生变化。用E-DAIC训练、在DAIC-WOZ测试集评估时,47/47测试参与者均出现在模型开发池中,相反方向则干净。有泄漏的声学协议AUROC为0.797,去重后为0.569,配对Delta AUROC +0.227(95%CI +0.106至+0.372;双侧p<0.001);匹配训练规模时Delta AUROC为+0.249。通用视觉分类器在混合外部集上AUROC为0.887,但在未见过的仅AI部分上为0.562。无泄漏参考基线约为AUROC 0.60。
相关度(医学AI研究者): 高。该研究揭示了临床AI基准重用中“继任版本参与者泄漏”这一可预防的评估失败,直接影响外部验证的可信度与模型比较的公平性,医学AI研究者在复用公开基准时需执行参与者身份验证、标签核对和去重。
一句话: 继任发布中的参与者泄漏会严重高估临床AI模型性能,应在评测前通过谱系审计与参与者去重来避免。


🔬 新增AI临床试验

近1日新增: 64 项 | 数据源: ClinicalTrials.gov

NCT05372627 NHLBI-Emory Advanced Cardiac CT Reconstruction

状态: NOT_YET_RECRUITING | 期: N/A | 赞助: | 招募: 1000 条件: Structural Heart Disease

NCT06893939 Limited Versus Extended Trophic Feeding (LET-FEED) Trial

状态: RECRUITING | 期: PHASE2, PHASE3 | 赞助: | 招募: 350 条件: Sepsis; Length of Stay; Mortality

NCT07766252 CCHW-delivered CFH-based Cancer Prevention Program Among Chinese Americans

状态: COMPLETED | 期: NA | 赞助: | 招募: 1129 条件: Cancer

NCT07572136 Anti-CRLF2-R/TSLPR Chimeric Antigen Receptor T Cells (TSLPR-CART) in Participants With Recurrent or Refractory CRLF2-R/TSLPR-Overexpressing B-Cell Acute Lymphoblastic Leukemia (B-ALL)

状态: NOT_YET_RECRUITING | 期: PHASE1 | 赞助: | 招募: 57 条件: B-All; Acute Lymphoblastic Leukemia

NCT07766694 Evaluating the Preferences and Tradeoffs of AI-based Electronic Consultations for Older Adults in Primary Care

状态: NOT_YET_RECRUITING | 期: NA | 赞助: | 招募: 220 条件: Primary Care Patients; Primary Care Provider


📄 医学AI预印本速览(arXiv 近3天)

1. SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization Weihan Meng, Hongzhu Guo, Yi Jing et al. | 2026-08-13 | arXiv Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavio…

2. Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology Yunsung Chung, Yingshuo Liu, Abboud F. Hassan et al. | 2026-08-13 | arXiv Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint. However, recovery after a procedure often unfolds as an irregular trajectory: clinical observations, medication changes, repeat interventions, a…

3. DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina et al. | 2026-08-13 | arXiv Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning…

4. MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination Saisha Shetty, Satvik Tripathi, Austin Lin et al. | 2026-08-13 | arXiv We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, an…

5. AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models Mohammed Ayman Habib, Rylan Hart, Morteza Fayazi | 2026-08-13 | arXiv Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that relies heavily on expert intuition. Among recent developments, LLMs have introduced a promising approach by bringing natural language reasoning to circuit design tas…

6. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification Daniel Perkins, John Squires, Janou Milligan et al. | 2026-08-13 | arXiv Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a m…

7. CAPRI: Contract-Aware Proof Repair for Isabelle Jim Woodcock, Gabriel Leite, Augusto Sampaio et al. | 2026-08-13 | arXiv We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an LLM changed only what the developer authorised. We present CAPRI, a contract-aware repair workflow in which Is…

8. LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles Md Wasiul Haque, Sagar Dasgupta, Mizanur Rahman et al. | 2026-08-13 | arXiv Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming exploitability requires execu…