苏菲的工房

每日医学AI论文简报 v0.3

[日报]

日期: 2026-08-12 来源: PubMed + arXiv | 分析: DeepSeek-v4-flash(PMC HTML解析,含表格+图注) 过滤: 近3天滚动窗口 | 期刊门槛(IF≥3/未收录放行) | 去重+状态追踪(PMID/DOI/标题) | 双层相关度(硬规则+LLM显式标准) | 研究设计识别 | PMC全文优先(不足摘要级补齐) | 复合排序


💡 今日学习推荐 CS231n 卷积神经网络(Stanford) — 一、卷积神经网络(CNN)架构(进度0%) ⏭ 下一步:CS229 机器学习讲义(Andrew Ng)


1. Benchmarking large language models for question answering on German clinical practice guidelines. 🆕

BMJ Health Care Inform (IF: 93.6) | 2026 Aug 11 | PubMed ⚠️ | DOI | ✅ 高置信 | 研究设计: Guideline 分析来源: PMC全文
领域: 医学人工智能 / 临床决策支持 / 大语言模型与检索增强生成(RAG)/ 德国临床指南问答
方法: 作者开发并验证了cpgQA-DE基准数据集,包含200道多项选择题,来源于10部德国现行临床实践指南,覆盖5个专科。所有题目均经专家审核,并标注指南来源、专科、相关性和难度。研究构建了基于德国指南语料的RAG问答流水线,评估多种大语言模型与检索器组合,并对RAG集成前后的准确率进行比较。
核心发现: RAG可显著提升所有测试模型的问答准确率。最佳配置为GPT-5结合multilingual-e5-large检索器,准确率达95%;开放权重模型gpt-oss-120b在RAG支持下达到90%准确率。该基准支持德语医疗环境下临床指南问答系统的可重复离线评估,并显示开放权重模型在隐私敏感场景中具有本地部署潜力。
相关度(医学AI研究者): 高。该研究填补了德语临床实践指南问答基准缺失的空白,为评估指南感知型大语言模型提供了可复现的方法,并系统证明了RAG对提升模型循证回答能力的重要性,对医学AI落地中的知识更新、可追溯性和隐私保护均有参考价值。
一句话: 该研究构建了基于10部德国临床指南的200题专家验证基准cpgQA-DE,证明RAG可显著提升LLM的指南问答准确率,最佳配置达95%,开放权重模型也可达90%。


🔬 新增AI临床试验

近1日新增: 64 项 | 数据源: ClinicalTrials.gov

NCT06887582 Crestal Sinus Lifting in Periodontally-Compromised Patients Utilizing Autologous Dentin Graft

状态: COMPLETED | 期: NA | 赞助: | 招募: 44 条件: Maxillary Sinus Disease; Alveolar Bone Loss

NCT03629743 Correlate of Surface Electroencephalogram (EEG) With Implanted EEG Recordings (ECOG)

状态: COMPLETED | 期: N/A | 赞助: | 招募: 20 条件: Anesthesia

NCT07751939 Artificial Intelligence Model to Predict Chronic Kidney Outcomes in a Vietnamese Cohort

状态: ACTIVE_NOT_RECRUITING | 期: N/A | 赞助: | 招募: 1182 条件: Chronic Kidney Disease

NCT07756853 Visual Attention, Object Recognition, and Eye Movements in Healthy Volunteers

状态: NOT_YET_RECRUITING | 期: N/A | 赞助: | 招募: 40 条件: Healthy Volunteer

NCT07326709 A Study to Investigate the Efficacy, Safety and Tolerability of Votoplam in Participants With Huntington’s Disease

状态: RECRUITING | 期: PHASE3 | 赞助: | 招募: 770 条件: Huntington Disease


📄 医学AI预印本速览(arXiv 近3天)

1. MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Changhao Xiang, Shangyu Xing, Zhen Wu et al. | 2026-08-11 | arXiv Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle …

2. Scheduling Mixed RL Rollouts Beyond Prefix Locality Zetao Hong, Song Yuan, Yuanhao Ding et al. | 2026-08-11 | arXiv Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it doe…

3. The Illusion of Cross-Lingual Safety in Low-Resource Languages Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay et al. | 2026-08-11 | arXiv Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual …

4. Attention-Path Fragility as an Uncertainty Signal in Large Language Models Minsoo Kim, Sungyoung Ji, Kisung Moon et al. | 2026-08-11 | arXiv We propose that a model’s uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We instantiate this as ASMI (Attention-Subnetwork Mutual …

5. 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment Alam Noor, Luis Almeida, Mohamed Daoudi | 2026-08-11 | arXiv Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically proven Sheep Pain Facial Expr…

6. TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification Jian Zhang, Zhuohao Yang, Songlin Lei et al. | 2026-08-11 | arXiv Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on large language models (LLMs) struggle to be efficiently applied to this task due to issues like lengthy prompt…

7. myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR Ye Kyaw Thu, Ye Bhone Lin, Thura Aung et al. | 2026-08-11 | arXiv Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by native speak…

8. SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training Zhuang Wang | 2026-08-11 | arXiv In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem l…