三篇 Nature 论文看 SHAP 的三重边界:归因、机制解释与因果识别
上一篇推文( SHAP解释边界:从“模型归因”到“因果识别” )中,我从 epistemic classification 的角度,将 SHAP 的解释划分为三个层次: model-based attribution、process-informed interpretati…
- Reference
- notes:n18
- Published
- 2026.04.10
- Series
- note
- Source
- Source ↗
上一篇推文( SHAP解释边界:从“模型归因”到“因果识别” )中,我从 epistemic classification 的角度,将 SHAP 的解释划分为三个层次: model-based attribution、process-informed interpretation 与 causality-informed identification 。 这篇文章尝试将这一分析框架放到具体论文之中,分析 SHAP 或 Shapley values 在各自证据链中究竟承担了什么 epistemic role。
一、 归因:当 SHAP 被严格限定为 model-based attribution
(Chen et al., Nature, 2026: Gene regulatory landscape dissected by single-cell four-omics sequencing)

图 1|SHAP 用于模型归因与调控元件排序(Chen et al., Nature , 2026) 。
这篇论文的核心任务,是在单细胞四组学数据框架下建立基因表达预测模型,并据此识别候选 enhancer–promoter linkages。文章使用 Shapley values 评估不同 genomic bins 对表达预测的贡献,并将预测结果与实验验证的 enhancer–promoter pairs 进行外部比较。
在这里,SHAP 的功能是清楚的: 它分解的是模型输出,而不是生物过程本身 。更准确地说,它回答的是:在当前训练好的预测模型中,哪些输入区域对表达预测更重要,哪些 bins 对某个 enhancer–promoter linkage prediction 贡献更大。它所揭示的是模型依赖的表示结构,而不是调控机制本体本身。论文中关于调控关系的讨论,也并不是由 SHAP 单独支撑,而是建立在 multi-omics integration、3D chromatin structure 和外部验证等多重证据之上。显然,这篇论文中的 SHAP 属于 model-based attribution 。它可以支持 contribution quantification 和 feature prioritization ,但不能单独推出“某个元件就是调控机制本身”。当然,作者也没有把这种 attribution 直接写成 mechanism identification。
二、 机制:SHAP 支持 process-informed interpretation
(Delavaux et al., Nature, 2023: Native diversity buffers against severity of non-native tree invasions)

图 2|SHAP 用于变量重要性与 机制解释( Delavaux et al., Nature , 202 3)。
这篇论文的核心任务,是分析生态变量、人类活动与非本地树种入侵强度之间的关系。作者使用 random forest 模型评估变量重要性,同时结合 GLM 检验关系的显著性与方向。文中明确说明,random forest 主要用于 variable importance 和 visualization。图中展示的 SHAP 值,正是在这一 predictive modeling 框架中用来表征变量贡献和 response pattern 的。
因此,从方法本身看,这里的 SHAP 仍然属于 predictive modeling 之后的 post-hoc attribution,而不是 causal inference。它首先做的,仍然是 attribution:告诉我们在模型中,native diversity、distance to ports 等变量对 invasion presence 或 invasion severity 的预测贡献有多大。这篇论文并没有停留在“模型如何使用变量”这一层,而是进一步把这些 attribution 结果嵌入到生态学理论语言中,例如 biotic resistance、niche filling 等机制性框架。这一步,就是 process-informed interpretation 。这里的 SHAP 不再只是描述 model behavior,而是被用来支持一种 theory-consistent reading :模型归因结果与既有生态机制是否一致,是否为这些机制提供额外支持。
三、因果:SHAP 嵌入 causality-informed identification
(Liu et al., Nature Plants, 2025: When and where soil dryness matters to ecosystem photosynthesis)

图 3| 因果结构 先于归因:regime-dependent causal framework(Liu et al., Nature Plants , 2025 ) 。

图 4|SHAP 嵌入因果框架的贡献分解(Liu et al., Nature Plants , 2025)。
与前两篇不同,不是先做普通 SHAP,再在讨论里附会因果语言,而是从方法设计一开始就把问题设定为一个因果问题。文章明确提出使用 causality-guided explainable AI framework ,并通过 regime distinction、causal chain graphs 以及干预语义来分析 soil moisture、VPD 等变量对 GPP 的作用。原文还特别对比了普通 SHAP 与 causal Shapley:前者依赖 predictor independence assumption,而后者借助 causal chain graphs 和 expert domain knowledge 来增强 causal interpretation。
这篇论文确实已经进入 causality-informed identification 的层级,但关键在于:这里的因果性并不是来自 SHAP 本身,而是由外加的 causal commitments 预先赋予的。真正承担识别负担的,是 regime distinction 是否合理,causal chain graphs 是否可信,intervention semantics 是否定义清楚,以及 confounding 是否得到了结构化处理。原文甚至主动承认,其 water-limited condition 下的 causal treatment 并不完美,因为 TA 对 VPD 的直接作用难以被 chain graph 完整表达;作者认为这一限制更可能高估 VPD 的重要性,而不会改变其主要结论。恰恰是这一点,使这篇论文在方法论上比很多“把 correlation 直接写成 causation”的论文更自觉:它没有把因果性偷渡给 SHAP,而是把因果前提显式写出来,并同时承担这些前提不成立时的解释风险。
写在最后
通过上述三篇论文可以看到,SHAP 在不同研究中的角色存在本质差异:
在 model-based attribution 层面,SHAP 是对模型输出的分解工具; 在 process-informed interpretation 层面,SHAP 支持 mechanism-consistent explanation,但不构成 mechanism identification; 在 causality-informed identification 层面,SHAP 仅在明确的 causal framework 下作为表达工具参与分析,而不构成因果性的来源。

References
Chen, Y., et al. (2026). Gene regulatory landscape dissected by single-cell four-omics sequencing . Nature.
Delavaux, C. S., et al. (2023). Native diversity buffers against severity of non-native tree invasions . Nature, 621, 773–781.
Liu, J., et al. (2025). When and where soil dryness matters to ecosystem photosynthesis . Nature Plants, 11, 1390–1400.