主页 文献库文献详情
PMID: 42296381 已发表 · ppublish 英语

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

Bioinformatics (Oxford, England) ·第 42 卷 ·第 7 期 ·2026-07-02

Kendiukhov(I)

摘要

Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96 892 ablation-derived edges, 80 191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2× enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1× above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33 301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A→B" each connected by at least one ablation edge in both models; 10.6× enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31 176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)→target pairs are enriched 2.06× (Fisher OR 5.84), markedly higher than 1.12× against TRRUST; direct ChIP-seq-supported target pairs show 10-30× larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations ρ≈0. Bootstrap and per-cell-type stability (N∈{50,100,200}; B cell, CD4 + T, macrophage) give Pearson r≥0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

文献信息
期刊
Bioinformatics (Oxford, England)
期刊简称
Bioinformatics
ISSN
1367-4811
发表日期
2026-07-02
语言
英语
国家/地区
England
NLM ID
9808944
分析服务
分析服务

联系地址

山东省济南市章丘区文博路2号

齐鲁师范学院 genelibs生信实验室

山东省济南市高新区舜华路750号

大学科技园北区F座4单元2楼

电话: 0531-88819269

微信公众号

关注微信订阅号,实时查看信息,关注医学生物学动态。


商务邮箱

E-mail: product@genelibs.com