主页 文献库文献详情
PMID: 42021407 已发表 · epublish 英语

Machine learning predictions surpass individual mRNAs as a proxy of single-cell protein expression.

Genome biology ·第 27 卷 ·第 1 期 ·2026-04-22

Fisher J, Wood O, Bullers S, Murray L, Li L, Jackson-Wood MA

摘要

Expansive repositories of scRNA-seq data are now available. These are often analysed assuming that mRNA abundance reflects expression of the cognate protein. However, post-transcriptional/translational regulation and the sparsity of measurements in single-cell data make mRNA an inadequate proxy for protein. Methods to quantify surface proteins alongside scRNA-seq exist but are less widely adopted. Machine learning approaches for protein imputation from scRNA-seq data have been published, which learn transcriptome-wide patterns that predict protein expression where data for both is available. These models can then be applied to infer surface protein expression on scRNA-seq only data sets, increasing their utility. We test 9 machine learning methods for predicting single-cell protein expression, comparing the accuracy between methods and compared to using cognate mRNAs alone. Overall, machine learning -based protein predictions across methods outperform direct inference from mRNAs, including cases where proteins absent by mRNA are successfully predicted by the wider transcriptome. When comparing models trained on restricted cell types and across different datasets/tissues, we find that the overlap in cell type composition of training and test data is an important determinant of prediction accuracy. We also compare computational resource requirements to guide method selection. These results reiterate that single-cell mRNA abundance is not a reliable proxy of cognate protein expression and that whole-transcriptome based imputations can improve upon them given appropriately trained models. However, limitations to the generalisability of these methods persist, notably a requirement for highly similar training data, which may limit the current scope of applications.

关键词
CITE-seq Learning Machine Prediction Protein Proteomics Single-cell Transcriptomics
文献信息
期刊
Genome biology
期刊简称
Genome Biol
ISSN
1474-760X
发表日期
2026-04-22
语言
英语
国家/地区
England
NLM ID
100960660
分析服务
分析服务

联系地址

山东省济南市章丘区文博路2号

齐鲁师范学院 genelibs生信实验室

山东省济南市高新区舜华路750号

大学科技园北区F座4单元2楼

电话: 0531-88819269

微信公众号

关注微信订阅号,实时查看信息,关注医学生物学动态。


商务邮箱

E-mail: product@genelibs.com