插补(统计学)
计算机科学
辍学(神经网络)
负二项分布
算法
数据挖掘
统计
缺少数据
数学
机器学习
泊松分布
作者
Zhun Miao,Xinyi Lin,Jiaqi Li,Ka Ho Mok,Qiuchen Meng,Xuegong Zhang
摘要
ABSTRACT High‐throughput single‐cell RNA‐seq (scRNA‐seq) data contains an excess of zero values, which can be contributed by unexpressed genes and detection signal dropouts. Existing imputation methods fail to distinguish between these two types of zeros. In this study, we introduce a statistical framework that effectively differentiates true zeros (lack of expression) from false zeros (dropouts). By focusing only on imputing the dropout zeros, we developed a new imputation tool, scRecover. Our approach utilizes a zero‐inflated negative binomial framework to model the gene expression of each gene in each cell, enabling the estimation of zero‐dropout probability. Additionally, we employ a modified version of the Good and Toulmin model to identify true zeros for each gene. To achieve imputation, scRecover is combined with other imputation methods such as scImpute, SAVER and MAGIC. Down‐sampling experiments show that it recovers dropout zeros with higher accuracy and avoids over‐imputing true zero values. Experiments conducted on real world data highlight the ability of scRecover to enhance downstream analysis and visualization.
科研通智能强力驱动
Strongly Powered by AbleSci AI