随机森林
计算机科学
数据挖掘
滤波器(信号处理)
背景(考古学)
水准点(测量)
机器学习
人工智能
特征(语言学)
独立性(概率论)
统计
数学
地理
语言学
哲学
考古
计算机视觉
大地测量学
出处
期刊:Current Bioinformatics
[Bentham Science]
日期:2022-02-22
卷期号:17 (4): 344-357
被引量:13
标识
DOI:10.2174/1574893617666220221120618
摘要
Background: Various feature (variable) screening approaches have been proposed in the past decade to mitigate the impact of ultra-high dimensionality in classification and regression problems, including filter based methods such as sure independence screening, and wrapper based methods such as random forest. However, the former type of methods rely heavily on strong modelling assumptions while the latter ones requires an adequate sample size to make the data speak for themselves. These requirements can seldom be met in biochemical studies in cases where we have only access to ultra-high dimensional data with a complex structure and a small number of observations. Objective: In this research, we want to investigate the possibility of combining both filter based screening methods and random forest based screening methods in the regression context. Method: We have combined four state-of-art filter approaches, namely, sure independence screening (SIS), robust rank correlation based screening (RRCS), high dimensional ordinary least squares projection (HOLP) and a model free sure independence screening procedure based on the distance correlation (DCSIS) from the statistical community with a random forest based Boruta screening method from the machine learning community for regression problems. Result: Among all the combined methods, RF-DCSIS performs better than the other methods in terms of screening accuracy and prediction capability on the simulated scenarios and real benchmark datasets. Conclusion: By empirical study from both extensive simulation and real data, we have shown that both filter based screening and random forest based screening have their pros and cons, while a combination of both may lead to a better feature screening result and prediction capability.
科研通智能强力驱动
Strongly Powered by AbleSci AI