Loading…

Accurate Markov Boundary Discovery for Causal Feature Selection

Causal feature selection has achieved much attention in recent years, which discovers a Markov boundary (MB) of the class attribute. The MB of the class attribute implies local causal relations between the class attribute and the features, thus leading to more interpretable and robust prediction mod...

Full description

Saved in:
Bibliographic Details
Published in:IEEE transactions on cybernetics 2020-12, Vol.50 (12), p.4983-4996
Main Authors: Wu, Xingyu, Jiang, Bingbing, Yu, Kui, Miao, chunyan, Chen, Huanhuan
Format: Article
Language:English
Subjects:
Citations: Items that this one cites
Items that cite this one
Online Access:Get full text
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:Causal feature selection has achieved much attention in recent years, which discovers a Markov boundary (MB) of the class attribute. The MB of the class attribute implies local causal relations between the class attribute and the features, thus leading to more interpretable and robust prediction models than the features selected by the traditional feature selection algorithms. Many causal feature selection methods have been proposed, and almost all of them employ conditional independence (CI) tests to identify MBs. However, many datasets from real-world applications may suffer from incorrect CI tests due to noise or small-sized samples, resulting in lower MB discovery accuracy for these existing algorithms. To tackle this issue, in this article, we first introduce a new concept of PCMasking to explain a type of incorrect CI tests in the MB discovery, then propose a cross-check and complement MB discovery (CCMB) algorithm to repair this type of incorrect CI tests for accurate MB discovery. To improve the efficiency of CCMB, we further design a pipeline machine-based CCMB (PM-CCMB) algorithm. Using benchmark Bayesian network datasets, the experiments demonstrate that both CCMB and PM-CCMB achieve significant improvements on the MB discovery accuracy compared with the existing methods, and PM-CCMB further improves the computational efficiency. The empirical study in the real-world datasets validates the effectiveness of CCMB and PM-CCMB against the state-of-the-art causal and traditional feature selection algorithms.
ISSN:2168-2267
2168-2275
DOI:10.1109/TCYB.2019.2940509