QuanDA: Quantile-Based Discriminant Analysis for High-Dimensional Imbalanced Classification

Qian Tang (University of Minnesota) · Yuwen Gu (University of Connecticut) · Boxiang Wang (University of Iowa)
binary classificationclass imbalancecost-sensitive classifiershigh-dimensional settingsimbalanced classeslow-sample-sizeminority class detectionnoisy featuresoverfittingpredictive performancequantile regressionquantile-based discriminant analysissimulation studiestheoretical analysisultra-high dimensional

Binary classification with imbalanced classes is a common and fundamental task, where standard machine learning methods often struggle to provide reliable predictive performance. Although numerous approaches have been proposed to address this issue, classification in low-sample-size and high-dimensional settings still remains particularly challenging. The abundance of noisy features in high-dimensional data limits the effectiveness of classical methods due to overfitting, and the minority class is even difficult to detect because of its severe underrepresentation with low sample size. To address this challenge, we introduce Quantile-based Discriminant Analysis (QuanDA), which builds upon a novel connection with quantile regression and naturally accounts for class imbalance through appropriately chosen quantile levels. We provide comprehensive theoretical analysis to validate QuanDA in ultra-high dimensional settings. Through extensive simulation studies and high-dimensional benchmark data analysis, we demonstrate that QuanDA overall outperforms existing classification methods for imbalanced data, including cost-sensitive large-margin classifiers, random forests, and SMOTE.