New Class Labeling and Evaluation Methodology for Balanced and Highly Imbalanced Data
| Year of Publication |
2024
|
|---|---|
| Author | |
| Abstract |
Acquiring labeled data in machine learning is challenging due to the high costs and need for domain expertise and human annotation, leaving much of the data unlabeled. In the healthcare domain, the prompt acquisition of labeled data is important for the early diagnosis of diseases such as dementia. Our research introduces an unsupervised method for generating binary class labels on two publicly available cognitive datasets, one balanced and one highly imbalanced (3% minority), derived from the Health and Retirement Study (HRS). Additionally, we employ an innovative evaluation approach that directly assesses the accuracy of the newly generated class labels, bypassing the need to train a supervised model to evaluate traditional performance metrics. This allows for an accurate quantification of the efficacy and reliability of the generated labels. Through this evaluation framework, we demonstrate the potential of our method to produce high-quality labels. Upon evaluation, our newly generated class labels significantly outperform our baseline method, Isolation Forest. © 2024 IEEE. |
| DOI |
10.1109/ICMLA61862.2024.00042
|
| Download citation |