Confident Labels: A Novel Approach to New Class Labeling and Evaluation on Highly Imbalanced Data Confident Labels: A Novel Approach to New Class Labeling and Evaluation on Highly Imbalanced Data 1 of 1 Confident Labels: A Novel Approach to New Class Lab

Year of Conference
2024
Author
Conference Name
36th International Conference on Tools with Artificial Intelligence
Number of Pages
232-239
Abstract

A common challenge in machine learning is obtaining readily available labeled data, as the majority of data remains unlabeled and thereby necessitates human annotation from domain experts, e.g. doctors. Thus, substantial cost is associated with labeling the data, e.g. healthcare diagnostics. The efficient and effective acquisition of labeled data is crucial for early detection, specifically, in relation to cognitive decline. In our work, we employ a novel unsupervised method to generate binary class labels on a publicly available, highly imbalanced cognition dataset derived from the Health and Retirement Study (HRS). To evaluate the efficacy of our newly generated class labels, we also employ a novel approach to evaluate the labels by directly comparing them to ground-truth labels, rather than the traditional approach which measures the performance of a supervised model trained on the generated labels. Our results demonstrate that the newly generated class labels significantly outperform the baseline method, IF, in terms of Balanced Accuracy (BA), Geometric Mean (GM), F-Measure (F1), and Matthew's Correlation Coefficient (MCC).

DOI
10.1109/ICTAI62512.2024.00042
Download citation