Synthesizing Class Labels for Balanced and Highly Imbalanced Cognition Data
| Year of Publication |
2024
|
|---|---|
| Author | |
| Abstract |
In machine learning, assembling labeled training datasets often incurs significant costs and requires expert human annotation, which is essential for effectively training supervised models. A vast majority of newly generated data is unlabeled, a situation that poses significant challenges in fields such as medical diagnosis where precise and timely labeling is critical for the early detection of cognitive conditions. In this work, we employ an unsupervised method in a novel context to synthesize class labels for two cognitive datasets derived from publicly available survey data in the Health and Retirement Study (HRS). To assess the quality and usability of the newly generated labels, we train six supervised classifiers and evaluate their classification performance. Our results indicate that the synthesized labels are of high enough quality and enable effective classifier training. Furthermore, these classifiers outperform a baseline learner on both balanced and highly imbalanced cognition datasets, as evidenced by their area under the precision-recall curve (AUPRC) scores. The results also demonstrate that the synthesized labels yield higher AUPRC scores on balanced data. © 2024 IEEE. |
| DOI |
10.1109/IRI62200.2024.00017
|
| Download citation |