Published 2024-10-30
How to Cite

This work is licensed under a Creative Commons Attribution 4.0 International License.
Abstract
This paper proposes a unified framework integrating semi-supervised learning and robust optimization for low-resource and high-noise data mining scenarios. This framework aims to improve discriminative reliability and training stability under conditions of scarce annotations and potential contamination of labels and features. The method uses a small number of highly reliable labeled samples as supervisory anchors and fully utilizes the structural information of a large number of unlabeled samples. Consistency constraints are used to strengthen representation invariance and reduce prediction drift caused by data augmentation and noise perturbation. To suppress the accumulation of pseudo-supervision errors and confirmation bias, a confidence gating strategy is introduced to filter pseudo-labels, incorporating only highly reliable pseudo-supervision signals into the optimization. Simultaneously, sample-level robust weighting and parameter regularization are employed to reduce the interference of abnormal samples and noise gradients on model updates, thereby achieving selective absorption of unlabeled information and controllable suppression of noise influence. The overall objective is composed of a combination of supervisory loss, consistency loss, and gated pseudo-supervision loss. The training process remains simple and reproducible in an end-to-end setting. Comparative experiments show that the proposed method outperforms multiple baselines and achieves better performance in terms of both discriminative accuracy and probabilistic reliability.