상세 보기
A Gated Attention-Based Semi-Supervised Learning Framework for Label-Scarce Tabular Data
- Jang, Yongeun;
- Jung, Yonggon;
- Oh, Ghil-Geun;
- Baek, Jun-Geol
WEB OF SCIENCE
0SCOPUS
0초록
Tabular data is a data format commonly used in various industries, including finance, medical services, and manufacturing. However, deep learning research for this data type has lagged behind the progress seen in computer vision and natural language processing. Real-world applications often face severe label scarcity due to high labeling costs. Traditional Gradient Boosting Decision Trees (GBDTs) and existing deep learning architectures struggle to achieve high predictive performance under these constraints. This study proposes the Gated Attention Tabular Encoder (GATE) and a two-stage semi-supervised learning pipeline. GATE effectively captures complex feature interactions by using global context from self-attention to dynamically regulate the flow of feature information. The proposed pipeline uses a teacher-student self-distillation framework to extract robust representations from unlabeled data. Then, it employs SoftNCA-based metric learning alongside high-confidence pseudo-labeling to optimize the embedding space with minimal labeled instances. Experimental results on public benchmark data showed that the proposed framework consistently achieved superior average performance across all label-ratio scenarios. The framework achieved an average accuracy of 62.23% with only 1% of the labeled data. It also achieved the highest average accuracy of 89.26% in a fully supervised setting, confirming its universal competitiveness. Through additional experiments, we confirmed that the proposed pipeline constructs a structurally superior representation space and achieves greater class separation compared to supervised learning models. In addition, in terms of computational efficiency, the proposed pipeline achieved this performance by using only 50% of the training parameters and 33% of the FLOPs compared to the comparative transformer-based model, FT-Transformer, and using only one-hundredth of the number of parameters compared to complex SSL models such as SAINT. This study presents a highly effective and computationally efficient paradigm for learning tabular representations in label-deficient industrial environments. The code is available at https://github.com/yongeun-jang/GATE
키워드
- 제목
- A Gated Attention-Based Semi-Supervised Learning Framework for Label-Scarce Tabular Data
- 저자
- Jang, Yongeun; Jung, Yonggon; Oh, Ghil-Geun; Baek, Jun-Geol
- 발행일
- 2026
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 90319 ~ 90333