상세 보기
CLAPMamba: CLAP-Guided Mamba for Language-Queried Target Sound Extraction
- Yang, Jeong Yeol;
- Park, Hyun Joon;
- Lee, Soyoung;
- Han, Sung Won
WEB OF SCIENCE
1SCOPUS
1초록
The extraction of particular sounds from noisy, overlapping mixtures of audio signals remains a significant practical challenge. Universal separation is computationally intractable in modern real-world situations, with query-conditioned extraction emerging as an indispensable targeted approach. Among conditioning signals, natural language exhibits the greatest expressivity, enabling the specification of objects, attributes, and relations without auxiliary references. Although existing contrastive language-audio pre-training (CLAP)-based models enhance language-audio alignment, they typically process hierarchical audio features in a layer-wise manner without modeling their semantic flow explicitly. Thus, their computational cost is identical to that of Transformer blocks. In this paper, we introduce CLAPMamba, a hybrid architecture that enhances conditioning performance while reducing the computational overhead simultaneously. In QueryNet, text embeddings enable conditioning for CLAP audio features, and the Mamba integration stage treats layer outputs sequentially to produce a single, structured conditioning vector that preserves hierarchical semantics. In SeparationNet, a Mamba Forward-Attention Block (MFAB) is used to combine the global context of self-attention with linear-time sequence modeling based on a bidirectional Mamba, capturing long-range dependencies more efficiently. Trained on AudioCaps and evaluated on multiple datasets, including AudioSet, ESC-50, FSDKaggle2018, and MUSIC21, CLAPMamba matches or outperforms Transformer baselines while incurring lower resource costs. Its SDRi and SI-SDRi are higher than those of CLAPSep by 0.25-0.46 and 0.09-0.85 dB, respectively, while using fewer parameters and MACs and exhibiting lower inference latency. Moreover, it increases semantic consistency, as reflected by text-audio similarity and layerwise alignment analyses. Ablation studies indicate complementary improvements based on Mamba-based conditioning in QueryNet and the forward-attention module in SeparationNet. Overall, the method exhibits high generalizability across datasets and sequence lengths, and achieves consistent separation quality with an efficiency suitable for on-device scenarios and latency-sensitive applications.
키워드
- 제목
- CLAPMamba: CLAP-Guided Mamba for Language-Queried Target Sound Extraction
- 저자
- Yang, Jeong Yeol; Park, Hyun Joon; Lee, Soyoung; Han, Sung Won
- 발행일
- 2026
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 17182 ~ 17195