CLAPMamba: CLAP-Guided Mamba for Language-Queried Target Sound Extraction

  • Yang, Jeong Yeol; 
  • Park, Hyun Joon; 
  • Lee, Soyoung; 
  • Han, Sung Won
Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

The extraction of particular sounds from noisy, overlapping mixtures of audio signals remains a significant practical challenge. Universal separation is computationally intractable in modern real-world situations, with query-conditioned extraction emerging as an indispensable targeted approach. Among conditioning signals, natural language exhibits the greatest expressivity, enabling the specification of objects, attributes, and relations without auxiliary references. Although existing contrastive language-audio pre-training (CLAP)-based models enhance language-audio alignment, they typically process hierarchical audio features in a layer-wise manner without modeling their semantic flow explicitly. Thus, their computational cost is identical to that of Transformer blocks. In this paper, we introduce CLAPMamba, a hybrid architecture that enhances conditioning performance while reducing the computational overhead simultaneously. In QueryNet, text embeddings enable conditioning for CLAP audio features, and the Mamba integration stage treats layer outputs sequentially to produce a single, structured conditioning vector that preserves hierarchical semantics. In SeparationNet, a Mamba Forward-Attention Block (MFAB) is used to combine the global context of self-attention with linear-time sequence modeling based on a bidirectional Mamba, capturing long-range dependencies more efficiently. Trained on AudioCaps and evaluated on multiple datasets, including AudioSet, ESC-50, FSDKaggle2018, and MUSIC21, CLAPMamba matches or outperforms Transformer baselines while incurring lower resource costs. Its SDRi and SI-SDRi are higher than those of CLAPSep by 0.25-0.46 and 0.09-0.85 dB, respectively, while using fewer parameters and MACs and exhibiting lower inference latency. Moreover, it increases semantic consistency, as reflected by text-audio similarity and layerwise alignment analyses. Ablation studies indicate complementary improvements based on Mamba-based conditioning in QueryNet and the forward-attention module in SeparationNet. Overall, the method exhibits high generalizability across datasets and sequence lengths, and achieves consistent separation quality with an efficiency suitable for on-device scenarios and latency-sensitive applications.

키워드

Transformers; Computational modeling; Semantics; Computational efficiency; Speech enhancement; Natural languages; Visualization; Vectors; Speech recognition; Solid modeling; Query-conditioned target sound extraction; Mamba; multimodal representation learning; SEPARATION
제목
CLAPMamba: CLAP-Guided Mamba for Language-Queried Target Sound Extraction
저자
Yang, Jeong Yeol; Park, Hyun Joon; Lee, Soyoung; Han, Sung Won
DOI
10.1109/ACCESS.2026.3656219
발행일
2026
유형
Article
저널명
IEEE Access
권
14
페이지
17182 ~ 17195