MTCM: Multi-context temporal consistent modeling for referring video object segmentation

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

1

초록

"Referring Video Object Segmentation (RVOS) focuses on segmenting objects in a video that is based on a provided text description. With recent advancements in transformers, many transformer-based RVOS methods have emerged to enhance interactions between the two modalities. However, these methods often struggle with temporal modeling due to issues with query consistency and limited context awareness. Query inconsistency could result in unstable masks that switch between different objects in the middle of the video, and insufficient context consideration could cause incorrect object segmentation due to a poor alignment with the textual description. To overcome the above challenges, we propose the Multi-context Temporal Consistency Module (MTCM), which integrates an Aligner and a Multi-Context Enhancer (MCE). The Aligner enhances query consistency by filtering out noise and aligning queries, while the MCE selects text-relevant queries through comprehensive context analysis. We applied MTCM to four distinct models, achieving performance improvements across all of them, including a J&F score of 47.6 on the MeViS dataset. The code is available in https://github.com/Choi58/MTCM. © 2025

키워드

Multi-contextReferring video object segmentationTemporal consistency
제목
MTCM: Multi-context temporal consistent modeling for referring video object segmentation
저자
"Choi, Sun-HyukJo, HayoungLee, Seong-Whan
DOI
10.1016/j.neunet.2025.107701
발행일
2025-10
유형
Article
저널명
Neural Networks
190