상세 보기
MTCM: Multi-context temporal consistent modeling for referring video object segmentation
- "Choi, Sun-Hyuk;
- Jo, Hayoung;
- Lee, Seong-Whan
WEB OF SCIENCE
2SCOPUS
1초록
"Referring Video Object Segmentation (RVOS) focuses on segmenting objects in a video that is based on a provided text description. With recent advancements in transformers, many transformer-based RVOS methods have emerged to enhance interactions between the two modalities. However, these methods often struggle with temporal modeling due to issues with query consistency and limited context awareness. Query inconsistency could result in unstable masks that switch between different objects in the middle of the video, and insufficient context consideration could cause incorrect object segmentation due to a poor alignment with the textual description. To overcome the above challenges, we propose the Multi-context Temporal Consistency Module (MTCM), which integrates an Aligner and a Multi-Context Enhancer (MCE). The Aligner enhances query consistency by filtering out noise and aligning queries, while the MCE selects text-relevant queries through comprehensive context analysis. We applied MTCM to four distinct models, achieving performance improvements across all of them, including a J&F score of 47.6 on the MeViS dataset. The code is available in https://github.com/Choi58/MTCM. © 2025
키워드
- 제목
- MTCM: Multi-context temporal consistent modeling for referring video object segmentation
- 저자
- "Choi, Sun-Hyuk; Jo, Hayoung; Lee, Seong-Whan
- 발행일
- 2025-10
- 유형
- Article
- 저널명
- Neural Networks
- 권
- 190