Bibliothèque / fiche n°117
Fusion Mechanism for Text Detection and Tracking in Video Scenes
Cet article n'a pas de fiche
Statut : exclu au tri — filtre mots-clés : aucun terme du groupe « origine ». Seuls les articles retenus au tri et dont le PDF est accessible sont lus en entier.
Résumé des auteurs
In video scene text detection and tracking,challenges such as motion blur, occlusion, and target deformation often lead to unstable detection results and tracking drift. To address these issues, we propose a joint optimization framework that integrates text detection and tracking with temporal modeling to enhance robustness. Specifically, (1) a temporal information fusion method is designed to ensure consistent text representation across frames; (2) a cross-task learning strategy reuses semantic features from detection for trajectory matching while dynamically refining detection thresholds with tracking feedback. Experimental results on the ICDAR2015 Video Text Dataset indicate the effectiveness of our approach: it improves MOTA by 2.8 % and IDF1 by 1.7 % compared with GoMatching, achieving superior performance with a favorable balance between accuracy and robustness.