Bibliothèque / fiche n°104
Violent Video Detection Based on Multimodal Video-Text Fusion
Cet article n'a pas de fiche
Statut : exclu au tri — filtre mots-clés : aucun terme du groupe « origine ». Seuls les articles retenus au tri et dont le PDF est accessible sont lus en entier.
Résumé des auteurs
This paper proposes a violent video detection method based on video-text multimodal fusion, which makes up for the semantic and perceptual limitations of single visual features in complex scenes. The method extracts keyframes via a frame difference-based sampling strategy, captures video spatiotemporal features with unfrozen EfficientNet and temporal attention (TA) module, generates video captions through a captioning model, and deeply fuses visual and text features by the Video Text Fusion (VTF) module. Fused features are input into a classifier for binary prediction. Experiments on three public datasets verify that the method significantly improves the detection accuracy and robustness of the model.