669Identifiées −38 doublons 631Uniques −352 exclues 61Retenues au tri 57 à lire 4Lues 4 à décider 0Incluses

Bibliothèque / fiche n°104

Violent Video Detection Based on Multimodal Video-Text Fusion

Li et Chenvenue non précisée2026 exclu au tri

Article

Cet article n'a pas de fiche

Statut : exclu au tri — filtre mots-clés : aucun terme du groupe « origine ». Seuls les articles retenus au tri et dont le PDF est accessible sont lus en entier.

Résumé des auteurs

This paper proposes a violent video detection method based on video-text multimodal fusion, which makes up for the semantic and perceptual limitations of single visual features in complex scenes. The method extracts keyframes via a frame difference-based sampling strategy, captures video spatiotemporal features with unfrozen EfficientNet and temporal attention (TA) module, generates video captions through a captioning model, and deeply fuses visual and text features by the Video Text Fusion (VTF) module. Fused features are input into a classifier for binary prediction. Experiments on three public datasets verify that the method significantly improves the detection accuracy and robustness of the model.