669Identifiées −38 doublons 631Uniques −352 exclues 61Retenues au tri 57 à lire 4Lues 4 à décider 0Incluses

Bibliothèque / fiche n°546

Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection

Thuau et al.arXiv2026 exclu au tri

Article PDF

Cet article n'a pas de fiche

Statut : exclu au tri — Détection d'anomalies dans les vidéos de surveillance, hors sujet.. Seuls les articles retenus au tri et dont le PDF est accessible sont lus en entier.

Résumé des auteurs

How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, weakly labeled, and resource-constrained? Most weakly supervised VAD methods assume centralized training; recent VLM-based extensions further rely on dense inference, generated explanations, or additional adaptation. We introduce a lightweight federated MIL-VLM cascade in which only a compact MIL scorer is trained across clients, while a frozen VLM verifies high-scoring suspect segments post hoc. We study two VLM feedback interfaces: parsed text-generation decisions and a logit-based interface that extracts a continuous anomaly score from next-token Yes/No probabilities. Experiments on UCF-Crime with InternVL3.5-2B and Qwen3-VL-2B-Instruct show that text-generation verification can improve frame-level AUC after diagnostic temporal post-processing, but remains sensitive to prompts, parsers, model choice, and smoothing. In contrast, the logit interface provides a fixed parser-free signal that improves both frame-level AUC and frame-level AP over the MIL baseline across both VLMs, without temporal post-processing in its main configuration. Since suspect segments are updated independently once available, next-token logit feedback provides a simple segment-local alternative to text-generation verification.