669Identifiées −38 doublons 631Uniques −352 exclues 61Retenues au tri 57 à lire 4Lues 4 à décider 0Incluses

Bibliothèque / fiche n°76

Zero-shot video highlight detection based on text descriptions and synthetic images

Byra et al.arXiv2026 exclu au tri

Article PDF

Cet article n'a pas de fiche

Statut : exclu au tri — Détection de moments forts (highlights) dans des vidéos, hors sujet.. Seuls les articles retenus au tri et dont le PDF est accessible sont lus en entier.

Résumé des auteurs

Detecting video highlights, the most informative or engaging moments in a video, is important for applications such as video summarization and content recommendation. We propose a zero-shot framework that combines CLIP, large language models (LLMs), and diffusion models. Given lightweight video metadata, such as a title or category, an LLM generates textual descriptions of likely highlight events. These descriptions are further converted into synthetic visual prototypes using a diffusion model. Textual and visual representations are matched to video frames using CLIP, enabling frame-level highlight detection without highlight annotations or dataset-specific training. Experiments on TVSum and SumMe demonstrate strong zero-shot performance, with particularly favorable results on TVSum. The proposed approach provides an effective framework for metadata-conditioned zero-shot video highlight detection.