669Identifiées −38 doublons 631Uniques −352 exclues 61Retenues au tri 57 à lire 4Lues 4 à décider 0Incluses

Diagramme de flux PRISMA 2020

Calculé à partir du journal des décisions. Il se met à jour à chaque veille et à chaque article gardé ou rejeté.

Notices identifiées (n = 669)Hugging Face Papers : 212arXiv : 211OpenAlex : 146Hugging Face Daily : 100Notices retirées avant le tri (n = 38)doublons : 38Notices triées (n = 413)en attente de tri : 218Notices exclues (n = 352)filtre par mots-clés : 333tri titre + résumé : 19Rapports recherchés (n = 61)en attente de lecture : 57Rapports non récupérés (n = 0)PDF indisponibleRapports évalués pour l'éligibilité (n = 4)en attente de ta décision : 4Rapports exclus (n = 0)Études incluses dans la revue (n = 0)IdentificationTriInclusion

Critères d'éligibilité

Inclusion

I1
Traite de la détection, l'attribution ou la localisation de vidéos générées ou manipulées par IA
I2
Propose une méthode, un benchmark / dataset, une étude d'évaluation ou une revue (survey)
I3
Publié à partir de 2022, rédigé en anglais

Exclusion

E1
Images fixes uniquement, aucune expérience sur vidéo
E2
Génération de vidéo sans volet détection
E3
Audio seul ou texte seul
E4
Aucune évaluation empirique (opinion, position paper)
E5
N'est pas un travail de recherche (tutoriel, annonce, brevet)
E6
Hors sujet

Modifiables dans les Réglages. Un changement de critères en cours de revue doit être justifié dans le mémoire.

Stratégie de recherche

SourceRequêteExécutionsNotices renvoyéesDernière exécutionÉtat
Hugging Face Daily(flux quotidien, sans requête)2 2002026-10-03 19:48 ok
Hugging Face PapersAI-generated video benchmark2 802026-10-03 19:48 ok
Hugging Face PapersAI-generated video detection2 802026-10-03 19:48 ok
Hugging Face Papersdeepfake video detection generalization2 802026-10-03 19:48 ok
Hugging Face Papersdiffusion video detection generalization2 802026-10-03 19:48 ok
Hugging Face Papersgenerated video forensics2 802026-10-03 19:48 ok
Hugging Face Paperssynthetic video detection2 802026-10-03 19:48 ok
Hugging Face Paperstext-to-video detection2 802026-10-03 19:48 ok
OpenAlexAI-generated video benchmark2 502026-10-03 19:48 ok
OpenAlexAI-generated video detection2 502026-10-03 19:48 ok
OpenAlexdeepfake video detection generalization2 502026-10-03 19:48 ok
OpenAlexdiffusion video detection generalization2 502026-10-03 19:48 ok
OpenAlexgenerated video forensics2 502026-10-03 19:48 ok
OpenAlexsynthetic video detection2 502026-10-03 19:48 ok
OpenAlextext-to-video detection2 502026-10-03 19:48 ok
Semantic ScholarAI-generated video detection2 02026-10-03 19:48 dernière erreur : HTTPStatusError: Client error '429 ' for
arXivAI-generated video benchmark1 402026-10-03 19:48 ok
arXivAI-generated video detection2 802026-10-03 19:47 ok
arXivdeepfake video detection generalization1 402026-10-03 19:48 ok
arXivdiffusion video detection generalization1 402026-10-03 19:48 ok
arXivgenerated video forensics2 402026-10-03 19:47 dernière erreur : ReadTimeout: The read operation timed ou
arXivsynthetic video detection2 402026-10-03 19:47 dernière erreur : ReadTimeout: The read operation timed ou
arXivtext-to-video detection2 402026-10-03 19:48 dernière erreur : HTTPStatusError: Client error '429 Too M

Articles exclus 352

Un article exclu à tort ? Ouvre-le et clique sur Garder : l'écart avec tes critères sera signalé dans les incohérences.

ArticleÉtapeCritèreMotif
Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
REMEDY: How Far Is Video Generation from Medical Education World Models?exclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understandingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
InternW0: A Foundational Physical World Model for Efficient Real-World Interactionsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflowsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalitiesexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agentsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecastingexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identificationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Companion-style QA Assistance in Ego-Visionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
From Visual Cues to Spoken Narration: Rethinking Audio Descriptionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstructionexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coachingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenesexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interactionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Frameworkexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understandingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
BreastMammo and DenseMammo: Benchmarks for Mammography Domain Generalizationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Predictionexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agentsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecastingexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
CAS2UML: A Handwritten Sketch-to-PlantUML Dataset for Class and Activity Diagramsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
MedForge-RSI: Medical Deepfake Detection via Recursive Self-Improvementexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arriveexclu au tri E4Position paper sans évaluation empirique directe d'une nouvelle méthode de détection.
Towards Generalizable Deepfake Image Detection with Vision Transformersexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modelingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Two-Stage Teacher-Student Reliable Prior Learning for Robust Underwater Image Enhancementexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Secret-Stego Dissimilarity as a Design Axis: Invertible Coverless Image Steganography with Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactionsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Enhancing Prostate Cancer Segmentation for Multi-Domain Generalization using a novel Parallel-Route Coherent Mixup Regularization Trainingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
High-Fidelity Synthetic Transmission Electron Microscopy Image Generation Using Diffusion Probabilistic Models for Data-Limited Semiconductor Metrologyexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
A Surface-based Multimodal Framework for Multitask Analysis in Alzheimer's Diseaseexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Q-ARVD: Quantizing Autoregressive Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
SCMAPR: Self-Correcting Multi-Agent Prompt Refinement for Complex-Scenario Text-to-Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristicsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancementexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggersexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interactionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwenexclu au tri E6Traite de la détection d'anomalies vidéo (crimes), pas de l'IA.
PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Back-Tracking from Clarity: Self-Learning to See Text from Afarexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speechexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Smartphone-Based Method for Automated Speed Enforcementexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Detecting Phone-Induced Pedestrian Distraction via a Multimodal Fusion Transformerexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Lesion-centered 3D mapping of colonoscopy procedures: validation of a hierarchical ensemble pipeline on public benchmark videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Online Video Agent Harness for Long Video Understandingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Detecting and Explaining Fake News Short Videos with Multimodal Content and Real-World Evidenceexclu au tri E6Détection de fake news vidéo générales, sans focus sur l'IA.
OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Modelsexclu au tri E6Détection d'hallucinations dans les MLLM, hors sujet.
Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detectionexclu au tri E6Détection d'anomalies dans les vidéos de surveillance, hors sujet.
NOVA: Normal-Side Modeling for Training-Free Zero-Shot Video Anomaly Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRIexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
A Data-Driven Multimodal Method for Early Detection of Coordinated Abnormal Behaviors in Live-Streaming Platformsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Training-Free Temporal Abstraction for General Video Understandingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmakuexclu au tri E6Détection de fake news vidéo via commentaires, hors sujet.
AEGIS: Preventing Cross-Domain Resource Abuse in MCPexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agentexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attributionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
A Multi-Domain Feature Fusion Framework for Generalizable Deepfake Detection Across Different Generatorsexclu au tri E1Concerne les images générées par IA, aucune expérience sur la vidéo.
Bridging the Modality Gap in Forensic Image Retrievalexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Synthetic-to-Real Transfer in Cerebral Microbleed Generation and Segmentationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Memory- and Bandwidth-Efficient SPAD-LiDAR Ranging via Coarse-to-Fine Spline Sketchingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Learned Compression of SAR Phase-History Data: A Rate-Honest Feasibility Study on GOTCHAexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
AutoRASOR: Autonomous Rapid Scanning Electron Microscope Operatorexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Conditioning noise is a free regularizer for LoRA fine-tuning: no pathology encoder required for diffusion-based artifact detection in histopathologyexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Spatial Attention Supervision for Defect Localization: Exploiting Ground-Truth Masks as Training Signal in Diffusion-Augmented Defect Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Videoexclu au tri E6Détection et localisation d'accidents de la route, hors sujet.
PERSIST: Persistent-State Discrimination for Shot Boundary Detectionexclu au tri E6Détection de transitions de plans dans les vidéos, hors sujet.
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-trainingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
SweepLSD: A One-Pass, O(width)-Memory Line Segment Detector with an Integer-Only Streaming Core and a Real-Time FPGA Realizationexclu au tri E6Détecteur de segments de lignes dans les images, hors sujet.
Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cellsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detectionexclu au tri E3Concerne uniquement la détection de deepfakes audio.
$\texttt{DisMorph}$: learning to disentangle technical distortions from true biological changeexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Accurate Localization of Road Traffic Objects on the Road Plane Using Surveillance Camera Imageryexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Synthetic data generation framework for quality control automation in gravure printingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detectionexclu au tri E6Étudie le benchmark de vidéos synthétiques pour la détection de travailleurs sous charge suspendue.
SCI-Mamba: Unsupervised Learning based Low-Light Image Enhancement for Non-Cooperative Spacecraftexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fogexclu au tri E6Se concentre sur la détection et le suivi de drones sous brouillard synthétique.
TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
MemLife: Curating and Reasoning over Long-Term Egocentric Video Memoriesexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetingsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Rubric Rewards from Item Response Theoryexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Prompted Identity Degrades Cooperation in Multi-Agent LLM Systemsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Aligning One-Step Generative Models with Reward-Weighted Transport Distillationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloningexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
BiasReducer: Adaptive Bias Mitigation for Reward Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
BIABench: Evaluating AI agents on real-world bioimage analysis tasksexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
DAGent: Evaluate-then-Grow Planning for Deep Research Agentsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolutionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformersexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Training LLM Judges from Language Feedback via Position-Selective Self-Distillationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Scaling and Distilling Text Embeddings for Better Diffusibilityexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RLexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
AutoGUIWorld: Image Generators as Visual World Models for GUI Agentexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
AutoDataBench: A Data-centric Testbed for Accelerating Auto Researchexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecastsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Rules to Tools: Executable Checks for LLM Agents in Scientific Computingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Hierarchical Continuous Diffusion Language Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesisexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
It Takes Workflows to Evolve Better Workflowsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoningexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
RPTune: Learned Context Curation for LLM Catalog Searchexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Align Then Reason: A Multimodal Lip-Sync Judge for Dubbingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief Statesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimizationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
ROWBench: Do Video Models Render What the Program Specifies?exclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agentsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agentsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
4Director: Controlling Video World Models with Rigid 3D Geometryexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Sharpening Tax in Post-Trainingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomyexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Agent Priors-guided Policy Learningexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Beyond the Current Scene: Event-Referential Grasping with Active View Selectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Controlled Decoding Attacks on Black-Box LLMsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
World Observer: Joint Actor-Observer Generation for Persistent World Modelingexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Personalized Image Generation with Reasoning and Reflectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructionsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
FlexRouter: Learning Complementary Model Sets for Flexible LLM Routingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Reviewexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimizationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflowsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RLexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spacesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?exclu au tri E6filtre mots-clés : aucun terme du groupe « video »
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Smaller Models, Better Rejects: Preference Distillation Scalingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Memorizon: Training World Models Beyond Their Context Windowexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Samplingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestrationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewardsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfindingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelinesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interactionexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokensexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learningexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
LOCI: Spatial Linear Memory for Streaming World Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plansexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transportexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
JevSpawn: Adaptive Agentic Inference through Compositional Action Spacesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Decoding Looped Transformers Better for (Almost) Freeexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Researchexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamicsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineersexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes Itexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Honeycomb: Constant-Size Scene Memory Representation for Video World Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalizationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectoriesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Video Generation Models: A Survey of Post-Training and Alignmentexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Persona Dosing: Calibrated Activation Steering for Graded Trait Controlexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimizationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestrationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
VMBench: A Benchmark for Perception-Aligned Video Motion Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative Systemexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
VI-Bench: Benchmarking Prompt Inversion from AIGC Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
VISTA: A Generative Egocentric Video Framework for Daily Assistanceexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?exclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AIexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effectsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AIexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
GAIA: Rethinking Action Quality Assessment for AI-Generated Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
RefCaptioner: Multi-Reference Image-Grounded Video Captioningexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restorationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
FilmBench: A Film-Grade Benchmark for Cinematic Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation Metricexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Benchmarking AIGC Video Quality Assessment: A Dataset and Unified Modelexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
HumanScore: Benchmarking Human Motions in Generated Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screeningexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Evaluating Deepfake Detectors in the Wildexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensemblesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Diffusion Deepfakeexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
ED^4: Explicit Data-level Debiasing for Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
The Deepfake Detection Challenge (DFDC) Preview Datasetexclu au tri I3publié avant 2022
SeeABLE: Soft Discrepancies and Bounded Contrastive Learning for Exposing Deepfakesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
AuthGuard: Generalizable Deepfake Detection via Language Guidanceexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Penny-Wise and Pound-Foolish in Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Combining EfficientNet and Vision Transformers for Video Deepfake Detectionexclu au tri I3publié avant 2022
UCF: Uncovering Common Features for Generalizable Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
The DeepFake Detection Challenge (DFDC) Datasetexclu au tri I3publié avant 2022
Adversarially robust deepfake media detection using fused convolutional neural network predictionsexclu au tri I3publié avant 2022
Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Robust AI-Generated Face Detection with Imbalanced Dataexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Diffusion Action Segmentationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
TPDiff: Temporal Pyramid Video Diffusion Modelexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
FreeInit: Bridging Initialization Gap in Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Efficient Video Diffusion Models: Advancements and Challengesexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transferexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
HARIVO: Harnessing Text-to-Image Models for Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Visual Bridge: Universal Visual Perception Representations Generatingexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Point Prompting: Counterfactual Tracking with Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Diffusion Models for Video Prediction and Infillingexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Generative Video Mattingexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
APLA: Additional Perturbation for Latent Noise with Adversarial Training Enables Consistencyexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Diffusion Classifiers Understand Compositionality, but Conditions Applyexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Diffusion Models in Low-Level Vision: A Surveyexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realismexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
LLM-grounded Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latentsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Diffusion Noise Feature: Accurate and Fast Generated Image Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Studying Image Diffusion Features for Zero-Shot Video Object Segmentationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
SinFusion: Training Diffusion Models on a Single Image or Videoexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Diffusion Models as Masked Autoencodersexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Optical-Flow Guided Prompt Optimization for Coherent Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Video Diffusion Alignment via Reward Gradientsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
VidGen-1M: A Large-Scale Dataset for Text-to-video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Learning Video Representations from Textual Web Supervisionexclu au tri I3publié avant 2022
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrievalexclu au tri I3publié avant 2022
Video and Text Matching with Conditioned Embeddingsexclu au tri I3publié avant 2022
MuSeD: A Multimodal Spanish Dataset for Sexism Detection in Social Media Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Text Detection and Recognition in the Wild: A Reviewexclu au tri I3publié avant 2022
Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
A Survey on Multimodal Disinformation Detectionexclu au tri I3publié avant 2022
Unveiling Hallucination in Text, Image, Video, and Audio Foundation Models: A Comprehensive Surveyexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Seeing, Signing, and Saying: A Vision-Language Model-Assisted Pipeline for Sign Language Data Acquisition and Curation from Social Mediaexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Combating Online Misinformation Videos: Characterization, Detection, and Future Directionsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Perception Test: A Diagnostic Benchmark for Multimodal Video Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Helping Hands: An Object-Aware Ego-Centric Video Recognition Modelexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
TubeDETR: Spatio-Temporal Video Grounding with Transformersexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answeringexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Formatexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoningexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Saliency-Guided DETR for Moment Retrieval and Highlight Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Trackingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Text-guided Fine-Grained Video Anomaly Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
DrawVideo: Generating Long Video from Storyboard Keyframe Sketchesexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Dynamic texture analysis for detecting fake faces in video sequencesexclu au tri I3publié avant 2022
LVD-2M: A Long-take Video Dataset with Temporally Dense Captionsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
FRAME: Forensic Routing and Adaptive Multi-path Evidence Fusion for Image Manipulation Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Seeing Fast and Slow: Learning the Flow of Time in Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
MesoNet: a Compact Facial Video Forgery Detection Networkexclu au tri I3publié avant 2022
Video Signature: In-generation Watermarking for Latent Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Movie Gen: A Cast of Media Foundation Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Media Forensics and DeepFakes: an overviewexclu au tri I3publié avant 2022
VPN: Video Provenance Network for Robust Content Attributionexclu au tri I3publié avant 2022
Forensic Self-Descriptions Are All You Need for Zero-Shot Detection, Open-Set Source Attribution, and Clustering of AI-generated Imagesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
SpeedNet: Learning the Speediness in Videosexclu au tri I3publié avant 2022
STEP: Segmenting and Tracking Every Pixelexclu au tri I3publié avant 2022
DynVFX: Augmenting Real Videos with Dynamic Contentexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Facesexclu au tri I3publié avant 2022
Classification Matters: Improving Video Action Detection with Class-Specific Attentionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Learning Video Representations without Natural Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understandingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding Systemexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
A Sanity Check for AI-generated Image Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Exposing DeepFake Videos By Detecting Face Warping Artifactsexclu au tri I3publié avant 2022
A Survey of AI-Generated Video Evaluationexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolutionexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Imagen Video: High Definition Video Generation with Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
FCMBench-Video: Benchmarking Document Video Intelligenceexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Occluded Video Instance Segmentation: A Benchmarkexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Face Consistency Benchmark for GenAI Videoexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
AIGVE-Tool: AI-Generated Video Evaluation Toolkit with Multifaceted Benchmarkexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and Benchmarkexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Deepfake detection across image, video, and audio: a comprehensive survey with empirical evaluation of generalization and robustnessexclu au tri E1Traite principalement d'images fixes et d'audio, aucune expérience spécifique sur vidéo.
aEYE: A deep learning system for video nystagmus detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
An overview of transformers for video anomaly detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Learning Video Salient Object Detection Progressively from Unlabeled Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Networking Systems for Video Anomaly Detection: A Tutorial and Surveyexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Video Diffusion Models: A Surveyexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Temporally consistent low-light face video enhancement via video-to-video conditional diffusionexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
A Survey on Video Diffusion Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « tache »
Unsupervised Video Anomaly Detection Based on Similarity with Predefined Text Descriptionsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Real-Time Turkish Video Text Detection and Recognitionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Fusion Mechanism for Text Detection and Tracking in Video Scenesexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Video text detection with multiattention feature fusion networkexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Sequential Transformer for End-to-End Video Text Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Emotion Detection from Video and Audio and Textexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
EMOTION DETECTION USING VIDEO AUDIO AND TEXTexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Video Text Detection With Robust Feature Representationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Farsi Text Detection and Localization in Videos and Imagesexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Cursive Caption Text Detection in Videosexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly Detectionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Violent Video Detection Based on Multimodal Video-Text Fusionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
1DCNN Video Tampering Detection in Automotive Dashcam Videos in Temporal Domainexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
Blockchain for video watermarking: An enhanced copyright protection approach for video forensics based on perceptual hash functionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
A New Forensic Video Database for Source Smartphone Identification: Description and Analysisexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
The Significance of Metadata and Video Compression for Investigating Video Files on Social Media Forensicexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
INTELLIGENT VIDEO ANALYTICS FOR CYBER SECURITY AND DIGITAL FORENSICSexclu au tri E6Survol généraliste de l'analytique vidéo pour la cybersécurité, hors sujet pour la détection spécifique de vidéos IA.
SAVAir: A Physics-Guided Temporal Synthetic Dataset for Satellite-Video Aircraft Detection and Trackingexclu au tri E6Détection d'avions et suivi d'objets, pas de falsification IA.
Zero-shot video highlight detection based on text descriptions and synthetic imagesexclu au tri E6Détection de moments forts (highlights) dans des vidéos, hors sujet.
C3DIR: A Deep Learning 3-Dimensional Cloud Property Retrieval Scheme for Passive Satellite Imagersexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Media Meets Communication in 6G: Fundamentals, Key Technologies, and Applicationsexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Calibration-Free 3D Multi-Camera People Tracking for Indoor Environmentexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answeringexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoningexclu au tri E6Analyse d'anomalies de trafic routier dans des vidéos, hors sujet.
A survey of AI-generated voices and their detectionexclu au tri E3Traite uniquement de la détection de voix synthétiques (audio seul).
AsymFeX: A Symmetry-Driven Framework for Ischemic Stroke Segmentation Across Imaging Modalities and Stroke Stagesexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Security Education in Higher Education through AI-Powered Gamificationexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understandingexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillanceexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Modelsexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »
PiMiX 2.02: Toward AI-Driven Data Fusion in Radiographic Imaging and Tomographyexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
ORION-CMR: On-scanner Reporting with Integrated Foundation Model for End-to-End Cardiac MRI Analysis and Interpretationexclu au tri E6filtre mots-clés : aucun terme du groupe « video »
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interactionexclu au tri E6filtre mots-clés : aucun terme du groupe « origine »