| Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| REMEDY: How Far Is Video Generation from Medical Education World Models? | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| InternW0: A Foundational Physical World Model for Efficient Real-World Interactions | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Companion-style QA Assistance in Ego-Vision | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| From Visual Cues to Spoken Narration: Rethinking Audio Description | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| BreastMammo and DenseMammo: Benchmarks for Mammography Domain Generalization | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| CAS2UML: A Handwritten Sketch-to-PlantUML Dataset for Class and Activity Diagrams | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| MedForge-RSI: Medical Deepfake Detection via Recursive Self-Improvement | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arrive | exclu au tri |
E4 | Position paper sans évaluation empirique directe d'une nouvelle méthode de détection. |
| Towards Generalizable Deepfake Image Detection with Vision Transformers | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Two-Stage Teacher-Student Reliable Prior Learning for Robust Underwater Image Enhancement | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Secret-Stego Dissimilarity as a Design Axis: Invertible Coverless Image Steganography with Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Enhancing Prostate Cancer Segmentation for Multi-Domain Generalization using a novel Parallel-Route Coherent Mixup Regularization Training | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| High-Fidelity Synthetic Transmission Electron Microscopy Image Generation Using Diffusion Probabilistic Models for Data-Limited Semiconductor Metrology | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| 2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| A Surface-based Multimodal Framework for Multitask Analysis in Alzheimer's Disease | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Q-ARVD: Quantizing Autoregressive Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| SCMAPR: Self-Correcting Multi-Agent Prompt Refinement for Complex-Scenario Text-to-Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interaction | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen | exclu au tri |
E6 | Traite de la détection d'anomalies vidéo (crimes), pas de l'IA. |
| PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Back-Tracking from Clarity: Self-Learning to See Text from Afar | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Smartphone-Based Method for Automated Speed Enforcement | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Detecting Phone-Induced Pedestrian Distraction via a Multimodal Fusion Transformer | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Lesion-centered 3D mapping of colonoscopy procedures: validation of a hierarchical ensemble pipeline on public benchmark videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Online Video Agent Harness for Long Video Understanding | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Detecting and Explaining Fake News Short Videos with Multimodal Content and Real-World Evidence | exclu au tri |
E6 | Détection de fake news vidéo générales, sans focus sur l'IA. |
| OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models | exclu au tri |
E6 | Détection d'hallucinations dans les MLLM, hors sujet. |
| Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection | exclu au tri |
E6 | Détection d'anomalies dans les vidéos de surveillance, hors sujet. |
| NOVA: Normal-Side Modeling for Training-Free Zero-Shot Video Anomaly Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| A Data-Driven Multimodal Method for Early Detection of Coordinated Abnormal Behaviors in Live-Streaming Platforms | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Training-Free Temporal Abstraction for General Video Understanding | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku | exclu au tri |
E6 | Détection de fake news vidéo via commentaires, hors sujet. |
| AEGIS: Preventing Cross-Domain Resource Abuse in MCP | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| A Multi-Domain Feature Fusion Framework for Generalizable Deepfake Detection Across Different Generators | exclu au tri |
E1 | Concerne les images générées par IA, aucune expérience sur la vidéo. |
| Bridging the Modality Gap in Forensic Image Retrieval | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Synthetic-to-Real Transfer in Cerebral Microbleed Generation and Segmentation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Memory- and Bandwidth-Efficient SPAD-LiDAR Ranging via Coarse-to-Fine Spline Sketching | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Learned Compression of SAR Phase-History Data: A Rate-Honest Feasibility Study on GOTCHA | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| AutoRASOR: Autonomous Rapid Scanning Electron Microscope Operator | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Conditioning noise is a free regularizer for LoRA fine-tuning: no pathology encoder required for diffusion-based artifact detection in histopathology | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Spatial Attention Supervision for Defect Localization: Exploiting Ground-Truth Masks as Training Signal in Diffusion-Augmented Defect Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video | exclu au tri |
E6 | Détection et localisation d'accidents de la route, hors sujet. |
| PERSIST: Persistent-State Discrimination for Shot Boundary Detection | exclu au tri |
E6 | Détection de transitions de plans dans les vidéos, hors sujet. |
| LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| SweepLSD: A One-Pass, O(width)-Memory Line Segment Detector with an Integer-Only Streaming Core and a Real-Time FPGA Realization | exclu au tri |
E6 | Détecteur de segments de lignes dans les images, hors sujet. |
| Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection | exclu au tri |
E3 | Concerne uniquement la détection de deepfakes audio. |
| $\texttt{DisMorph}$: learning to disentangle technical distortions from true biological change | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Accurate Localization of Road Traffic Objects on the Road Plane Using Surveillance Camera Imagery | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Synthetic data generation framework for quality control automation in gravure printing | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection | exclu au tri |
E6 | Étudie le benchmark de vidéos synthétiques pour la détection de travailleurs sous charge suspendue. |
| SCI-Mamba: Unsupervised Learning based Low-Light Image Enhancement for Non-Cooperative Spacecraft | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog | exclu au tri |
E6 | Se concentre sur la détection et le suivi de drones sous brouillard synthétique. |
| TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Rubric Rewards from Item Response Theory | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Aligning One-Step Generative Models with Reward-Weighted Transport Distillation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| BiasReducer: Adaptive Bias Mitigation for Reward Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| BIABench: Evaluating AI agents on real-world bioimage analysis tasks | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| DAGent: Evaluate-then-Grow Planning for Deep Research Agents | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Training LLM Judges from Language Feedback via Position-Selective Self-Distillation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Scaling and Distilling Text Embeddings for Better Diffusibility | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| AutoGUIWorld: Image Generators as Visual World Models for GUI Agent | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| AutoDataBench: A Data-centric Testbed for Accelerating Auto Research | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Rules to Tools: Executable Checks for LLM Agents in Scientific Computing | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Hierarchical Continuous Diffusion Language Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| It Takes Workflows to Evolve Better Workflows | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| RPTune: Learned Context Curation for LLM Catalog Search | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| ROWBench: Do Video Models Render What the Program Specifies? | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| 4Director: Controlling Video World Models with Rigid 3D Geometry | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Sharpening Tax in Post-Training | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Agent Priors-guided Policy Learning | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Beyond the Current Scene: Event-Referential Grasping with Active View Selection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Controlled Decoding Attacks on Black-Box LLMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| World Observer: Joint Actor-Observer Generation for Persistent World Modeling | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Personalized Image Generation with Reasoning and Reflection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Does Native 3D Texture Generation Necessarily Require 3D Assets for Training? | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Smaller Models, Better Rejects: Preference Distillation Scaling | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Memorizon: Training World Models Beyond Their Context Window | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| LOCI: Spatial Linear Memory for Streaming World Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Decoding Looped Transformers Better for (Almost) Free | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Honeycomb: Constant-Size Scene Memory Representation for Video World Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Video Generation Models: A Survey of Post-Training and Alignment | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Persona Dosing: Calibrated Activation Steering for Graded Trait Control | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| VMBench: A Benchmark for Perception-Aligned Video Motion Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| VI-Bench: Benchmarking Prompt Inversion from AIGC Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| VISTA: A Generative Egocentric Video Framework for Daily Assistance | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| UVE: Are MLLMs Unified Evaluators for AI-Generated Videos? | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| GAIA: Rethinking Action Quality Assessment for AI-Generated Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| RefCaptioner: Multi-Reference Image-Grounded Video Captioning | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| FilmBench: A Film-Grade Benchmark for Cinematic Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation Metric | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Benchmarking AIGC Video Quality Assessment: A Dataset and Unified Model | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| HumanScore: Benchmarking Human Motions in Generated Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Evaluating Deepfake Detectors in the Wild | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Diffusion Deepfake | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| ED^4: Explicit Data-level Debiasing for Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| The Deepfake Detection Challenge (DFDC) Preview Dataset | exclu au tri |
I3 | publié avant 2022 |
| SeeABLE: Soft Discrepancies and Bounded Contrastive Learning for Exposing Deepfakes | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| AuthGuard: Generalizable Deepfake Detection via Language Guidance | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Penny-Wise and Pound-Foolish in Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Combining EfficientNet and Vision Transformers for Video Deepfake Detection | exclu au tri |
I3 | publié avant 2022 |
| UCF: Uncovering Common Features for Generalizable Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| The DeepFake Detection Challenge (DFDC) Dataset | exclu au tri |
I3 | publié avant 2022 |
| Adversarially robust deepfake media detection using fused convolutional neural network predictions | exclu au tri |
I3 | publié avant 2022 |
| Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Robust AI-Generated Face Detection with Imbalanced Data | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Diffusion Action Segmentation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| TPDiff: Temporal Pyramid Video Diffusion Model | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| FreeInit: Bridging Initialization Gap in Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Efficient Video Diffusion Models: Advancements and Challenges | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| HARIVO: Harnessing Text-to-Image Models for Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Visual Bridge: Universal Visual Perception Representations Generating | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Point Prompting: Counterfactual Tracking with Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Diffusion Models for Video Prediction and Infilling | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Generative Video Matting | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| APLA: Additional Perturbation for Latent Noise with Adversarial Training Enables Consistency | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Diffusion Classifiers Understand Compositionality, but Conditions Apply | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Diffusion Models in Low-Level Vision: A Survey | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| LLM-grounded Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Diffusion Noise Feature: Accurate and Fast Generated Image Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Studying Image Diffusion Features for Zero-Shot Video Object Segmentation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| SinFusion: Training Diffusion Models on a Single Image or Video | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Diffusion Models as Masked Autoencoders | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Optical-Flow Guided Prompt Optimization for Coherent Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Video Diffusion Alignment via Reward Gradients | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| VidGen-1M: A Large-Scale Dataset for Text-to-video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Learning Video Representations from Textual Web Supervision | exclu au tri |
I3 | publié avant 2022 |
| Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval | exclu au tri |
I3 | publié avant 2022 |
| Video and Text Matching with Conditioned Embeddings | exclu au tri |
I3 | publié avant 2022 |
| MuSeD: A Multimodal Spanish Dataset for Sexism Detection in Social Media Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Text Detection and Recognition in the Wild: A Review | exclu au tri |
I3 | publié avant 2022 |
| Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| A Survey on Multimodal Disinformation Detection | exclu au tri |
I3 | publié avant 2022 |
| Unveiling Hallucination in Text, Image, Video, and Audio Foundation Models: A Comprehensive Survey | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Seeing, Signing, and Saying: A Vision-Language Model-Assisted Pipeline for Sign Language Data Acquisition and Curation from Social Media | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Combating Online Misinformation Videos: Characterization, Detection, and Future Directions | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Perception Test: A Diagnostic Benchmark for Multimodal Video Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Helping Hands: An Object-Aware Ego-Centric Video Recognition Model | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| TubeDETR: Spatio-Temporal Video Grounding with Transformers | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Saliency-Guided DETR for Moment Retrieval and Highlight Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Tracking | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Text-guided Fine-Grained Video Anomaly Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| DrawVideo: Generating Long Video from Storyboard Keyframe Sketches | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Dynamic texture analysis for detecting fake faces in video sequences | exclu au tri |
I3 | publié avant 2022 |
| LVD-2M: A Long-take Video Dataset with Temporally Dense Captions | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| FRAME: Forensic Routing and Adaptive Multi-path Evidence Fusion for Image Manipulation Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Seeing Fast and Slow: Learning the Flow of Time in Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| MesoNet: a Compact Facial Video Forgery Detection Network | exclu au tri |
I3 | publié avant 2022 |
| Video Signature: In-generation Watermarking for Latent Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Movie Gen: A Cast of Media Foundation Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Media Forensics and DeepFakes: an overview | exclu au tri |
I3 | publié avant 2022 |
| VPN: Video Provenance Network for Robust Content Attribution | exclu au tri |
I3 | publié avant 2022 |
| Forensic Self-Descriptions Are All You Need for Zero-Shot Detection, Open-Set Source Attribution, and Clustering of AI-generated Images | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| SpeedNet: Learning the Speediness in Videos | exclu au tri |
I3 | publié avant 2022 |
| STEP: Segmenting and Tracking Every Pixel | exclu au tri |
I3 | publié avant 2022 |
| DynVFX: Augmenting Real Videos with Dynamic Content | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces | exclu au tri |
I3 | publié avant 2022 |
| Classification Matters: Improving Video Action Detection with Class-Specific Attention | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Learning Video Representations without Natural Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| A Sanity Check for AI-generated Image Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Exposing DeepFake Videos By Detecting Face Warping Artifacts | exclu au tri |
I3 | publié avant 2022 |
| A Survey of AI-Generated Video Evaluation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Imagen Video: High Definition Video Generation with Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| FCMBench-Video: Benchmarking Document Video Intelligence | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Occluded Video Instance Segmentation: A Benchmark | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Face Consistency Benchmark for GenAI Video | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| AIGVE-Tool: AI-Generated Video Evaluation Toolkit with Multifaceted Benchmark | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and Benchmark | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Deepfake detection across image, video, and audio: a comprehensive survey with empirical evaluation of generalization and robustness | exclu au tri |
E1 | Traite principalement d'images fixes et d'audio, aucune expérience spécifique sur vidéo. |
| aEYE: A deep learning system for video nystagmus detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| An overview of transformers for video anomaly detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Learning Video Salient Object Detection Progressively from Unlabeled Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Networking Systems for Video Anomaly Detection: A Tutorial and Survey | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Video Diffusion Models: A Survey | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Temporally consistent low-light face video enhancement via video-to-video conditional diffusion | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| A Survey on Video Diffusion Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « tache » |
| Unsupervised Video Anomaly Detection Based on Similarity with Predefined Text Descriptions | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Real-Time Turkish Video Text Detection and Recognition | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Fusion Mechanism for Text Detection and Tracking in Video Scenes | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Video text detection with multiattention feature fusion network | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Sequential Transformer for End-to-End Video Text Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Emotion Detection from Video and Audio and Text | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| EMOTION DETECTION USING VIDEO AUDIO AND TEXT | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Video Text Detection With Robust Feature Representation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Farsi Text Detection and Localization in Videos and Images | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Cursive Caption Text Detection in Videos | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly Detection | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Violent Video Detection Based on Multimodal Video-Text Fusion | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| 1DCNN Video Tampering Detection in Automotive Dashcam Videos in Temporal Domain | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| Blockchain for video watermarking: An enhanced copyright protection approach for video forensics based on perceptual hash function | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| A New Forensic Video Database for Source Smartphone Identification: Description and Analysis | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| The Significance of Metadata and Video Compression for Investigating Video Files on Social Media Forensic | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| INTELLIGENT VIDEO ANALYTICS FOR CYBER SECURITY AND DIGITAL FORENSICS | exclu au tri |
E6 | Survol généraliste de l'analytique vidéo pour la cybersécurité, hors sujet pour la détection spécifique de vidéos IA. |
| SAVAir: A Physics-Guided Temporal Synthetic Dataset for Satellite-Video Aircraft Detection and Tracking | exclu au tri |
E6 | Détection d'avions et suivi d'objets, pas de falsification IA. |
| Zero-shot video highlight detection based on text descriptions and synthetic images | exclu au tri |
E6 | Détection de moments forts (highlights) dans des vidéos, hors sujet. |
| C3DIR: A Deep Learning 3-Dimensional Cloud Property Retrieval Scheme for Passive Satellite Imagers | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Media Meets Communication in 6G: Fundamentals, Key Technologies, and Applications | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning | exclu au tri |
E6 | Analyse d'anomalies de trafic routier dans des vidéos, hors sujet. |
| A survey of AI-generated voices and their detection | exclu au tri |
E3 | Traite uniquement de la détection de voix synthétiques (audio seul). |
| AsymFeX: A Symmetry-Driven Framework for Ischemic Stroke Segmentation Across Imaging Modalities and Stroke Stages | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Security Education in Higher Education through AI-Powered Gamification | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |
| PiMiX 2.02: Toward AI-Driven Data Fusion in Radiographic Imaging and Tomography | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| ORION-CMR: On-scanner Reporting with Integrated Foundation Model for End-to-End Cardiac MRI Analysis and Interpretation | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « video » |
| Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction | exclu au tri |
E6 | filtre mots-clés : aucun terme du groupe « origine » |