669Identifiées −38 doublons 631Uniques −352 exclues 61Retenues au tri 57 à lire 4Lues 4 à décider 0Incluses

Bibliothèque / fiche n°507

VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent

Fu et al.arXiv2026 exclu au tri

Article PDF

Cet article n'a pas de fiche

Statut : exclu au tri — filtre mots-clés : aucun terme du groupe « origine ». Seuls les articles retenus au tri et dont le PDF est accessible sont lus en entier.

Résumé des auteurs

World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance with the laws of physics, thus opening a compelling application: fusing multimodal legal evidence to re-create a crime scene and re-enact how an offence could have been committed. However, feeding the raw, unorganized evidence into a world model fails in forensic use: it silently drops evidence, glosses over contradictory testimony, and produces motion that violates the evidentiary record. This paper presents VeriScene, an agent that orchestrates the world model: it reconstructs crime scenes from forensic photographs and witness statements of varying reliability, keeping every claim traceable to evidence and every motion physically plausible. VeriScene iteratively fuses the evidence into a cited narrative under an auditing loop, verifies the hypothesized dynamics via probe rollouts in the world model with corrective constraint injection, and renders the offence as a re-enactment video from a fused keyframe. On a benchmark of 25 crime scenarios across 7 physically-driven case types (139 forensic-style photographs and 65 statements with planted unreliability), VeriScene attains 0.9014 evidence coverage and 0.7217 factual consistency (0-1 scale) on the 20 test scenes, outperforming an end-to-end multimodal-LLM baseline by 20.35% in factual consistency and 34.88% in temporal coherence, while generalizing across four LLM orchestration backends at USD 1.82 per scene.