arXiv · 2607.15285
Proof-Carrying Multimodal Timelines: Finite-Trace Modal Certificates for Video-Audio Consistency
Abstract
Multimodal video systems often expose clip-level scores while hiding the local temporal failure that makes a video inconsistent. We formulate video-audio consistency as finite-trace modal monitoring over synchronized visual, audio, and subtitle/OCR atoms. A formula library specifies speech-speaker agreement, audio-visual event agreement, subtitle-video agreement, scene continuity, and edit-induced temporal shock. A certificate records the trace hash, formula identifier, verdict, violating indices, defect score, first counterexample window, execution engine, and certificate hash. The certificate is not a proof that a detector is correct; it is a reproducible witness that the finite-trace checker can independently reconstruct once the atom trace is fixed. The artifact decodes real MP4 video, extracts CLIP visual atoms, extracts AST audio atoms, runs dense counterfactual perturbation sweeps, emits CSV traces and JSON certificates, and renders manuscript figures from those files.
Explore related subjects
Keep this discovery
Faruk Alpay, Hamdi Alakkad. 2026-06-16. Proof-Carrying Multimodal Timelines: Finite-Trace Modal Certificates for Video-Audio Consistency. https://arxiv.org/abs/2607.15285
Cite the original work for its findings. Save a collection to share your selection of sources.