arXiv · 2609.20001
E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews
Abstract
Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Haoshen Wang, Dongbo Che, Zeyi Xie, Yuanjie Du, Shicheng Hua, Xingyu Wang. 2026-09-17. E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews. https://arxiv.org/abs/2609.20001
Cite the original work for its findings. Save a collection to share your selection of sources.