arXiv · 2509.14001
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
Abstract
Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Lightweight models often struggle in this setting due to their weak semantic priors, while large vision-language models (VLMs) offer strong object-level understanding but are too computationally demanding for real-time or on-device applications. We introduce MOCHA (Multi-modal Objects-aware Cross-arcHitecture Alignment), a distillation framework that transfers multimodal region-level knowledge from a frozen VLM teacher into a lightweight vision-only detector. MOCHA extracts fused visual and textual teacher's embeddings and uses them to guide student training through a dual-objective loss that enforces accurate local alignment and global relational consistency across regions. This process enables efficient transfer of semantics without the need for teacher modifications or textual input at inference. MOCHA consistently outperforms prior baselines across four personalized detection benchmarks under strict few-shot regimes, yielding a +10.1 average improvement, with minimal inference cost.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Elena Camuffo, Francesco Barbato, Mete Ozay, Simone Milani, Umberto Michieli. 2025-09-17. MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment. https://arxiv.org/abs/2509.14001
Cite the original work for its findings. Save a collection to share your selection of sources.