arXiv · 2412.09072
Cross-View Completion Models are Zero-shot Correspondence Estimators
Abstract
In this work, we explore new perspectives on cross-view completion learning by drawing an analogy to self-supervised correspondence learning. Through our analysis, we demonstrate that the cross-attention map within cross-view completion models captures correspondence more effectively than other correlations derived from encoder or decoder features. We verify the effectiveness of the cross-attention map by evaluating on both zero-shot matching and learning-based geometric matching and multi-frame depth estimation. Project page is available at https://cvlab-kaist.github.io/ZeroCo/.
Explore related subjects
Keep this discovery
Honggyu An, Jinhyeon Kim, Seonghoon Park, Jaewoo Jung, Jisang Han, Sunghwan Hong, Seungryong Kim. 2024-12-12. Cross-View Completion Models are Zero-shot Correspondence Estimators. https://arxiv.org/abs/2412.09072
Cite the original work for its findings. Save a collection to share your selection of sources.