arXiv · 2609.34719
ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation
Abstract
Autonomous excavation requires a spatial representation that jointly captures terrain geometry and task-relevant semantics. Existing excavation mapping is largely elevation-centric, while generic semantic models remain unstable in unstructured outdoor scenes. We present ExcavaTwin, a pure-vision geometry-guided semantic elevation mapping framework without excavation-specific training. Given multi-view RGB images, the framework: 1) reconstructs scene geometry and semantic observations using frozen vision models; 2) derives terrain and non-terrain geometric support; 3) performs geometry-constrained multi-view semantic fusion to suppress implausible predictions and recover incomplete observations; and 4) projects the fused state into a task-oriented semantic elevation map. Experiments on public datasets and real excavation scenes demonstrate reliable geometric and semantic perception. In real excavation, the system achieved an average update interval of approximately 1.4 s and a mean elevation error of 12.74cm in dynamically modified regions. Larger errors mainly occur during rapid terrain changes and transient visual disturbances caused by machine motion.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yu Deng, Lingshan Zeng, Tong Hu, Rushi Dai. 2026-09-28. ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation. https://arxiv.org/abs/2609.34719
Cite the original work for its findings. Save a collection to share your selection of sources.