arXiv · 2408.11347
Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments
Abstract
We used a 3D simulator to create artificial video data with standardized annotations, aiming to aid in the development of Embodied AI. Our question answering (QA) dataset measures the extent to which a robot can understand human behavior and the environment in a home setting. Preliminary experiments suggest our dataset is useful in measuring AI's comprehension of daily life. \end{abstract}
Explore related subjects
Keep this discovery
Takanori Ugai, Kensho Hara, Shusaku Egami, Ken Fukuda. 2024-08-21. Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments. https://arxiv.org/abs/2408.11347
Cite the original work for its findings. Save a collection to share your selection of sources.