TY - RPRT TI - What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions AU - Brian Chen AU - Nina Shvetsova AU - Andrew Rouditchenko AU - Daniel Kondermann AU - Samuel Thomas AU - Shih-Fu Chang AU - Rogerio Feris AU - James Glass AU - Hilde Kuehne PY - 2024 UR - https://arxiv.org/abs/2303.16990 ID - 2303.16990 ER -