SearcharxivSearch

arXiv subjects

Frangil Ramirez

Publications and source records attributed to Frangil Ramirez.

3 recordsLinked to original sources

A solution to generalized learning from small training sets found in infants repeated visual experiences of individual objects

One-year-old infants rapidly form and generalize categories from idiosyncratic experiences of very few exemplars of those categories. Here we provide evidence on the statistics of infants daily-life visual experiences for 8 object categories. Using a corpus of infant head-camera images recorded at mealtimes (87 mealtimes,14 infants), we measure the frequency of the unique instances of each category and the variability of the visual experiences within and across instances of the same category. The frequency distributions of instances for individual infants are highly skewed, containing many images of the same few objects along with fewer images of other instances. Graph theoretic measures of individual category experiences for individual children reveal a lumpy mix of high similarity and high variability, organized into multiple but interconnected clusters of high-similarity images. In computational experiments, we show that artificially created training sets characterized by an interconnected mix of high and low similarity support generalization to novel instances after limited training. We discuss implications for category recognition, and for learning more generally, by both humans and machines.

cs.CV

EgoVIS@CVPR: What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning

Understanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Yet, existing work on procedure-aware video representations fails to explicitly learned the state changes (scene transformations). In this work, we study procedure-aware video representation learning by incorporating state-change descriptions generated by LLMs as supervision signals for video encoders. Moreover, we generate state-change counterfactuals that simulate hypothesized failure outcomes, allowing models to learn by imagining the unseen ``What if'' scenarios. This counterfactual reasoning facilitates the model's ability to understand the cause and effect of each step in an activity. To verify the procedure awareness of our model, we conduct extensive experiments on procedure-aware tasks, including temporal action segmentation, error detection, and more. Our results demonstrate the effectiveness of the proposed state-change descriptions and their counterfactuals, and achieve significant improvements on multiple tasks.

cs.CV

What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning

Understanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Existing work has studied procedure-aware video representations by modeling the temporal order of actions, but has not explicitly learned the state changes (scene transformations). In this work, we study procedure-aware video representation learning by incorporating state-change descriptions generated by Large Language Models (LLMs) as supervision signals for video encoders. Moreover, we generate state-change counterfactuals that simulate hypothesized failure outcomes, allowing models to learn by imagining unseen "What if" scenarios. This counterfactual reasoning facilitates the model's ability to understand the cause and effect of each step in an activity. We conduct extensive experiments on procedure-aware tasks, including temporal action segmentation, error detection, action phase classification, frame retrieval, multi-instance retrieval, and action recognition. Our results demonstrate the effectiveness of the proposed state-change descriptions and their counterfactuals, and achieve significant improvements on multiple tasks.

cs.CV