SearcharxivSearch

arXiv subjects

Joey Huang

Publications and source records attributed to Joey Huang.

6 recordsLinked to original sources

Natural Language Camera Movement Understanding

Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing vision-language models (VLMs) fail this task in surprising ways, frequently confusing translation with rotation, left with right, and object movement with camera movement. To address these limitations, we establish natural language camera movement understanding as a standalone research task. We introduce a two-level cinematographic taxonomy and an extensive, atomic benchmark featuring both real and synthetic videos. Furthermore, we curate a large-scale, multi-source training set enhanced by targeted camera movement augmentation. Our fine-tuned VLM-8B outperforms Gemini 3.1 Pro by 10% and 11% on our benchmark's real and synthetic videos, respectively. Despite these gains, a significant gap remains relative to human performance, underscoring the need to promote and facilitate future research on natural language camera movement understanding.

cs.CV

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-language modeling that extensively improves upon BabyVLM-V1 through a longitudinal, multifaceted pretraining set, a versatile model, and, most importantly, DevCV Toolbox for cognitive evaluation. The pretraining set maximizes coverage while minimizing curation of a longitudinal, infant-centric audiovisual corpus, yielding video-utterance, image-utterance, and multi-turn conversational data that mirror infant experiences. DevCV Toolbox adapts all vision-related measures of the recently released NIH Baby Toolbox into a benchmark suite of ten multimodal tasks, covering spatial reasoning, memory, and vocabulary understanding aligned with early children's capabilities. Experimental results show that a compact model pretrained from scratch can achieve competitive performance on DevCV Toolbox, outperforming GPT-4o on some tasks. We hope the principled, unified BabyVLM-V2 framework will accelerate research in developmentally plausible pretraining of vision foundation models.

cs.CV

A Simple Mathematical Model of Politics (II)

In this paper, some main eigenvalues and eigenvectors of the politics matrix are investigated. The number of upper-class families in a society is the number of eigenvalues which are very close to 1. An algorithm to identify all the upper-class families from the right and left eigenvectors of those eigenvalues is developed.

cs.SI

Key principles for workforce upskilling via online learning: a learning analytics study of a professional course in additive manufacturing

Effective adoption of online platforms for teaching, learning, and skill development is essential to both academic institutions and workplaces. Adoption of online learning has been abruptly accelerated by COVID19 pandemic, drawing attention to research on pedagogy and practice for effective online instruction. Online learning requires a multitude of skills and resources spanning from learning management platforms to interactive assessment tools, combined with multimedia content, presenting challenges to instructors and organizations. This study focuses on ways that learning sciences and visual learning analytics can be used to design, and to improve, online workforce training in advanced manufacturing. Scholars and industry experts, educational researchers, and specialists in data analysis and visualization collaborated to study the performance of a cohort of 900 professionals enrolled in an online training course focused on additive manufacturing. The course was offered through MITxPro, MIT Open Learning is a professional learning organization which hosts in a dedicated instance of the edX platform. This study combines learning objective analysis and visual learning analytics to examine the relationships among learning trajectories, engagement, and performance. The results demonstrate how visual learning analytics was used for targeted course modification, and interpretation of learner engagement and performance, such as by more direct mapping of assessments to learning objectives, and to expected and actual time needed to complete each segment of the course. The study also emphasizes broader strategies for course designers and instructors to align course assignments, learning objectives, and assessment measures with learner needs and interests, and argues for a synchronized data infrastructure to facilitate effective just in time learning and continuous improvement of online courses.

cs.HC

Pseudo-real-time retinal layer segmentation for high-resolution adaptive optics optical coherence tomography

We present a pseudo-real-time retinal layer segmentation for high-resolution Sensorless Adaptive Optics-Optical Coherence Tomography (SAO-OCT). Our pseudo-real-time segmentation method is based on Dijkstra's algorithm that uses the intensity of pixels and the vertical gradient of the image to find the minimum cost in a geometric graph formulation within a limited search region. It segments six retinal layer boundaries in an iterative process according to their order of prominence. The segmentation time is strongly correlated to the number of retinal layers to be segmented. Our program permits en face images to be extracted during data acquisition to guide the depth specific focus control and depth dependent aberration correction for high-resolution SAO-OCT systems. The average processing times for our entire pipeline for segmenting six layers in a retinal B-scan of 496x400 pixels and 240x400 pixels are around 25.60 ms and 13.76 ms, respectively. When reducing the number of layers segmented to only two layers, the time required for a 240x400 pixel image is 8.26 ms.

eess.IV