SearcharxivSearch

arXiv subjects

Tiantian Zheng

Publications and source records attributed to Tiantian Zheng.

4 recordsLinked to original sources

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos

Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models (VLMs) to this challenging domain with Compositional Context Fine-Tuning (CCFT), a method that decomposes assembly actions into semantic elements (Verb, Object, Tool) and fine-tunes VLMs to recognize each action element using templated question-answering pairs. This approach ensures near-deterministic outputs. To enable efficient and effective multi-task learning under limited data, a Layer-Partitioned Alternating Training (LP-AT) method is presented, which assigns distinct model layers to recognize specific action elements through element-specific low-rank adapters. LP-AT alternates weight updates across element-specific adapters, reducing cross-task interference while enabling per-adapter hyperparameter optimization. Furthermore, we create HA-ViD-VQA and IKEA-ASM-VQA datasets from existing assembly video datasets. Extensive experiments on these datasets demonstrate that our method consistently outperforms strong action recognition baselines while providing interpretable element-level predictions that can support diverse downstream applications.

cs.CV

Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning

Dual-hand action segmentation, densely predicting actions for both hands from untrimmed videos, is essential for understanding complex bimanual activities. However, it poses several unique challenges: complex inter-hand dependencies, visual asymmetry between hands, representation conflicts where the dominant hand monopolizes gradients, and semantic ambiguity in fine-grained actions. We propose Polyphony, a three-stage method to address these challenges through: (1) an Alternating Dual-Hand Vision Transformer that alternates training between left- and right-hand mini-batches to ensure balanced gradient contributions from both hands while sharing a spatio-temporal encoder; (2) Semantic Feature Conditioning that aligns visual features with structured, compositional action descriptions to enhance discrimination of semantically similar actions; and (3) Diffusion-Based Segmentation with cross-hand feature fusion for inter-hand coordination and adaptive loss weighting for balancing performance. Polyphony achieves state-of-the-art on both dual-hand datasets (HA-ViD, ATTACH) with improvements up to 16.8 points, and on the single-stream Breakfast dataset (82.5%), outperforming the prior best method that uses a 12x larger backbone. Notably, our unified model with a single shared backbone surpasses baselines requiring separate per-hand models. Code is at https://github.com/x-labs-xyz/Polyphony-Dual-hand-Action-Segmentation.

cs.CV

Exploration of Embodied Space Experience through Umbilical Interaction: A Grounded Theory Approach

This paper critiques the limits of human-centered design in HCI, proposing a shift toward Interface-Centered Design. Drawing on Hookway's philosophy of interfaces, phenomenology, and embodied interaction, we created Umbilink, an umbilical interaction device simulating a uterine environment with tactile sensors and rhythmic feedback to induce a pre-subjectivized state of sensory reduction. Participants' experiences were captured through semi-structured interviews and analyzed with grounded theory. Our contributions are: (1) introducing the novel interface type of Umbilical Interaction; (2) demonstrating the cognitive value of materialized interfaces in a human-interface-environment relation; (3) highlighting the design role of wearing rituals as liminal experiences. As a pilot study, this design suggests imaginative applications in healing, meditation, and sleep, while offering a speculative tool for future interface research.

cs.HC

P-MaNGA: Gradients in Recent Star Formation Histories as Diagnostics for Galaxy Growth and Death

We present an analysis of the data produced by the MaNGA prototype run (P-MaNGA), aiming to test how the radial gradients in recent star formation histories, as indicated by the 4000AA-break (D4000), Hdelta absorption (EW(Hd_A)) and Halpha emission (EW(Ha)) indices, can be useful for understanding disk growth and star formation cessation in local galaxies. We classify 12 galaxies observed on two P-MaNGA plates as either centrally quiescent (CQ) or centrally star-forming (CSF), according to whether D4000 measured in the central spaxel of each datacube exceeds 1.6. For each galaxy we generate both 2D maps and radial profiles of D4000, EW(Hd_A) and EW(Ha). We find that CSF galaxies generally show very weak or no radial variation in these diagnostics. In contrast, CQ galaxies present significant radial gradients, in the sense that D4000 decreases, while both EW(Hd_A) and EW(Ha) increase from the galactic center outward. The outer regions of the galaxies show greater scatter on diagrams relating the three parameters than their central parts. In particular, the clear separation between centrally-measured quiescent and star-forming galaxies in these diagnostic planes is largely filled in by the outer parts of galaxies whose global colors place them in the green valley, supporting the idea that the green valley represents a transition between blue-cloud and red-sequence phases, at least in our small sample. These results are consistent with a picture in which the cessation of star formation propagates from the center of a galaxy outwards as it moves to the red sequence.

astro-ph.GA