SearcharxivSearch

arXiv subjects

Kaixiang Yao

Publications and source records attributed to Kaixiang Yao.

8 recordsLinked to original sources

RopeFormer: Cross-Trial Adaptation from Interaction History for Dynamic Rope Manipulation

Dynamic rope manipulation is highly sensitive to unknown object dynamics: the same robot motion can produce substantially different responses across ropes, while explicitly identifying the relevant physical properties is difficult. We present RopeFormer, a history-conditioned framework that uses prior task interaction as context for subsequent control. The policy retains cross-trial action-response history while keeping its weights fixed and requires no explicit online rope-parameter estimation. In matched simulation evaluations across sustained single-arm rotation, bimanual rotation, and transient whipping, retaining context improves subsequent control relative to resetting the same checkpoint, with the benefit varying across rope dynamics and observation settings. We further deploy the frozen policies on a Unitree H1-2 with previously unseen physical ropes. From T1 to T3, target-acquisition time decreases by 30.9% for Rope Swing and 33.9% for Rope Twirl, while mean Rope Whip target hits increase from 0.2 to 2.3 out of three. These results show that prior interaction can provide effective control context for dynamic deformable-object manipulation. Robot videos, code, and data are available at https://ropeformer.github.io/.

cs.RO

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.

cs.AI

TOGEARI: Interaction-Space Preconditioning for Condensed Finite-Element Systems with IPC Contact

TOGEARI constructs a compressed contact correction for condensed finite-element equations with incremental potential contact (IPC). A core-conditioned interaction matrix ranks contact combinations by their mechanical responses; a raw-factor proxy acquires only the retained responses. For a symmetric positive-definite core, the reference analysis gives the exact spectrum, scaled inverse error and optimal worst omitted interaction. The selected responses are maintained through signed contact changes using supported operator images and a small coarse matrix. A complementary Schur spectrum and perturbation bounds connect this construction to the full current linear equation, damped local Newton steps and compatible episode reuse. On a 325,260-coordinate finger/cup problem, eight selected directions reduce Krylov work from eleven columns to five, matching full contact. Two reversed-order load studies give matching anchored displacements and accepted normal reactions, with 28.8-29.3 percent less warm time and 5.1-5.3 percent less complete time than per-step direct refactorization. A completed twelve-episode run saves 18.8 percent of total time against direct and 2.8 percent against a held factor. Full response caching has comparable time with 7.5 times the response storage. A compressed repeat terminates at an energy line search; separate diagnostics identify constitutive cancellation and terminal-backtracking sensitivity. Larger-contact, held-out-load and frictional comparisons show where compression or preparation loses its advantage. The complete Newton equation remains the residual operator throughout; the convergence guarantees retain their stated reference and local hypotheses.

math.NA

RUPA: Nonlinear volume consistency, constraint geometry and singular penalty limits in finite elements

Volume quadrature can change nonlinear finite-element constraints while preserving their reference-state derivatives. We connect an explicit determinant defect to feasible-set geometry and singular mechanical response. For affine tensor elements of coordinate degree $p\ge3$ with $n\ge p+1$ Gauss points per coordinate, determinant volume is exact precisely when $2n\ge3p$. Below that threshold we construct a boundary-fixed cubic defect at every order. The same directions yield a full-space cube-root residual--distance bound under explicit cell-support, coefficient and physical-norm assumptions, with mesh-uniform upper constants at fixed order. With all cell-pressure equations retained, the volume Jacobian gains rank at nearby feasible states despite agreement through second derivatives at rest; the cube-root exponent is sharp on each fixed mesh. A general localized-minimum theorem shows that the first reduced compatibility term contributes its weighted square to the leading energy in a joint small-load, large-bulk limit. Cubic and quadratic defects therefore produce sextic and quartic terms. The full cubic-element interior space has an exact normal form and sharp local error exponents. Curved quadratic tetrahedra supply the second-order contrast, a sparse rational witness and an exact four-Jacobian volume formula. Finite-strain tensor calculations illustrate normalized response separation, with explicit stationary-point, extreme-bulk and pressure-recovery qualifications. The constructive correction preserves exact cell volumes, so quadrature feasibility remains distinct from physical volume preservation.

math.NA

MORTIS: Quadrature-compatible cofactor gauges and optimal material margins in finite elasticity

We develop a discrete design theory for cofactor reference forms that preserve the complete finite-element elasticity tangent. A fixed continuous potential generates the coefficient, while the variation space and quadrature determine compatibility. We classify the exact local potential spaces for all-order tensor elements, anisotropic spaces, cubic serendipity and quadratic tetrahedra. At a stress-free Mooney-Rivlin state with a common positive material-coefficient sum, a convex loss expresses the attainable quadrature-point material margin. The optimum is attained; under stated space and sampling assumptions, its ideal value is attained precisely when the physical identity potential is compatible. A conforming fixed-candidate example proves a strict margin loss, linear in curvature amplitude and uniform in cell size, under displacement enrichment. Explicit positive candidates accompany this restriction. A Gram identity gives the sharp constant-coefficient defect at a prescribed margin in the stated norm. Compatible and reference-defect-corrected forms retain the original tangent through a deformation-independent external-boundary Hessian; equal potential traces give equal exact assembled gauges. Quadratic tensor geometry admits a stress-free physical coercivity bound uniform in mesh size and displacement order under explicit material, shape, quadrature and boundary assumptions. Certified curved tetrahedral shear families with zero recovered cell pressure establish finite-deformation regimes, while a loaded pressure-curvature counterexample identifies their limitation. The theory and observations distinguish local exactness, material certificates, complete-reference coercivity and unchanged equations.

math.NA

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.

cs.CV

SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition

Recent advances in Large Vision-Language Models (LVLMs) have enabled training-free Fine-Grained Visual Recognition (FGVR). However, effectively exploiting LVLMs for FGVR remains challenging due to the inherent visual ambiguity of subordinate-level categories. Existing methods predominantly adopt either retrieval-oriented or reasoning-oriented paradigms to tackle this challenge, but both are constrained by two fundamental limitations:(1) They apply the same inference pipeline to all samples without accounting for uneven recognition difficulty, thereby leading to suboptimal accuracy and efficiency; (2) The lack of mechanisms to consolidate and reuse error-specific experience causes repeated failures on similar challenging cases. To address these limitations, we propose SARE, a Sample-wise Adaptive textbfREasoning framework for training-free FGVR. Specifically, SARE adopts a cascaded design that combines fast candidate retrieval with fine-grained reasoning, invoking the latter only when necessary. In the reasoning process, SARE incorporates a self-reflective experience mechanism that leverages past failures to provide transferable discriminative guidance during inference, without any parameter updates. Extensive experiments across 14 datasets substantiate that SARE achieves state-of-the-art performance while substantially reducing computational overhead.

cs.CV

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents

Our research uncovers a novel privacy risk associated with multimodal large language models (MLLMs): the ability to infer sensitive personal attributes from audio data -- a technique we term audio private attribute profiling. This capability poses a significant threat, as audio can be covertly captured without direct interaction or visibility. Moreover, compared to images and text, audio carries unique characteristics, such as tone and pitch, which can be exploited for more detailed profiling. However, two key challenges exist in understanding MLLM-employed private attribute profiling from audio: (1) the lack of audio benchmark datasets with sensitive attribute annotations and (2) the limited ability of current MLLMs to infer such attributes directly from audio. To address these challenges, we introduce AP^2, an audio benchmark dataset that consists of two subsets collected and composed from real-world data, and both are annotated with sensitive attribute labels. Additionally, we propose Gifts, a hybrid multi-agent framework that leverages the complementary strengths of audio-language models (ALMs) and large language models (LLMs) to enhance inference capabilities. Gifts employs an LLM to guide the ALM in inferring sensitive attributes, then forensically analyzes and consolidates the ALM's inferences, overcoming severe hallucinations of existing ALMs in generating long-context responses. Our evaluations demonstrate that Gifts significantly outperforms baseline approaches in inferring sensitive attributes. Finally, we investigate model-level and data-level defense strategies to mitigate the risks of audio private attribute profiling. Our work validates the feasibility of audio-based privacy attacks using MLLMs, highlighting the need for robust defenses, and provides a dataset and framework to facilitate future research.

cs.CR