SearcharxivSearch

arXiv subjects

Haibo Lu

Publications and source records attributed to Haibo Lu.

11 recordsLinked to original sources

Local maximal-canard threshold shifts under Runge--Kutta discretization: an observable-specific order condition

Near a planar fast--slow fold, a local maximal canard is selected by the parameter at which the attracting and repelling slow manifolds meet. We compare this threshold for a physical flow and a Runge--Kutta map, using actual invariant manifolds on a common fold section. An order-two Runge--Kutta method has two independent order-three rooted-tree defects. Both enter the pointwise one-step residual, but Gaussian fold transport acts on their leading contribution by $(\alpha,\beta)\mapsto-3\beta\Xi(J)/8$. Here $\Xi(J)$ is an explicit functional of the fold jet. Thus the singular passage filters the numerical defect space: it annihilates the bushy-tree direction and can retain only the chain-tree direction. For compact analytic classes of affinely normalizable folds and every fixed compact, uniformly finite-stage family of real Runge--Kutta methods of order at least two, the actual flow and map splittings obtained from independent continuations have unique roots whose displacement satisfies a uniform absolute estimate throughout the full small-step rectangle. Whenever the step-independent, exponentially small selection ambiguity is $o(h^2\varepsilon^2)$, the joint-fold law is $\lambda_{\mathrm{RK}}-\lambda_{\mathrm{flow}}=K_\theta(J)h^2\varepsilon^2+o(h^2\varepsilon^2)$. Fold-matched continuations additionally give ordinary second-order convergence as $h\to0$ with $\varepsilon$ fixed. The leading joint-fold bias therefore vanishes under $b^T A c=1/6$, without classical third order. This is cancellation in one nonlinear observable, not an increase in trajectory order or, in general, in fixed-$\varepsilon$ threshold order. Affine covariance transfers the coefficient to physical fold germs, and a van der Pol invariant-graph computation illustrates the sign change, cancellation, and fixed-$\varepsilon$ convergence.

math.DS

Neutral Returns at High-Order Grazing: Sharp Cyclicity, Physical Codimension, and Weighted Crossover

We study a neutral periodic orbit of a continuous two-field flow that is tangent to a switching seam while nearby orbits enter the second field. We separate the roles of two governing integers: the neutral-return order $n\geq2$ controls orbit count, whereas the even contact order $\nu$ controls grazing codimension, passage scale, and transverse crossover. Continuity factors the branch mismatch as $X^+-X^-=hW$; an active excursion of duration $O(q^{1/\nu})$ therefore produces a correction of order $q_+^{1+1/\nu}$. A parameter-uniform passage theorem and a cross-seam Hermite zero theorem then give the sharp fixed-stratum cyclicity: $n$ or $n+1$ periodic orbits, according to the sign coupling of smooth and active terms. A rank-$n$ physical unfolding attains the applicable bound. Contact-jet incidence gives ambient codimension $\nu-2$ for order-$\nu$ grazing and $n+\nu-2$ for simultaneous neutral grazing. Transverse contact parameters replace the central power by a weighted positive-part cap whose onset exponent crosses from $1+1/\nu$ to the generic quadratic-contact value $3/2$. The same sharp bound holds for both the exact cap and a prepared physical multiwell return, including births, mergers, and multiple entry--exit pairs. A value-only $C^1$ counterexample shows that finite cyclicity requires event-derivative information. For every even $\nu\geq4$, a closed polynomial family realizes the prescribed contact and rank conditions and the sharp periodic-orbit configurations through actual finite-time first-return maps.

math.DS

Neutral Entry--Exit Cycles with Quadratic Grazing: Uniform Return Reduction and Local Two-Parameter Bifurcations

We study planar continuous piecewise-smooth slow--fast return circuits in which an invariant-line entry--exit passage is followed by a separated quadratic grazing and the reference cycle has unit multiplier. We first prove that the physical entry--exit map for positive slow parameter extends, uniformly to any prescribed finite order, to the singular parameter, including nonvertical endpoint fibers. In the exact moving penetration coordinate, composition with the grazing passage gives the extended Poincare displacement $\Delta(q,p)=S(q,p)+q_+^{3/2}K(\sqrt{q_+},q,p)$. Under a rank-two unfolding, we use the exact constant and linear coefficients of $\Delta$ as parameters and construct, for every sufficiently small fixed positive slow parameter, a unique nonpenetrating fold, a unique penetrating fold, and the grazing-incidence stratum. We give the complete marked chamber decomposition by nonpenetrating, grazing, and penetrating cycles and determine their stability from the Poincare multiplier. The open chambers contain zero or two cycles when the smooth and grazing coefficients have the same sign, and one or three when their signs are opposite; the corresponding cyclicity bounds are sharp corollaries. A compact polynomial family realizes both sign classes. In a cutoff-Gause family, interval certificates isolate one singular balanced-neutral point in each of two parameter boxes and verify the required signs and rank; an analytic local pullback transfers the diagram to the response parameters. Thus one local classification records cycle number, itinerary, and stability across the grazing and fold boundaries.

math.DS

Impedance Control of Ship-Borne Manipulators via Optimization-based Task-Space Inverse Dynamics

Ship-borne manipulators operating in maritime environments are subject to stochastic wave-induced base motions that introduce kinematic disturbances and dynamic coupling, degrading trajectory tracking accuracy and complicating safe, contact-rich manipulation. This paper proposes a torque-level optimization-based control framework that integrates high-precision trajectory tracking with task-space impedance for ship-borne manipulators. The controller is formulated using task-space inverse dynamics (TSID) and solved via quadratic programming to explicitly compensate for the dynamic coupling introduced by base motion. To enable accurate feedforward compensation, an error-state Kalman filter (ESKF) is developed to estimate the base state by fusing inertial measurements with end-effector pose feedback. The framework is validated in simulation and real-world experiments using a 7-DOF manipulator mounted on a 6-DOF Stewart platform. The proposed method reduces real-world end-effector position tracking error by over 25.7% compared with the best baseline. Furthermore, the controller enables dynamic peg-in-hole insertion with 1~mm clearance under base motion, increasing the success rate while reducing average contact forces by 45%, demonstrating precise and compliant manipulation in contact-rich environments.

cs.RO

Local Uniform Finite Cyclicity of the $H_{14}^{3}$ Semihyperbolic Hemicycle in Quadratic Systems

We prove local uniform finite cyclicity of the labelled $H_{14}^{3}$ hemicycle in a full neighborhood of its base field in the twelve-dimensional space of planar quadratic vector fields. Thus there is one fixed two-sided annular neighborhood of the compactified graphic in which the number of isolated limit cycles is bounded uniformly for every sufficiently close quadratic field. A local analytic slice theorem is part of the result: the displayed five-parameter source-normalized family is transverse at the $B=0$ field to the seven-dimensional action of affine phase changes and positive constant time rescalings. This removes the normalization and transfers the same cyclicity bound to the full quadratic coefficient space. The normalized problem combines a noncompact period annulus, two semihyperbolic endpoints at infinity, and a degeneration at the upper equatorial point. We replace the unavailable global Poincare map by finitely many stopped transition maps and prove an exact multiplicity-preserving correspondence between collar cycles and displacement zeros. At the noncompact source a matched nonlinear return on a common physical domain, controlled through six derivatives, gives a two-zero bound by Rolle's theorem. At the two saddle-nodes an exhaustive physical incidence and scale decomposition reduces all limits to source, mixed, hyperbolic, central, and lips estimates. A finite specialization argument then extends these bounds across coefficient and identity strata. The resulting bound is existential and is not claimed to be sharp. In the terminology of the quadratic finite-cyclicity program, this completes the labelled $H_{14}^{3}$ open case.

math.DS

SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation

Real-world robotic manipulation demands spatial grounding, task-aware reasoning, and precise control. Learning such capabilities becomes particularly challenging in the low-data regime. Prior methods often trade off scalable task-level reasoning and explicit physical structure: video-based approaches can drift geometrically over long horizons, 3D approaches often require depth sensing, and many flow/trajectory interfaces emphasize motion without an explicit RGB-only geometric representation. We introduce SSI-Policy, a modular framework built around a Structured Scene Interface (SSI) -- a unified, RGB-only intermediate representation that jointly encodes monocular depth features, language-grounded object layouts, and instruction-conditioned 2D motion trajectories. Critically, SSI is robot-agnostic and trainable from action-free video, decoupling perception from control so that the downstream policy can learn from few demonstrations. On the LIBERO benchmark with only 10 demonstrations per task, SSI-Policy improves over the strongest prior method by nearly 15\% and remains competitive with 50-demo methods that leverage large-scale external pretraining. Ablations show that geometric and motion cues provide complementary benefits within the shared interface. We further validate on 13 real-world tasks spanning spatial reasoning, cross-embodiment transfer, and contact-rich manipulation.

cs.RO

MinInter: Minimizing Trajectory Interpolation During Data Augmentation for Imitation Learning

Imitation learning enables robots to acquire complex manipulation skills from demonstrations, but its effectiveness is limited by the cost of collecting high-quality data. Trajectory-level data augmentation methods alleviate this challenge by recombining expert demonstrations under varied initial states. However, such methods typically insert interpolations or other non-expert transition segments between disjoint parts, and such non-expert segments could reduce the quality of the generated data. This paper introduces Minimizing Interpolation (MinInter), an effective trajectory selection method that, for each sampled initial configuration, chooses the source demonstration requiring the least interpolation to form a complete trajectory. By explicitly minimizing interpolations during data generation, MinInter produces higher-quality synthetic demonstrations while remaining compatible with existing data generation frameworks. Experiments on 12 manipulation tasks with 26 variants from the MimicGen benchmark show that MinInter consistently improves both data generation success rates and policy success rates, with the largest gains on contact-rich, long-horizon and high-variance settings. Compared to the recent SkillGen framework, MinInter achieves higher policy success rates despite its conceptual simplicity, underscoring the value of interpolation minimization for data augmentation.

cs.RO

FastViDAR: Real-Time Omnidirectional Depth Estimation via Alternative Hierarchical Attention

In this paper we propose FastViDAR, a novel framework that takes four fisheye camera inputs and produces a full $360^\circ$ depth map along with per-camera depth, fusion depth, and confidence estimates. Our main contributions are: (1) We introduce Alternative Hierarchical Attention (AHA) mechanism that efficiently fuses features across views through separate intra-frame and inter-frame windowed self-attention, achieving cross-view feature mixing with reduced overhead. (2) We propose a novel ERP fusion approach that projects multi-view depth estimates to a shared equirectangular coordinate system to obtain the final fusion depth. (3) We generate ERP image-depth pairs using HM3D and 2D3D-S datasets for comprehensive evaluation, demonstrating competitive zero-shot performance on real datasets while achieving up to 20 FPS on NVIDIA Orin NX embedded hardware. Project page: \href{https://3f7dfc.github.io/FastVidar/}{https://3f7dfc.github.io/FastVidar/}

cs.CV

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the state-of-the-art in this very dynamic area. Meanwhile, a growing number of testbeds have boosted the evolution of general-purpose large language models. Thus, this year's MARS2 focuses on real-world and specialized scenarios to broaden the multimodal reasoning applications of MLLMs. Our organizing team released two tailored datasets Lens and AdsQA as test sets, which support general reasoning in 12 daily scenarios and domain-specific reasoning in advertisement videos, respectively. We evaluated 40+ baselines that include both generalist MLLMs and task-specific models, and opened up three competition tracks, i.e., Visual Grounding in Real-world Scenarios (VG-RS), Visual Question Answering with Spatial Awareness (VQA-SA), and Visual Reasoning in Creative Advertisement Videos (VR-Ads). Finally, 76 teams from the renowned academic and industrial institutions have registered and 40+ valid submissions (out of 1200+) have been included in our ranking lists. Our datasets, code sets (40+ baselines and 15+ participants' methods), and rankings are publicly available on the MARS2 workshop website and our GitHub organization page https://github.com/mars2workshop/, where our updates and announcements of upcoming events will be continuously provided.

cs.CV

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

In the past year, video-based large language models (Video LLMs) have achieved impressive progress, particularly in their ability to process long videos through extremely extended context lengths. However, this comes at the cost of significantly increased computational overhead due to the massive number of visual tokens, making efficiency a major bottleneck. In this paper, we identify the root of this inefficiency as the high redundancy in video content. To address this, we propose a novel pooling strategy that enables aggressive token compression while retaining instruction-relevant visual semantics. Our model, Prompt-guided Pooling LLaVA (PPLLaVA), introduces three key components: a CLIP-based visual-prompt alignment module that identifies regions of interest based on user instructions, a prompt-guided pooling mechanism that adaptively compresses the visual sequence using convolution-style pooling, and a clip context extension module tailored for processing long and complex prompts in visual dialogues. With up to 18x token reduction, PPLLaVA maintains strong performance across tasks, achieving state-of-the-art results on diverse video understanding benchmarks-ranging from image-to-video tasks such as captioning and QA to long-form video reasoning-while significantly improving inference throughput. Codes have been available at https://github.com/farewellthree/PPLLaVA.

cs.CV

Coordinated Defense Allocation in Reach-Avoid Scenarios with Efficient Online Optimization

In this paper, we present a dual-layer online optimization strategy for defender robots operating in multiplayer reach-avoid games within general convex environments. Our goal is to intercept as many attacker robots as possible without prior knowledge of their strategies. To balance optimality and efficiency, our approach alternates between coordinating defender coalitions against individual attackers and allocating coalitions to attackers based on predicted single-attack coordination outcomes. We develop an online convex programming technique for single-attack defense coordination, which not only allows adaptability to joint states but also identifies the maximal region of initial joint states that guarantees successful attack interception. Our defense allocation algorithm utilizes a hierarchical iterative method to approximate integer linear programs with a monotonicity constraint, reducing computational burden while ensuring enhanced defense performance over time. Extensive simulations conducted in 2D and 3D environments validate the efficacy of our approach in comparison to state-of-the-art approaches, and show its applicability in wheeled mobile robots and quadcopters.

cs.RO