SearcharxivSearch

arXiv subjects

Yicheng Ma

Publications and source records attributed to Yicheng Ma.

6 recordsLinked to original sources

SID: Sliding into Distribution for Robust Few-Demonstration Manipulation

Generalizing robotic manipulation across object poses, viewpoints, and dynamic disturbances is difficult, especially with only a few demonstrations. End-to-end visuomotor policies are expressive but data-hungry, while planning and optimization satisfy explicit constraints but do not directly capture the interaction strategies demonstrated by humans. We propose Sliding into Distribution (SID), a structured framework that learns an object-centric motion field from canonicalized demonstrations to iteratively slide the system toward the demonstrated manifold and into the reliable operating region of a lightweight egocentric execution policy, mitigating out-of-distribution (OOD) execution. The motion field provides large corrective motions when far from the demonstration manifold and naturally vanishes near convergence, enabling robust reaching under substantial pose and viewpoint shifts. Within the reached regime, an egocentric policy trained with conditioned flow matching performs task-specific manipulation, supported by kinematically consistent point-cloud reprojection augmentation that preserves action-observation consistency. Across six real-world tasks, SID achieves approximately 90% success under OOD initializations with only two demonstrations, with under a 10% drop under distractors and external disturbances. Overall, SID provides a new paradigm for few-shot manipulation: explicitly managing distribution shift via online distribution recovery.

cs.RO

FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation

Force sensing is a crucial modality for Vision-Language-Action (VLA) frameworks, as it enables fine-grained perception and dexterous manipulation in contact-rich tasks. We present Force-Distilled VLA (FD-VLA), a novel framework that integrates force awareness into contact-rich manipulation without relying on physical force sensors. The core of our approach is a Force Distillation Module (FDM), which distills force by mapping a learnable query token, conditioned on visual observations and robot states, into a predicted force token aligned with the latent representation of actual force signals. During inference, this distilled force token is injected into the pretrained VLM, enabling force-aware reasoning while preserving the integrity of its vision-language semantics. This design provides two key benefits: first, it allows practical deployment across a wide range of robots that lack expensive or fragile force-torque sensors, thereby reducing hardware cost and complexity; second, the FDM introduces an additional force-vision-state fusion prior to the VLM, which improves cross-modal alignment and enhances perception-action robustness in contact-rich scenarios. Surprisingly, our physical experiments show that the distilled force token outperforms direct sensor force measurements as well as other baselines, which highlights the effectiveness of this force-distilled VLA approach.

cs.RO

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

We present Uni-MoE 2.0 from the Lychee family. As a fully open-source omnimodal large model (OLM), it substantially advances Lychee's Uni-MoE series in language-centric multimodal understanding, reasoning, and generating. Based on the dense LLM, we build Uni-MoE-2.0-Omni from scratch through three core contributions: dynamic-capacity Mixture-of-Experts (MoE) design, a progressive training strategy enhanced with an iterative reinforcement strategy, and a carefully curated multimodal data matching technique. It is capable of omnimodal understanding, as well as generating images, text, and speech. Architecturally, our new MoE framework balances computational efficiency and capability for 10 cross-modal inputs using shared, routed, and null experts, while our Omni-Modality 3D RoPE ensures spatio-temporal cross-modality alignment in the self-attention layer. For training, following cross-modal pretraining, we use a progressive supervised fine-tuning strategy that activates modality-specific experts and is enhanced by balanced data composition and an iterative GSPO-DPO method to stabilise RL training and improve reasoning. Data-wise, the base model, trained on approximately 75B tokens of open-source multimodal data, is equipped with special speech and image generation tokens, allowing it to learn these generative tasks by conditioning its outputs on linguistic cues. Extensive evaluation across 85 benchmarks demonstrates that our model achieves SOTA or highly competitive performance against leading OLMs, surpassing Qwen2.5-Omni (trained with 1.2T tokens) on over 50 of 76 benchmarks. Key strengths include video understanding (+7% avg. of 8), omnimodallity understanding (+7% avg. of 4), and audiovisual reasoning (+4%). It also advances long-form speech processing (reducing WER by 4.2%) and leads in low-level image processing and controllable generation across 5 metrics.

cs.CL

Preferred Synthesis of Armchair Transition Metal Dichalcogenide Nanotubes

In this work, we present the synthesis of transition-metal dichalcogenide (TMDC) nanotubes with a preferred chiral angle. SnS2, MoS2, and WS2 are formed with high yield and structural purity inside the channels of boron nitride nanotubes. Atomic-resolution imaging, nano-area electron diffraction, and Circular Dichroism spectroscopy reveal that these synthesized TMDC nanotubes prefer to have an armchair configuration, with a probability up to 84%. Density functional theory reveals a negligible difference in the formation energy between armchair and zigzag nanotubes, suggesting that the chirality preference does not originate from the differences in structural stability. However, a detailed TEM investigation revealed that these TMDC nanotubes formed via a transition state of nanoribbons, and these nanoribbons are energetically more stable in a zigzag configuration. Subsequent machine learning potential molecular dynamics simulations verify that zigzag nanoribbons do roll up to form an armchair SnS2 nanotubes. Finally, this "zigzag nanoribbon to armchair nanotube" transition process is directly observed in real time by in-situ transmission electron microscopy. This work demonstrates the first, but likely general, experimental strategy for synthesizing chirality-preferred TMDC nanotubes.

cond-mat.mtrl-sci

Low-Temperature Synthesis of Weakly Confined Carbyne inside Single-Walled Carbon Nanotubes

Carbyne, a one-dimensional (1D) carbon allotrope with alternating triple and single bonds, has the highest known mechanical strength but is unstable to bending, limiting synthesis to short linear chains. Encapsulation within carbon nanotubes (CNTs) stabilizes carbyne, forming confined carbyne (CC), thus enabling further research concerning attractive 1D physics and materials properties of carbyne. While CC has been synthesized in multi-walled CNTs (MWCNTs) using the arc-discharge method and in double-walled CNTs (DWCNTs) via high-temperature high-vacuum (HTHV) method, synthesis in single-walled CNTs (SWCNTs) has been challenging due to their fragility under such conditions. In this work, we report a low-temperature method to synthesize CC inside SWCNTs (CC@SWCNT). By annealing SWCNTs containing ammonium deoxycholate (ADC) at 400{\deg}C, ADC is converted into CC without damaging the SWCNTs. Raman spectroscopy revealed a strong CC phonon (CC-mode) peak at 1860-1870 cm^-1, much stronger than the SWCNT G-band peak, confirming a high fraction of CC in the resulting material. The Raman mapping result showed the uniformity of the CC-mode signal across the entire film sample, proving the high efficiency of this method in synthesizing CC in every SWCNT of appropriate size. Notably, the CC-mode peaks of CC@SWCNT (above 1860 cm^-1) are higher than those reported in previous CC@CNT samples (mostly less than 1856 cm^-1). This is attributed to larger SWCNT diameters (over 0.95 nm) used in this study, compared to the typical 0.6-0.8 nm range. Larger diameters result in reduced confinement, allowing carbyne to closely resemble free-standing carbyne while remaining stabilized. This low-temperature synthesis of long-chain, nearly free-standing carbyne within large-diameter SWCNTs offers new opportunities for exploring 1D physics and the unique properties of carbyne for potential applications.

physics.chem-ph

Janus MoSSe nanotubes on one-dimensional SWCNT-BNNT van der Waals heterostructures

2D Janus TMDC layers with broken mirror symmetry exhibit giant Rashba splitting and unique excitonic behavior. For their 1D counterparts, the Janus nanotubes possess curvature, which introduce an additional degree of freedom to break the structural symmetry. This could potentially enhance these effects or even give rise to novel properties. In addition, Janus MSSe nanotubes (M=W, Mo), with diameters surpassing 40 {\AA} and Se positioned externally, consistently demonstrate lower energy states than their Janus monolayer counterparts. However, there have been limited studies on the preparation of Janus nanotubes, due to the synthesis challenge and limited sample quality. Here we first synthesized MoS2 nanotubes based on SWCNT-BNNT heterostructure and then explored the growth of Janus MoSSe nanotubes from MoS2 nanotubes with the assistance of H2 plasma at room temperature. The successful formation of the Janus structure was confirmed via Raman spectroscopy, and microscopic morphology and elemental distribution of the grown samples were further characterized. The synthesis of Janus MoSSe nanotubes based on SWCNT-BNNT enables the further exploration of novel properties in Janus TMDC nanotubes.

cond-mat.mtrl-sci