SearcharxivSearch

arXiv subjects

Shuai Shi

Publications and source records attributed to Shuai Shi.

At least 19 recordsLinked to original sources

SOLO: Stable Omni-terrain Long-Horizon Perceptive Humanoid Locomotion

Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as perception and control errors accumulate. We present SOLO, a unified framework addressing two compounding causes of this long-horizon fragility: dense terrain reconstruction smooths action-critical details, and pointwise imitation lacks temporal credit assignment. Its Query Reconstructor (QR) uses Fourier-encoded cell queries to retrieve spatially specific evidence from depth-proprioception tokens, preserving sharp terrain boundaries. Trajectory-Aware MSE (TA-MSE) Distillation adds next-state teacher-student disagreement to the PPO reward, enabling Generalized Advantage Estimation to propagate future disagreement penalties to preceding actions. In simulation, QR reduces height-map L1 error by factors of 3.3-4.0, while TA-MSE surpasses PPO and MSE+PPO in curriculum progression. On stress-test terrains, SOLO achieves 97.5% mean traversal success and 96% stepping-stone success, versus 75.0-75.6% and 0-3% for dense-reconstructor variants. Deployed zero-shot with only a chest-mounted depth camera and proprioception, SOLO completes a continuous 1.5-km outdoor route and an indoor mixed-terrain course. Project page: https://sunpihai-up.github.io/solo/

cs.RO

Luna-TTS Family Technical Report

Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Quantization (RVQ) token grid. We propose Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean. The family is built by progressive adaptation of a pretrained AR text LLM, from causal to bidirectional and finally to block-causal attention, and comprises two variants sharing a single tokenizer, data pipeline, and 0.6B backbone lineage. Luna-TTS is fully non-autoregressive: it generates the entire RVQ token grid in a fixed number of parallel refinement steps, with zero-shot voice cloning and speech editing arising natively as infilling. Luna-TTS Realtime, derived by continual training, is autoregressive over blocks of 32 codec frames (1.28s) while denoising each block in parallel; it supports KV-cached blockwise generation and incremental audio delivery, achieving an end-to-end RTF of 0.0240 and 41.6 ms local first-block latency under the warmed serving protocol. An annealed fine-tuning stage adds explicit control over emotion and non-verbal vocalizations (NVVs), and a reinforcement-learning stage applies GRPO with policy ratios computed over the realized denoising trajectory. On Seed-TTS-Eval, Luna-TTS achieves the best results on all four metrics among compared open-source and commercial systems (0.73 CER / 79.7 SIM on test-zh, 1.49 WER / 76.8 SIM on test-en); on the harder in-the-wild CV3-Eval, it posts the lowest Mandarin and English error rates in our comparison. Against leading commercial systems, it achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.

cs.SD

Observation of Moir\'e Time Crystal in Floquet-driven Rydberg Atomic Gases

A Moir\'e time crystal is a non-equilibrium quantum phase emerging from the coherent interference of two distinct frequencies, at least one being the intrinsic oscillation of a symmetry-broken time crystal. Its hallmark is an ultra-long beat period, reflecting a time-domain mapping of the Moir\'e fringes that arise from mismatched spatial lattices. However, to date, no experimental realization of such a Moir\'e time crystal has been reported. In this work, by applying a bichromatic driving field with two distinct frequencies, we demonstrate that the interplay between long-range Rydberg interactions and dissipation gives rise to a unique comb-like Moir\'e pattern characterized by a beat-note comb, which superimposes subharmonic periodicity and fundamental frequencies. This Moir\'e pattern formed by two mismatched drives is staggered in the spectrum as the frequency of one driver changes. We experimentally map the phase diagram of the system and identify a robust region where the Moir\'e temporal order persists against perturbations in laser detuning. The reported Moir\'e time crystal not only provides a controllable platform for exploring emergent slow-fast dynamics and synthetic space-time symmetries but also opens avenues for engineering complex temporal order in driven quantum many-body systems.

cond-mat.quant-gas

Toroidal helical pulses

Toroidal topologies and helicity are pervasive in nature and hold basic importance in scientific research. In particular, the interplay between these features gives rise to fascinating toroidal helical electromagnetic excitations. Here, we present a theoretical framework and experimental realization to introduce a family of toroidal helical pulses, exploring the intersection of the helicity and propagating toroidal modes. For this purpose, we propose a configuration combining a coaxial horn emitter and an equiangular spiral grating to directly generate such single-cycle pulses. In addition to their inherent non-transverse toroidal topology and space-time nonseparability, such pulses also possess controllable helicity. This work gives rise to a helical version of propagating toroidal electrodynamics, thereby paving the way for advanced applications, such as nontrivial light-matter interactions and data transfer.

physics.optics

SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding

Large language models incur high inference latency due to sequential autoregressive decoding. Speculative decoding alleviates this bottleneck by using a lightweight draft model to propose multiple tokens for batched verification. However, its adoption has been limited by the lack of high-quality draft models and scalable training infrastructure. We introduce SpecForge, an open-source, production-oriented framework for training speculative decoding models with full support for EAGLE-3. SpecForge incorporates target-draft decoupling, hybrid parallelism, optimized training kernels, and integration with production-grade inference engines, enabling up to 9.9x faster EAGLE-3 training for Qwen3-235B-A22B. In addition, we release SpecBundle, a suite of production-grade EAGLE-3 draft models trained with SpecForge for mainstream open-source LLMs. Through a systematic study of speculative decoding training recipes, SpecBundle addresses the scarcity of high-quality drafts in the community, and our draft models achieve up to 4.48x end-to-end inference speedup on SGLang, establishing SpecForge as a practical foundation for real-world speculative decoding deployment.

cs.LG

Topological Tunneling Magnetoresistance Driven by Type-II Weyl-Like States in the Room-Temperature Half-Metal Mn2PC Monolayer

We predict the tetragonal Mn2PC monolayer to be a room-temperature ferromagnetic half-metal with a Curie temperature of 554 K. The spin-up channel hosts type-II Weyl-like crossings at the Fermi level with highly anisotropic band dispersion, whereas the spin-down channel is a wide-gap semiconductor. Topological edge states obtained from tight-binding calculations confirm the non-trivial bulk topology. Spin-orbit coupling opens a small gap of 11.2 meV at the Weyl-like crossings, generating pronounced Berry curvature and a sizable anomalous Hall conductivity near the Fermi level. Based on these properties, we propose topological tunneling magnetoresistance in a Mn2PC-based magnetic tunnel junction: the parallel configuration conducts through fully spin-polarized Weyl-like carriers, while the antiparallel configuration is suppressed by the half-metallic gap, yielding a giant magnetoresistance ratio. The concurrent anomalous Hall effect in the conducting state provides an experimentally accessible signature of the topological carriers. These results identify the Mn2PC monolayer as a promising platform for room-temperature topological spintronic devices.

cond-mat.other

MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction

Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-like behaviors. However, the high dimensionality and intricate dynamics of humanoid robots make manual motion design impractical, leading to a heavy reliance on expensive motion capture (MoCap) data. These datasets are not only costly to acquire but also frequently lack the necessary geometric context of the surrounding physical environment. Consequently, existing motion synthesis frameworks often suffer from a decoupling of motion and scene, resulting in physical inconsistencies such as contact slippage or mesh penetration during terrain-aware tasks. In this work, we present MeshMimic, an innovative framework that bridges 3D scene reconstruction and embodied intelligence to enable humanoid robots to learn coupled "motion-terrain" interactions directly from video. By leveraging state-of-the-art 3D vision models, our framework precisely segments and reconstructs both human trajectories and the underlying 3D geometry of terrains and objects. We introduce an optimization algorithm based on kinematic consistency to extract high-quality motion data from noisy visual reconstructions, alongside a contact-invariant retargeting method that transfers human-environment interaction features to the humanoid agent. Experimental results demonstrate that MeshMimic achieves robust, highly dynamic performance across diverse and challenging terrains. Our approach proves that a low-cost pipeline utilizing only consumer-grade monocular sensors can facilitate the training of complex physical interactions, offering a scalable path toward the autonomous evolution of humanoid robots in unstructured environments.

cs.RO

Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

Humanoid robot technology is advancing rapidly, with manufacturers introducing diverse heterogeneous visual perception modules tailored to specific scenarios. Among various perception paradigms, occupancy-based representation has become widely recognized as particularly suitable for humanoid robots, as it provides both rich semantic and 3D geometric information essential for comprehensive environmental understanding. In this work, we present Humanoid Occupancy, a generalized multimodal occupancy perception system that integrates hardware and software components, data acquisition devices, and a dedicated annotation pipeline. Our framework employs advanced multi-modal fusion techniques to generate grid-based occupancy outputs encoding both occupancy status and semantic labels, thereby enabling holistic environmental understanding for downstream tasks such as task planning and navigation. To address the unique challenges of humanoid robots, we overcome issues such as kinematic interference and occlusion, and establish an effective sensor layout strategy. Furthermore, we have developed the first panoramic occupancy dataset specifically for humanoid robots, offering a valuable benchmark and resource for future research and development in this domain. The network architecture incorporates multi-modal feature fusion and temporal information integration to ensure robust perception. Overall, Humanoid Occupancy delivers effective environmental perception for humanoid robots and establishes a technical foundation for standardizing universal visual modules, paving the way for the widespread deployment of humanoid robots in complex real-world scenarios.

cs.RO

Occupancy World Model for Robots

Understanding and forecasting the scene evolutions deeply affect the exploration and decision of embodied agents. While traditional methods simulate scene evolutions through trajectory prediction of potential instances, current works use the occupancy world model as a generative framework for describing fine-grained overall scene dynamics. However, existing methods cluster on the outdoor structured road scenes, while ignoring the exploration of forecasting 3D occupancy scene evolutions for robots in indoor scenes. In this work, we explore a new framework for learning the scene evolutions of observed fine-grained occupancy and propose an occupancy world model based on the combined spatio-temporal receptive field and guided autoregressive transformer to forecast the scene evolutions, called RoboOccWorld. We propose the Conditional Causal State Attention (CCSA), which utilizes camera poses of next state as conditions to guide the autoregressive transformer to adapt and understand the indoor robotics scenarios. In order to effectively exploit the spatio-temporal cues from historical observations, Hybrid Spatio-Temporal Aggregation (HSTA) is proposed to obtain the combined spatio-temporal receptive field based on multi-scale spatio-temporal windows. In addition, we restructure the OccWorld-ScanNet benchmark based on local annotations to facilitate the evaluation of the indoor 3D occupancy scene evolution prediction task. Experimental results demonstrate that our RoboOccWorld outperforms state-of-the-art methods in indoor 3D occupancy scene evolution prediction task. The code will be released soon.

cs.CV

RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots

3D occupancy prediction enables the robots to obtain spatial fine-grained geometry and semantics of the surrounding scene, and has become an essential task for embodied perception. Existing methods based on 3D Gaussians instead of dense voxels do not effectively exploit the geometry and opacity properties of Gaussians, which limits the network's estimation of complex environments and also limits the description of the scene by 3D Gaussians. In this paper, we propose a 3D occupancy prediction method which enhances the geometric and semantic scene understanding for robots, dubbed RoboOcc. It utilizes the Opacity-guided Self-Encoder (OSE) to alleviate the semantic ambiguity of overlapping Gaussians and the Geometry-aware Cross-Encoder (GCE) to accomplish the fine-grained geometric modeling of the surrounding scene. We conduct extensive experiments on Occ-ScanNet and EmbodiedOcc-ScanNet datasets, and our RoboOcc achieves state-of the-art performance in both local and global camera settings. Further, in ablation studies of Gaussian parameters, the proposed RoboOcc outperforms the state-of-the-art methods by a large margin of (8.47, 6.27) in IoU and mIoU metric, respectively. The codes will be released soon.

cs.RO

Observation of helical pulses

Ultrafast spatiotemporal vortex pulses constitute a category within spatiotemporal topological waves. Nevertheless, the experimental realization of helical pulses single or few cycle short vortex pulses characterized by space time nonseparability remains elusive to date. Here, we introduce two complementary methods for experimentally generating such space time nonseparable helical pulses (SNHPs) in the optical and microwave spectral regimes. We achieve few cycle quasi linearly polarized SNHPs by decomposing the optical toroidal pulses into their polarization components. We also generated single cycle nontransverse SNHPs directly from a microwave ultrawideband spiral emitter. These approaches enable the experimental realization of SNHPs and provide a platform for further investigation into their properties and applications, such as nontrivial light-matter interactions and optical communications.

physics.optics

Double-Helix Singularity and Vortex-Antivortex Annihilation in Space-Time Helical Pulses

Topological structures reveal the hidden secrets and beauty in nature, such as the double helix in DNA, whilst, the manipula-tion of which in physical fields, especially in ultrafast struc-tured light, draw booming attention. Here we introduce a new family of spatiotemporal light fields, i.e. helical pulses, carry-ing sophisticated double-helix singularities in its electromag-netic topological structures. The helical pulses were solved from Maxwell's equation as chiral extensions of toroidal light pulses but with controlled angular momentum dependence. We unveil that the double helix singularities can maintain their topological invariance during propagation and the field exhibits paired generation and annihilation of vortices and antivortices in ultrafast space-time, so as to be potential information carriers beating previous conventional vortex structured light.

physics.optics

Hybrid electromagnetic toroidal vortices

The ubiquitous occurrence of toroidal vortices or vortex rings in fluid-dynamic scenarios in nature has garnered significant attention of scientific frontier, whilst, the electromagnetic counterparts of which were only proposed recently with two distinct manifestations: vector toroidal pulses [Nat. Photon. 16, 523 (2022)] and scalar phase toroidal vortices [Nat. Photon. 16, 519 (2022)]. This dichotomy in the understanding of toroidal vortex phenomena has prompted a reassessment of their fundamental nature. Herein, we theoretically propose a novel form of electromagnetic toroidal vortex solutions, that uniquely integrate both scalar and vector characteristics, challenging the prevailing notion of their mutual exclusivity. We also present the experimental generation of the hybrid toroidal vortex pulses by a compact coaxial horn emitter augmented with a metasurface. This methodology not only demonstrates the feasibility of creating such complex vortex structures but also endows the resulting pulses with unique properties, including the coexistence of transverse orbital angular momentum, electromagnetic vortex streets, and topological skyrmion textures. These attributes introduce new dimensions in topologically complex structured waves, opening avenues for enhanced free-space information transmission, topologically nontrivial light-matter interaction and microscopy techniques.

physics.optics

Propagation-invariant strongly longitudinally polarized toroidal pulses

Recent advancements in optical, terahertz, and microwave systems have unveiled non-transverse optical toroidal pulses characterized by skyrmionic topologies, fractal-like singularities, space-time nonseparability, and anapole-exciting ability. Despite this, the longitudinally polarized fields of canonical toroidal pulses notably lag behind their transverse counterparts in magnitude. Interestingly, although mushroom-cloud-like toroidal vortices with strong longitudinal fields are common in nature, they remain unexplored in the realm of electromagnetics. Here, we present strongly longitudinally polarized toroidal pulses (SLPTPs) which boast a longitudinal component amplitude exceeding that of the transverse component by over tenfold. This unique polarization property endows SLPTPs with robust propagation characteristics, showcasing nondiffracting behavior. The propagation-invariant strongly longitudinally polarized field holds promise for pioneering light-matter interactions, far-field superresolution microscopy, and high-capacity wireless communication utilizing three polarizations.

physics.optics

VMambaCC: A Visual State Space Model for Crowd Counting

As a deep learning model, Visual Mamba (VMamba) has a low computational complexity and a global receptive field, which has been successful applied to image classification and detection. To extend its applications, we apply VMamba to crowd counting and propose a novel VMambaCC (VMamba Crowd Counting) model. Naturally, VMambaCC inherits the merits of VMamba, or global modeling for images and low computational cost. Additionally, we design a Multi-head High-level Feature (MHF) attention mechanism for VMambaCC. MHF is a new attention mechanism that leverages high-level semantic features to augment low-level semantic features, thereby enhancing spatial feature representation with greater precision. Building upon MHF, we further present a High-level Semantic Supervised Feature Pyramid Network (HS2PFN) that progressively integrates and enhances high-level semantic information with low-level semantic information. Extensive experimental results on five public datasets validate the efficacy of our approach. For example, our method achieves a mean absolute error of 51.87 and a mean squared error of 81.3 on the ShangHaiTech\_PartA dataset. Our code is coming soon.

cs.CV

Free-Space Propagation and Skyrmion Topology of Toroidal Electromagnetic Pulses

Toroidal electromagnetic pulses have been recently reported as nontransverse, space-time nonseparable topological excitations of free space [Nat. Photon. 16, 523-528 (2022)]. However, their propagation dynamics and topological configurations have not been comprehensively experimentally characterized. Here, we report that microwave toroidal pulses can be launched by a broadband conical horn antenna. We experimentally map their skyrmionic textures and demonstrate how that during propagation the pulses evolves towards stronger space-time nonseparability and closer proximity to the canonical Hellwarth and Nouchi toroidal pulses.

physics.class-ph

Towards V2I Age-aware Fairness Access: A DQN Based Intelligent Vehicular Node Training and Test Method

Vehicles on the road exchange data with base station (BS) frequently through vehicle to infrastructure (V2I) communications to ensure the normal use of vehicular applications, where the IEEE 802.11 distributed coordination function (DCF) is employed to allocate a minimum contention window (MCW) for channel access. Each vehicle may change its MCW to achieve more access opportunities at the expense of others, which results in unfair communication performance. Moreover, the key access parameters MCW is the privacy information and each vehicle are not willing to share it with other vehicles. In this uncertain setting, age of information (AoI) is an important communication metric to measure the freshness of data, we design an intelligent vehicular node to learn the dynamic environment and predict the optimal MCW which can make it achieve age fairness. In order to allocate the optimal MCW for the vehicular node, we employ a learning algorithm to make a desirable decision by learning from replay history data. In particular, the algorithm is proposed by extending the traditional DQN training and testing method. Finally, by comparing with other methods, it is proved that the proposed DQN method can significantly improve the age fairness of the intelligent node.

cs.NI

A multi-band atomic candle with microwave-dressed Rydberg atoms

Stabilizing important physical quantities to atom-based standards lies at the heart of modern atomic, molecular and optical physics, and is widely applied to the field of precision metrology. Of particular importance is the atom-based microwave field amplitude stabilizer, the so-called atomic candle. Previous atomic candles are realized with atoms in their ground state, and hence suffer from the lack of frequency band tunability and small stabilization bandwidth, severely limiting their development and potential applications. To tackle these limitations, we employ microwave-dressed Rydberg atoms to realize a novel atomic candle that features multi-band frequency tunability and large stabilization bandwidth. We demonstrate amplitude stabilization of microwave field from C-band to Ka-band, which could be extended to quasi-DC and terahertz fields by exploring abundant Rydberg levels. Our atomic candle achieves stabilization bandwidth of 100 Hz, outperforming previous ones by more than two orders of magnitude. Our simulation indicates the stabilization bandwidth can be further increased up to 100 kHz. Our work paves a route to develop novel electric field control and applications with a noise-resilient, miniaturized, sensitive and broadband atomic candle.

physics.atom-ph