SearcharxivSearch

arXiv subjects

Shuai Yuan

Publications and source records attributed to Shuai Yuan.

At least 19 recordsLinked to original sources

A Finite-Entropy Criterion for the Entropic Conditional Central Limit Theorem

We prove a finite-entropy criterion for the entropic conditional central limit theorem. Let $(ξ_i,η_i)_{i\geq 1}$ be independent copies of a pair $(ξ,η)$, and set $W_n=n^{-1/2}\sum_{i=1}^n ξ_i$ and $\boldsymbolη_n=(η_1,\ldots,η_n)$. Under the assumptions that $\mathbb{E}\operatorname{Var}(ξ\midη)<\infty$ and that the conditional law of $ξ$ given $η$ is absolutely continuous almost surely, we show that $\mathbb{E}h(W_n\mid\boldsymbolη_n)$ converges to the Gaussian entropy $\frac12\log(2πeσ^2)$, where $σ^2=\mathbb{E}\operatorname{Var}(ξ\midη)$, if and only if $\mathbb{E}h(W_{n_0}\mid\boldsymbolη_{n_0})>-\infty$ for some $n_0$. The main technical ingredient is a continuity theorem for Fisher information under Gaussian smoothing, which allows us to replace the finite expected conditional Fisher-information assumption by a necessary and sufficient finite-entropy condition.

math.PR

Candidate for a Fractional Topological Insulator in Twisted MoTe2

The interplay among electronic correlation, topology, and time-reversal-symmetry (TRS) often leads to exotic quantum states of matter, as highlighted by the discoveries of fractional Chern insulators (FCIs) in twisted bilayer MoTe2 (tMoTe2). Among the FCIs in tMoTe2, the most robust is at a hole filling factor of v=-2/3 per moiré unit cell. Here, employing pump-probe circular dichroism (CD) measurement on tMoTe2 at twist angles (3.9 and 3.7 degrees), we show that a correlated state at v =-4/3 exhibits an unusual Ising antiferromagnet behavior. The v =-4/3 state with no net magnetization undergoes first order phase transitions at extremely low magnetic fields of ~ 2-6 mT to partially valley polarized (PVP) states. This behavior is notably absent for all other correlated states in tMoTe2 and also disappears for v =-4/3 at higher or lower twist angles (4.0 or 3.3 degree). The observed magnetic signature is consistent with a theoretically proposed fractional topological insulator (FTI), consisting of two copies of v =-2/3 FCIs with opposite chirality in the two K valleys. The experimental results are supported by interacting continuum model calculations that reveal the extreme closeness in energy ( < 1 meV) between the putative FTI and PVP states. Our findings present a candidate FTI with TRS and call for advanced transport and imaging measurements to establish the quantized helical edge modes.

cond-mat.str-el

Perturbation Power Selection for First-Error Delay Maximization in Enhanced SC Decoding

In this paper, we analyze the effect of perturbation power in delaying the first error position, i.e., the first information bit incorrectly decoded by the successive cancellation (SC) decoding. It is conducted over the finite-length perturbation-enhanced SC (PE-SC) decoding paradigm. We show that the FEP delaying probability exhibits a non-monotonic dependence on the perturbation power \(σ_{p}^{2}\). Based on this property, an efficient perturbation power selection algorithm that maximizes the delay probability is proposed to enhance the perturbation efficiency. It results in a more efficient perturbation power selection in finite-length PE-SC decoding.

cs.IT

MSP-Net: Manifold-Guided Spectral Prompt Network for Hyperspectral Object Tracking

Hyperspectral object tracking leverages abundant spectral information to provide unique advantages for target discrimination in complex scenes. However, existing methods typically treat hyperspectral images as multi-channel extensions of RGB images, performing feature fusion in fixed band order. This approach leads to models dependent on specific sensor configurations while neglecting manifold relationships between bands, making generalization to heterogeneous sensors difficult. Moreover, the discriminative contribution of bands dynamically changes with target attributes and scene variations, further limiting the representational capacity of static fusion strategies. To address this, we propose the Manifold-Guided Spectral Prompt Network (MSP-Net). This network first reconstructs band relationships and forms adaptive spectral grouping through graph-driven manifold routing, then jointly integrates grouped spectral statistics with template appearance to construct target-related dynamic conditional prompts, enhancing target features while suppressing background interference. Furthermore, as tracking progresses, spectral conditions continuously evolve based on intermediate target representations, enabling target prompts to adapt in real-time to appearance and scene changes. Meanwhile, reliable historical states are used to constrain target localization and scale fluctuations, significantly improving temporal stability in cross-sensor tracking. Experiments on HOT2020 and HOT2023 demonstrate that MSP-Net achieves AUC and Precision exceeding 0.80 and 0.96, respectively, exhibiting exceptional robustness under heterogeneous sensors, target deformation, and complex background conditions. The code will be released at https://github.com/GGML668897/MSP-Net.

cs.CV

Speculative Pipeline Decoding: Higher-Accuracy Drafting with Hidden Latency via Pipeline Parallelism

Speculative Decoding (SD) accelerates low-concurrency LLM inference with a draft-then-verify paradigm. Mainstream methods, however, rely on multi-token prediction, which incurs compounding prediction difficulty and exposed draft latency. We propose Speculative Pipeline Decoding (SPD), which partitions the target LLM into $n$ pipeline stages so that $n$ tokens of a single sequence advance in parallel. To keep the pipeline saturated, a Pipeline Draft Module (PDM) aggregates multi-depth target features to predict the next token and runs concurrently with each pipeline step, yielding bounded prediction difficulty, higher acceptance, and hidden draft latency. Experiments show that SPD achieves higher theoretical and wall-clock speedup than EAGLE-3 at moderate pipeline width, while more aggressive widths still leave room for further gains. Our code is available at https://github.com/yuyijiong/speculative_pipeline_decoding

cs.CL

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking

Dynamic target tracking is essential for Unmanned Aerial Vehicles (UAVs) operating in complex urban environments, where both the target and the camera viewpoint change continuously. Existing Vision-Language-Action (VLA) policies can track visible targets effectively, but their performance often degrades when buildings, vegetation, or roadside objects block the line of sight. During sustained occlusion, a policy may lose the target state, execute actions toward an incorrect region, and amplify this error through subsequent observations until re-acquisition becomes impossible. To this end, we present CosFly-VLA, a spatially aware VLA model that jointly grounds the target, estimates its visibility, and generates continuous flight actions through a structured prediction interface. To train this policy, we use a large-scale recipe over diverse data sources. Spatially Grounded Continued Pretraining (CPT) on a 500k mixed pool injects UAV-view depth, distance, and 3-D spatial reasoning. A three-stage Curriculum-based Supervised Fine-Tuning (SFT) process then specializes the tracker through multi-head warm-up followed by two-stage curriculum learning over natural and hard / long-occlusion data. Chain-of-Thought (CoT) training subsequently teaches recovery-oriented reasoning traces before structured answers. Finally, a closed-loop Reinforcement Learning (RL) stage optimizes tracking behavior with a multi-component reward covering stand-off tracking, grounding quality, collision avoidance, and task success. Relative to OpenVLA, CosFly-VLA-0.8B reduces open-loop Average Displacement Error (ADE) by 34.1% on seen-test and 35.3% on unseen-test. Closed-loop optimization improves Success Rate (SR) by 29.8% and 2.5%, respectively. These results demonstrate progress from visible-frame imitation toward spatially grounded action-closed-loop control, evaluated under a shared oracle state history.

cs.RO

OnePath: Efficient and Privacy-Preserving Decision Tree Inference in the Cloud

The vast storage capacity and computational power of cloud servers have led to the widespread outsourcing of machine learning inference services. While offering significant operational benefits, this practice also introduces privacy risks, such as the exposure of proprietary models and sensitive user data. In this paper, we present OnePath, a framework for secure and efficient decision tree inference in cloud environments. Unlike existing methods that traverse all internal nodes of a decision tree, our traversal protocol processes only the nodes on the prediction path, significantly improving inference efficiency while preserving privacy. To further optimize privacy and performance, OnePath is the first to employ functional encryption for evaluating decision tree nodes. Notably, our protocol enables both model providers and users to remain offline during the inference phase, offering a crucial advantage for practical deployment. We provide formal security analysis to demonstrate that OnePath provides comprehensive privacy protections during the model inference process. Extensive experimental results show that our approach processes query data in microseconds, highlighting its efficiency. OnePath offers a practical solution that strikes a balance between security and performance, making it a promising option for a wide range of cloud-based decision tree inference applications.

cs.CR

Electron-beam Writing of Spectrally Uniform Green Single-photon Emitters in Hexagonal Boron Nitride

Scalable quantum photonic technologies require single-photon emitters whose positions and emission energies can be engineered simultaneously. Hexagonal boron nitride (hBN) is an attractive room-temperature host, but deterministic creation of spectrally reproducible emitters remains challenging. Here, we use a standard scanning electron microscope as a direct-writing tool to activate bright green single-photon emitters in hBN at predefined sites, without ion implantation or post-fabrication thermal annealing. The written emitters exhibit reproducible zero-phonon-line emission centered near 536 nm, room-temperature antibunching with g(2)(0) as low as 0.08, high brightness, strong linear polarization, and stable emission. Thickness-dependent activation, stacking experiments, cathodoluminescence spectroscopy, and first-principles calculations support a carbon-related defect complex as the most plausible origin of the emission. As a proof of nanophotonic compatibility, we further activate emitters in a nanoparticle-on-mirror plasmonic nanocavity and observe photoluminescence enhancement accompanied by shortened emission lifetimes. These results establish electron-beam direct writing as a practical route to site-selective, spectrally uniform green quantum emitters in hBN, offering a promising basis for integrated room-temperature quantum photonic architectures.

physics.optics

Nonlinearity-Aware LoRA: Structured Gate Adaptation under Low-Rank Constraints

Low-rank adaptation (LoRA) is commonly viewed as an update-space approximation to full fine-tuning, yet this view is incomplete for self-gated Transformer feed-forward networks. In gated FFNs, a low-rank residual can change not only projected features but also the nonlinear selection weights that determine which channels contribute to the output. We formalize this effect as selection misalignment and connect it to the local effective homogeneity of self-gated activations. This motivates a nonlinearity-aware principle for parameter-efficient fine-tuning: low-rank updates should allocate capacity to gate channels whose nonlinear states remain responsive and should shape the temporal evolution of selection. We propose NA-LoRA, a training-only method with two lightweight mechanisms: a derivative-based temporal-importance mask for gate-related LoRA updates and an activation-specific step-scaling rule when a meaningful coarse effective-homogeneity partition is available. NA-LoRA adds no auxiliary loss and incurs no inference-time overhead. Experiments on language-model fine-tuning and vision-language transfer benchmarks show that NA-LoRA consistently improves over vanilla LoRA and is competitive with or better than strong PEFT variants.

cs.LG

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as undifferentiated conditioning signals, and existing benchmarks focus mainly on reconstruction accuracy rather than whether forecasts respond correctly to changed weather forcing.We introduce EO-WM, a video diffusion transformer for multispectral EO forecasting. EO-WM incorporates a physically informed conditioning framework that represents meteorological forcing through a climatological baseline, weather anomalies, and cumulative physical stress signals. Specifically, it separates baseline and anomaly through distinct conditioning pathways, and accumulates anomalous forcing over time to capture sustained heat and drought stress. To evaluate weather-response behavior beyond standard metrics, we introduce two diagnostic benchmarks: an Extreme Summer Benchmark for severity-aware prediction of vegetation degradation under extreme weather, and a Seasonal Matched-Pair Benchmark for testing response fidelity under changed weather forcing. Experiments show that EO-WM reduces the error in predicted Normalized Difference Vegetation Index (NDVI) decline amplitude by a relative 5.63% and improves directional hit rate by a relative 7.80%, while remaining competitive on standard pixel-level metrics. The benchmarks and model will be made open-source at https://github.com/Luo-Z13/EO-WM.

cs.AI

Electrically Programmable Correlated Topology and Magnetism in a Moiré Trilayer

Strong electron-electron interactions underlie a wide range of quantum many-body phenomena, including magnetism, superconductivity, and charge fractionalization. A central goal is to achieve in situ control over lattice geometry, bandwidth, and band topology within a single platform. Here we realize such an electrically programmable quantum many-body system in an alternating twisted trilayer MoTe$_2$, where an out-of-plane displacement field continuously modifies the layer polarization, effective lattice, and topology of the moiré bands. At zero displacement field, the system realizes a triangular lattice hosting a correlated insulator at one hole per moiré unit cell ($ν= -1$). Doping this state produces strongly asymmetric magnetic responses: double-exchange-like ferromagnetism for $|ν| > 1$, and signatures of spin polarons and antiferromagnetism for $|ν| < 1$. At large displacement field, interlayer hybridization reconstructs the electronic structure into a honeycomb lattice with a flat Chern band, supporting integer and fractional Chern insulators. Magneto-optical measurements further reveal the signatures of gap closure and Landau-level formation from a spin-polarized Fermi surface near the crossover between the two regimes. These results establish a unified, electrically tunable platform in which correlated magnetism and topological states emerge from a single controllable band structure.

cond-mat.mes-hall

SurroundNEXO: Ego-Centric Metric Bridging for Spatially Consistent Geometry in Autonomous Driving

Modern autonomous driving depends on accurate metric 3D understanding for perception, reconstruction, and planning, which in turn requires reliable multi-camera depth prediction. However, the outward-facing nature of vehicle-mounted surround-view camera rigs inherently limits visual overlap across views, challenging the correspondence-based assumptions that underpin conventional multi-view geometry. To bridge this gap, we present SurroundNEXO, named after the Spanish word nexo for a geometric link, a low-overlap multi-camera metric depth framework that grounds cross-view reasoning in ego-centric geometry rather than dense visual correspondences. Instead of directly enforcing early global fusion, SurroundNEXO first assigns image tokens globally comparable ego-frame viewing directions through Ego-Ray Positional Encoding, then uses sparse LiDAR measurements as metric anchors to propagate absolute scale cues, and finally expands feature interaction progressively from view-local modeling to decomposed spatio-temporal reasoning and global integration. This design enables metric-scale depth prediction with improved spatial consistency across weakly overlapping cameras. Across low-overlap autonomous driving benchmarks, including NuScenes, Waymo and DDAD, SurroundNEXO reduces single-view error by 33.2%, improves cross-view consistency by 10.5%, and enhances metric reconstruction quality by 25.6% compared with SOTA methods. It further remains robust under extremely sparse depth prompts and exhibits strong zero-shot generalization to unseen camera layouts.

cs.CV

STGBD-Net: Spatio-temporal Gradient Basis Decomposition Network for Infrared Small Target Detection

A key challenge in infrared small target detection (IRSTD) is that weak target signal responses are easily obscured by strong background clutter, frequently resulting in missed detections. While traditional gradient-based methods attempt to capture fine details, their robustness is limited by the static fusion of multi-directional gradient features. In this paper, we rethink feature fusion from the perspective of Basis Decomposition Theory and propose a novel framework that reformulates the process into an explicit and adaptive decomposition-and-reconstruction paradigm. Specifically, we introduce the Basis Decomposition Module (BDM) and its specialized variant, the Gradient Decomposition Module (GDM) for IRSTD. GDMs treat the normalized gradient features as basis vectors to reconstruct a new feature, thereby maintaining detailed structures and highlighting infrared small targets. By integrating GDMs into a lightweight three-stage U-Net, we develop two unified architectures: the Spatial Gradient Basis Decomposition Network for single-frame detection and the Spatio-temporal Gradient Basis Decomposition Network for multi-frame scenarios. Extensive experiments demonstrate that our networks achieve state-of-the-art (SOTA) performance across multiple benchmarks, offering a superior balance between detection accuracy and computational efficiency. Our codes will be made public at: https://github.com/greekinRoma/IRSTD_HC_Platform.

cs.CV

CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization

Recent aerial vision-language navigation (VLN) datasets have grown rapidly, but they primarily address goal-oriented navigation to static destinations, leaving UAV visual tracking -- continuously following a moving target while maintaining visibility -- largely without dedicated training data. We introduce CosFlyTrack, a large-scale multi-modal dataset and scalable generation pipeline for UAV visual tracking in urban environments. The dataset provides approximately 12,000 expert and perturbed UAV trajectories generated from 6,000 pedestrian paths, comprising 2.4 million timesteps (approximately 334 hours) with seven aligned data channels: RGB, metric depth, semantic segmentation, six-degree-of-freedom drone pose, target state with visibility flag, bilingual (Chinese-English) instructions, and trajectory-pair metadata. To generate high-quality expert trajectories, we develop MuCO, a multi-constraint optimizer that plans directly in continuous three-dimensional space with BVH-accelerated collision and visibility queries, jointly enforcing target visibility, viewpoint quality, collision avoidance, smoothness, and kinematic feasibility, avoiding the discretization artifacts and post-hoc smoothing of grid-based planners. Fine-tuning experiments on seven vision-language models show that CosFlyTrack improves tracking performance to 78.3 to 95.6 percent SR@1 meter, a 53 to 69 percentage point gain over zero-shot baselines, supporting the dataset as a training resource for dynamic target-following agents. The dataset is publicly available at https://huggingface.co/datasets/AutelRobotics/CosFly; evaluation scripts and pre-trained checkpoints are hosted at https://huggingface.co/AutelRobotics/CosFly-Track.

cs.RO

CosFly: Plan in the Matrix, Fly in the World

We present CosFly, a box-structured planning and multimodal simulation pipeline for aerial tracking, together with CosFly-Track, a large-scale UAV dataset for dynamic target tracking across diverse environments including urban centers, highways, rural landscapes, forests, and coastal towns. In our current implementation on CARLA, CosFly provides a modular 7-step construction pipeline that converts complex 3D worlds into structured obstacle representations for planning, then projects the resulting trajectories back into multi-modal sensor data -- including RGB images, high-precision depth maps, and semantic segmentation masks -- paired with natural language navigation instructions. A key feature is the support for configurable fixed-FOV zoom levels (one FOV setting drawn per trajectory and held constant throughout), enabling simulation of various focal lengths through camera-intrinsic adjustments. The pipeline covers the complete workflow from 3D map export through grid simplification, pedestrian and drone trajectory planning, multi-modal rendering with 6-DOF pose annotations, quality inspection, and teacher-student caption generation. We analyze two trajectory-planning paradigms for aerial target tracking: a conventional two-stage pipeline with front-end candidate generation and backend refinement, and a direct gradient-based formulation that optimizes multiple tracking constraints in a single objective. The public CosFly-Track release contains 250 validated trajectories and approximately 100,000 rendered images with complete 6-DOF drone pose annotations (position x, y, z and orientation yaw, pitch, roll). Together, the pipeline and dataset establish a scalable foundation for aerial-ground collaborative research, supporting dynamic target tracking, UAV navigation, and multi-modal perception across diverse environments.

cs.RO

Cascade of fractional quantum Hall states in 2D system

The observation of the fractional quantum Hall (FQH) effect in 2D electron gases ushered in investigations of topological phases driven by strong electron correlations. Their remarkable features include fractionalized elementary excitations, gapless boundary states, and non-trivial quantum entanglement patterns. Thanks to persistent efforts in the building of new platforms and making higher-quality samples, a diverse plethora of FQH states have been unveiled in experiments. We report a systematic study of ultrahigh-quality GaAs/AlGaAs quantum wells with mobility up to 3.7*10^7 cm^2/V/s using quantum transport measurements in nuclear adiabatic demagnetization and dilution refrigerators down to 1 mK. In addition to many FQH states that have already been identified in previous work, new longitudinal resistance dips are observed at filling factors 17/33 and 15/31. The application of an in-plane magnetic field causes disparate variations of the FQH states. The theoretical foundation of these states is discussed in the framework of composite fermion theory. While most fractions can be explained as non-interacting composite fermions forming integer quantum Hall states, a few states correspond to FQH states of composite fermions that arise from residual interaction between them. We summarize the observed fractions in the range of 0 < ν < 2 and propose a pattern to account for their experimental appearance that provides an intuitive picture about the relative strengths of different FQH states.

cond-mat.mes-hall

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications

Large language models (LLMs) hold great promise for business applications, yet business analysis remains inherently complex, demanding rigorous reasoning and the integration of diverse knowledge sources. Existing benchmarks typically target narrow tasks and thus leave a fundamental question unanswered: how can LLMs be reliably applied in business, and how are these applications grounded in underlying theoretical capabilities? To address this gap, we introduce BizCompass, a benchmark explicitly designed to connect theoretical foundations with practical business knowledge and applications. At the knowledge level, BizCompass covers four core domains--finance, economics, statistics, and operations management. At the application level, it structures tasks around three representative roles: the analyst, the trader, and the consultant. This dual-axis design not only exposes performance differences across realistic scenarios but also diagnoses which foundational capabilities enable or constrain success. We systematically evaluate both open-source and commercial LLMs, revealing how theoretical knowledge translates into practical performance in business. The results provide actionable insights for model selection and training optimization in real-world business contexts. All datasets and evaluation code are publicly released to support reproducibility and future research: https://bizcompass.dev.ypemc.com.

cs.CE

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap

Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on enabling UAVs to interpret high-level human commands and execute long-horizon tasks in complex 3D environments. This paper provides a comprehensive and structured survey of the field, from its formal task definition to the current state of the art. We establish a methodological taxonomy that charts the technological evolution from early modular and deep learning approaches to contemporary agentic systems driven by large foundation models, including Vision-Language Models (VLMs), Vision-Language-Action (VLA) models, and the emerging integration of generative world models with VLA architectures for physically-grounded reasoning. The survey systematically reviews the ecosystem of essential resources simulators, datasets, and evaluation metrics that facilitates standardized research. Furthermore, we conduct a critical analysis of the primary challenges impeding real-world deployment: the simulation-to-reality gap, robust perception in dynamic outdoor settings, reasoning with linguistic ambiguity, and the efficient deployment of large models on resource-constrained hardware. By synthesizing current benchmarks and limitations, this survey concludes by proposing a forward-looking research roadmap to guide future inquiry into key frontiers such as multi-agent swarm coordination and air-ground collaborative robotics.

cs.RO