SearcharxivSearch

arXiv subjects

Yiduo Wang

Publications and source records attributed to Yiduo Wang.

At least 19 recordsLinked to original sources

Evidence for Three-component Interlayer Coherent Exciton Condensation

Increasing the number of internal components in a quantum many-body system can host collective orders inaccessible to simpler settings. Quantum Hall bilayers provide a canonical realization of interlayer exciton condensation, yet extending such coherence across three independently addressable electronic fluids has remained elusive. Here we report evidence for three-component interlayer coherent exciton condensation in triple-layer graphene system. Using Rydberg excitons in an adjacent WSe2 monolayer as a layer-sensitive optical probe, we resolve interaction-induced incompressibility at zeroth-Landau-level crossings for all three pairwise layer combinations, establishing top-middle, middle-bottom and top-bottom exciton condensate channels within the same device. Independent control of displacement field and interlayer bias continuously tunes these pairwise states towards a regime where Landau levels from all three layers approach simultaneous degeneracy. At their convergence, incompressibility persists while the exciton energy and spectral weight evolve smoothly between the pairwise limits, suggesting coherent participation of all three layers in a single three-component state. More broadly, the ability to independently control layer potentials and engineer interlayer interactions establishes multilayer graphene as a programmable synthetic dimension for exploring higher-component quantum Hall order and simulating strongly correlated quantum matter.

cond-mat.mes-hall

Small-time annealed large deviations principle for one-dimensional diffusions in a random environment

In this paper, we establish a small-time annealed path large deviation principle for one-dimensional diffusions in a random environment associated with the generator ${\mathcal L}_W f(x)=e^{-\rho(x,W)}(e^{a(x,W)}f'(x))'$. The coefficients $\{\rho(x,\cdot):x\in\mathbb R\}$ and $\{a(x,\cdot):x\in\mathbb R\}$ are random. We assume that for each fixed realization of the environment, $\rho$ and $a$ are continuous and locally exponentially integrable, and that the support of the associated intrinsic coordinates is compact and non-collapsing. This framework includes the extensively studied Brox diffusion $dX_t=dB_t-\frac12\dot W(X_t)\,dt$, where $B$ is a standard Brownian motion and $W$ is an independent two-sided Brownian motion representing the environment. The It\^o--McKean representation of the diffusions and the estimates of the first exit probabilities derived via Moser iteration play a crucial role.

math.PR

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the energy-optimal frequencies of Attention and FFN (A/F) differ and vary with the inference phase, workload, and system configurations. However, runtime variability and independent A/F frequency control create a large search space and high communication overhead. To address these challenges, we present AFlex, a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving. AFlex introduces a global scheduler and a local operator-level dynamic voltage and frequency scaling (DVFS) controller to determine A/F resource allocations and frequencies. It further introduces an interleaved A/F pipeline with dynamic microbatch depth and adaptive request batching to reduce pipeline bubbles. We implement AFlex in SGLang and evaluate it on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8$\times$7B under production Conversation and Coding traces. \AFlex reduces energy per token by up to 49\% over state-of-the-art disaggregated serving and 48\% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.

cs.DC

Destructive interference of second harmonic generation in AA stacked MoTe$_2$/WSe$_2$

The stacking configuration of two-dimensional materials critically governs their optical and electronic responses. Monolayer transition-metal dichalcogenides (TMDC) lack inversion symmetry and exhibit exciton-enhanced second-harmonic generation (SHG). In TMDC bilayers, 60{\deg} (0{\deg}) stacking is conventionally expected to suppress (enhance) SHG owing to destructive (constructive) interference of the layer-resolved nonlinear polarizations. Here, we report an unconventional destructive SHG interference in nearly 0{\deg}-stacked (AA-stacked) MoTe2/WSe2 heterobilayers using two independent probes: atomic-resolution imaging and stacking-sensitive exciton hybridization measurements. Supported by ab initio GW and Bethe-Salpeter equation calculations, we show that distinct two-photon resonances associated with the WSe2 C exciton and the MoTe2 D exciton generate a nearly $\pi$ phase difference ($\Delta\phi$) in their second-order nonlinear susceptibilities $\chi^{(2)}$, leading to the anomalous destructive interference. We further demonstrate that in small-angle twisted MoTe2/WSe2, the SHG polarization state is governed by the interplay between twist angle $\alpha$ and phase difference $\Delta\phi$, and can be mapped onto trajectories on the Poincar\'e sphere. At excitation energies satisfying $\Delta\phi$ + 3$\alpha$ = 180{\deg}, the SHG output becomes nearly circularly polarized (ellipticity ~ 0.91) and undergoes an abrupt 90{\deg} azimuthal rotation, corresponding to a geometric polarization singularity in the parameter space. Our findings open new routes for exciton-resonance engineered nonlinear photonics and stacking-resolved optical functionality in moir\'e materials.

cond-mat.mes-hall

DynoJEPP: Joint Estimation, Prediction and Planning in Dynamic Environments

DynoJEPP is a factor-graph-based framework that jointly formulates and simultaneously optimizes estimation, prediction, and planning in dynamic environments. In conventional factor-graph-based approaches that jointly formulate estimation, prediction, and planning, information from prediction and planning feeds back into state estimation, yielding corrupted estimates, undesired behaviors, and unsafe plans. To address this, DynoJEPP introduces a novel directed factor that enforces directional information flow within the factor graph, preventing prediction and planning from corrupting state estimation. We evaluate the impact of directed factors on inter-module interactions during navigation in both static and dynamic environments. Our results demonstrate that these factors are critical for safe operation, as without them, the robot collides in the majority of experiments. Building on this, we further introduce Cooperative DynoJEPP, which enables the ego robot to incorporate cooperative object behavior into its prediction and trajectory planning.

cs.RO

MooD: Perception-Enhanced Efficient Affective Image Editing via Continuous Valence-Arousal Modeling

Affective Image Editing (AIE) aims to modify visual content to evoke targeted emotions. Although current approaches achieve impressive editing quality, they often overlook inference efficiency, which limits their applicability in computational social scenarios. Moreover, most methods depend on discrete emotion representations, which hinder the continuous modeling of complex human emotions and constrain expressive capabilities in interactive scenarios. To tackle these gaps, we propose MooD, the first framework that directly leverages continuous Valence-Arousal (VA) values as editing instruction for fine-grained and efficient AIE in computational social systems. Specifically, we first introduce a VA-Aware retrieval strategy to bridge vague affective values and detailed visual semantics. Building upon this, MooD integrates visual transfer and perception-enhanced semantic guidance to achieve controllable AIE. Furthermore, considering that existing VA-annotated datasets mainly focus on social scenarios and largely overlook natural scenes, we therefore construct AffectSet, a comprehensive VA-annotated dataset covering diverse scenarios, to support model optimization and evaluation. Extensive qualitative and quantitative experimental results demonstrate that our MooD achieves superior performance in both affective controllability and visual fidelity while maintaining high efficiency. A series of ablation studies further reveal the crucial factors of our design.

cs.CV

I/O Optimizations for Graph-Based Disk-Resident Approximate Nearest Neighbor Search: A Design Space Exploration

Approximate nearest neighbor (ANN) search on SSD-backed indexes is increasingly I/O-bound (I/O accounts for 70--90\% of query latency). We present an I/O-first framework for disk-based ANN that organizes techniques along three dimensions: memory layout, disk layout, and search algorithm. We introduce a page-level complexity model that explains how page locality and path length jointly determine page reads, and we validate the model empirically. Using consistent implementations across four public datasets, we quantify both single-factor effects and cross-dimensional synergies. We find that (i) memory-resident navigation and dynamic width provide the strongest standalone gains; (ii) page shuffle and page search are weak alone but complementary together; and (iii) a principled composition, OctopusANN, substantially reduces I/O and achieves 4.1--37.9\% higher throughput than the state-of-the-art system Starling and 87.5--149.5\% higher throughput than DiskANN at matched Recall@10=90\%. Finally, we distill actionable guidelines for selecting storage-centric or hybrid designs across diverse concurrency levels and accuracy constraints, advocating systematic composition rather than isolated tweaks when pushing the performance frontier of disk-based ANN.

cs.DB

RePose: A Real-Time 3D Human Pose Estimation and Biomechanical Analysis Framework for Rehabilitation

We propose a real-time 3D human pose estimation and motion analysis method termed RePose for rehabilitation training. It is capable of real-time monitoring and evaluation of patients'motion during rehabilitation, providing immediate feedback and guidance to assist patients in executing rehabilitation exercises correctly. Firstly, we introduce a unified pipeline for end-to-end real-time human pose estimation and motion analysis using RGB video input from multiple cameras which can be applied to the field of rehabilitation training. The pipeline can help to monitor and correct patients'actions, thus aiding them in regaining muscle strength and motor functions. Secondly, we propose a fast tracking method for medical rehabilitation scenarios with multiple-person interference, which requires less than 1ms for tracking for a single frame. Additionally, we modify SmoothNet for real-time posture estimation, effectively reducing pose estimation errors and restoring the patient's true motion state, making it visually smoother. Finally, we use Unity platform for real-time monitoring and evaluation of patients' motion during rehabilitation, and to display the muscle stress conditions to assist patients with their rehabilitation training.

cs.CV

Disentangling Hardness from Noise: An Uncertainty-Driven Model-Agnostic Framework for Long-Tailed Remote Sensing Classification

Long-Tailed distributions are pervasive in remote sensing due to the inherently imbalanced occurrence of grounded objects. However, a critical challenge remains largely overlooked, i.e., disentangling hard tail data samples from noisy ambiguous ones. Conventional methods often indiscriminately emphasize all low-confidence samples, leading to overfitting on noisy data. To bridge this gap, building upon Evidential Deep Learning, we propose a model-agnostic uncertainty-aware framework termed DUAL, which dynamically disentangles prediction uncertainty into Epistemic Uncertainty (EU) and Aleatoric Uncertainty (AU). Specifically, we introduce EU as an indicator of sample scarcity to guide a reweighting strategy for hard-to-learn tail samples, while leveraging AU to quantify data ambiguity, employing an adaptive label smoothing mechanism to suppress the impact of noise. Extensive experiments on multiple datasets across various backbones demonstrate the effectiveness and generalization of our framework, surpassing strong baselines such as TGN and SADE. Ablation studies provide further insights into the crucial choices of our design.

cs.CV

Online Dynamic SLAM with Incremental Smoothing and Mapping

Dynamic SLAM methods jointly estimate for the static and dynamic scene components, however existing approaches, while accurate, are computationally expensive and unsuitable for online applications. In this work, we present the first application of incremental optimisation techniques to Dynamic SLAM. We introduce a novel factor-graph formulation and system architecture designed to take advantage of existing incremental optimisation methods and support online estimation. On multiple datasets, we demonstrate that our method achieves equal to or better than state-of-the-art in camera pose and object motion accuracy. We further analyse the structural properties of our approach to demonstrate its scalability and provide insight regarding the challenges of solving Dynamic SLAM incrementally. Finally, we show that our formulation results in problem structure well-suited to incremental solvers, while our system architecture further enhances performance, achieving a 5x speed-up over existing methods.

cs.RO

Plasmon-driven Ultrafast and Highly Efficient Saturable Absorption for Ultrashort Pulse Generation Based on 2D V2C

Plasmon-driven ultrafast nonlinearities hold promise for advanced photonics but remain challenging to harness in two-dimensional materials at telecommunication wavelengths. Here, we demonstrate few-layer V2C MXene as a high-performance saturable absorber by leveraging its tailored surface plasmon resonance. Combining transient absorption spectroscopy and first-principles calculations, we unveil a plasmon-driven relaxation mechanism dominated by interfacial high-energy hot electron generation (~100 fs), enabling giant ultrafast nonlinearities. Crucially, at the communication band (1550 nm), V2C exhibits a high saturable absorption coefficient of -1.35 cm/GW. Integrating this into an erbium-doped fiber laser, we generate mode-locked pulses with a duration of 486 fs at 1569 nm, a 39.51 MHz repetition rate, and exceptional stability (92 dB SNR). This work establishes plasmonic MXenes as a paradigm for tailored ultrafast photonic devices.

physics.optics

DynoSAM: Open-Source Smoothing and Mapping Framework for Dynamic SLAM

Traditional Visual Simultaneous Localization and Mapping (vSLAM) systems focus solely on static scene structures, overlooking dynamic elements in the environment. Although effective for accurate visual odometry in complex scenarios, these methods discard crucial information about moving objects. By incorporating this information into a Dynamic SLAM framework, the motion of dynamic entities can be estimated, enhancing navigation whilst ensuring accurate localization. However, the fundamental formulation of Dynamic SLAM remains an open challenge, with no consensus on the optimal approach for accurate motion estimation within a SLAM pipeline. Therefore, we developed DynoSAM, an open-source framework for Dynamic SLAM that enables the efficient implementation, testing, and comparison of various Dynamic SLAM optimization formulations. DynoSAM integrates static and dynamic measurements into a unified optimization problem solved using factor graphs, simultaneously estimating camera poses, static scene, object motion or poses, and object structures. We evaluate DynoSAM across diverse simulated and real-world datasets, achieving state-of-the-art motion estimation in indoor and outdoor environments, with substantial improvements over existing systems. Additionally, we demonstrate DynoSAM utility in downstream applications, including 3D reconstruction of dynamic scenes and trajectory prediction, thereby showcasing potential for advancing dynamic object-aware SLAM systems. DynoSAM is open-sourced at https://github.com/ACFR-RPG/DynOSAM.

cs.RO

Observation of polaronic state assisted sub-bandgap saturable absorption

Polaronic effects involving stabilization of localized charge character by structural deformations and polarizations have attracted considerable investigations in soft lattice lead halide perovskites. However, the concept of polaron assisted nonlinear photonics remains largely unexplored, which has a wide range of applications from optoelectronics to telecommunications and quantum technologies. Here, we report the first observation of the polaronic state assisted saturable absorption through subbandgap excitation with a redshift exceeding 60 meV. By combining photoluminescence, transient absorption measurements and density functional theory calculations, we explicate that the anomalous nonlinear saturable absorption is caused by the transient picosecond timescale polaronic state formed by strong carrier exciton phonon coupling effect. The bandgap fluctuation can be further tuned through exciton phonon coupling of perovskites with different Young's modulus. This suggests that we can design targeted soft lattice lead halide perovskite with a specific structure to effectively manipulate exciton phonon coupling and exciton polaron formation. These findings profoundly expand our understanding of exciton polaronic nonlinear optics physics and provide an ideal platform for developing actively tunable nonlinear photonics applications.

physics.optics

DynORecon: Dynamic Object Reconstruction for Navigation

This paper presents DynORecon, a Dynamic Object Reconstruction system that leverages the information provided by Dynamic SLAM to simultaneously generate a volumetric map of observed moving entities while estimating free space to support navigation. By capitalising on the motion estimations provided by Dynamic SLAM, DynORecon continuously refines the representation of dynamic objects to eliminate residual artefacts from past observations and incrementally reconstructs each object, seamlessly integrating new observations to capture previously unseen structures. Our system is highly efficient (~20 FPS) and produces accurate (~10 cm) reconstructions of dynamic objects using simulated and real-world outdoor datasets.

cs.RO

Quasi-Distribution Appraisal Based on Piecewise B\'ezier Curves: An Objective Evaluation Method about Finite Element Analysis

A class of quasi-distribution evaluation criteria based on piecewise Bezier curves is proposed to address the issue of the inability to objectively evaluate finite element models. During the optimization design of mechanical parts, finite element modeling is performed on their stress deformation, and the mesh node shape variable values are converted into distribution histogram data for piecewise Bezier curve fitting. Being dealt with area normalization method, the fitting curve could be regarded as a kind of probability density function (PDF), and its variance could be used to evaluate the finite element modeling results. The situation with the minimum variance is the optimal choice for overall deformation. Numerical experiments have indicated that the new method demonstrated the intrinsic characteristics of the finite element models of difference mechanical parts. As an objective appraisal method for evaluating finite element models, it is both effective and feasible.

math.NA

The Importance of Coordinate Frames in Dynamic SLAM

Most Simultaneous localisation and mapping (SLAM) systems have traditionally assumed a static world, which does not align with real-world scenarios. To enable robots to safely navigate and plan in dynamic environments, it is essential to employ representations capable of handling moving objects. Dynamic SLAM is an emerging field in SLAM research as it improves the overall system accuracy while providing additional estimation of object motions. State-of-the-art literature informs two main formulations for Dynamic SLAM, representing dynamic object points in either the world or object coordinate frame. While expressing object points in a local reference frame may seem intuitive, it may not necessarily lead to the most accurate and robust solutions. This paper conducts and presents a thorough analysis of various Dynamic SLAM formulations, identifying the best approach to address the problem. To this end, we introduce a front-end agnostic framework using GTSAM that can be used to evaluate various Dynamic SLAM formulations.

cs.RO

3D Lidar Reconstruction with Probabilistic Depth Completion for Robotic Navigation

Safe motion planning in robotics requires planning into space which has been verified to be free of obstacles. However, obtaining such environment representations using lidars is challenging by virtue of the sparsity of their depth measurements. We present a learning-aided 3D lidar reconstruction framework that upsamples sparse lidar depth measurements with the aid of overlapping camera images so as to generate denser reconstructions with more definitively free space than can be achieved with the raw lidar measurements alone. We use a neural network with an encoder-decoder structure to predict dense depth images along with depth uncertainty estimates which are fused using a volumetric mapping system. We conduct experiments on real-world outdoor datasets captured using a handheld sensing device and a legged robot. Using input data from a 16-beam lidar mapping a building network, our experiments showed that the amount of estimated free space was increased by more than 40% with our approach. We also show that our approach trained on a synthetic dataset generalises well to real-world outdoor scenes without additional fine-tuning. Finally, we demonstrate how motion planning tasks can benefit from these denser reconstructions.

cs.RO

The Newer College Dataset: Handheld LiDAR, Inertial and Vision with Ground Truth

In this paper we present a large dataset with a variety of mobile mapping sensors collected using a handheld device carried at typical walking speeds for nearly 2.2 km through New College, Oxford. The dataset includes data from two commercially available devices - a stereoscopic-inertial camera and a multi-beam 3D LiDAR, which also provides inertial measurements. Additionally, we used a tripod-mounted survey grade LiDAR scanner to capture a detailed millimeter-accurate 3D map of the test location (containing $\sim$290 million points). Using the map we inferred centimeter-accurate 6 Degree of Freedom (DoF) ground truth for the position of the device for each LiDAR scan to enable better evaluation of LiDAR and vision localisation, mapping and reconstruction systems. This ground truth is the particular novel contribution of this dataset and we believe that it will enable systematic evaluation which many similar datasets have lacked. The dataset combines both built environments, open spaces and vegetated areas so as to test localization and mapping systems such as vision-based navigation, visual and LiDAR SLAM, 3D LIDAR reconstruction and appearance-based place recognition. The dataset is available at: ori.ox.ac.uk/datasets/newer-college-dataset

cs.RO