SearcharxivSearch

arXiv subjects

Ran Zheng

Publications and source records attributed to Ran Zheng.

9 recordsLinked to original sources

dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models

Vision-Language-Action (VLA) models have established a powerful paradigm for generalist robotic manipulation by grounding control into the semantic reasoning of VLMs. Prevailing architectures typically model actions continuously via diffusion or flow processes, or discretely through either autoregressive generation or parallel decoding. Recently, Discrete Diffusion VLAs (dVLAs) have emerged as a distinct alternative, unifying vision, language, and action into a single discrete token space via masked generative modeling. While combining iterative refinement with unified representations, its training has thus far been restricted to Supervised Fine-Tuning (SFT), leaving the potential of Reinforcement Learning (RL) for further policy refinement largely unexplored. A fundamental challenge in RL for dVLAs is that the marginal probability of the final action generated by dVLAs remains intractable. To solve this problem, we propose \textbf{dVLA-RL}, shifting the learning objective from the marginal action probability to the joint probability of the sampled generation path. Specifically, by modeling the denoising process as a Markov Decision Process (MDP), we mathematically formulate this path probability as a product of step-wise transitions. This trajectory-level objective provides a unified formulation that natively accommodates variable denoising steps. Leveraging this intrinsic fexibility, we introduce a unified step scheduling approach for complex multi-task learning, tailoring denoising steps to specific task complexities to maximize both success rates and computational effciency. Extensive evaluations demonstrate that our approach achieves a success rate of \textbf{99.7\%} on LIBERO. Furthermore, it establishes strong VLA-based results on RoboTwin 2.0 by delivering a \textbf{30.6\%} improvement over the SFT baseline, remaining competitive with strong World-Action Model baselines.

cs.RO

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing world-action models couple world prediction and action execution at the same temporal resolution, forcing the world branch to model near-term frame variations that are redundant and weakly informative. We posit that strictly binding world prediction and action execution to the same temporal rhythm may underutilize the potential of the video branch for embodied control. Therefore, we propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model built on a dual Diffusion Transformer (DiT) architecture that reorganizes world-action modeling around this temporal asymmetry. AHA-WAM instantiates the video DiT as a low-frequency world planner that maintains rolling key-value memory over past observations and exposes reusable layerwise latent context encoding long-horizon scene evolution, while a high-frequency action DiT executes short action chunks in closed loop by querying this context through layerwise joint attention. To support asynchronous execution, we introduce horizon-adaptive offset training and Observation-Guided Video-Context Routing (OVCR), which together let the action expert exploit long-horizon world context while remaining responsive to real-time execution state without rerunning the video DiT. Experiments on RoboTwin and real-world manipulation tasks show that AHA-WAM achieves state-of-the-art performance without any robot-data pretraining, attaining 92.80% average success on RoboTwin and 78.3% success across 4 real-world tasks, while reaching 24.17 Hz closed-loop control with a 4.59x speedup over Fast-WAM.

cs.RO

Two-Stage Active Distribution Network Voltage Control via LLM-RL Collaboration: A Hybrid Knowledge-Data-Driven Approach

The growing integration of distributed photovoltaics (PVs) into active distribution networks (ADNs) has exacerbated operational challenges, making it imperative to coordinate diverse equipment to mitigate voltage violations and enhance power quality. Although existing data-driven approaches have demonstrated effectiveness in the voltage control problem, they often require extensive trial-and-error exploration and struggle to incorporate heterogeneous information, such as day-ahead forecasts and semantic-based grid codes. Considering the operational scenarios and requirements in real-world ADNs, in this paper, we propose a hybrid knowledge-data-driven approach that leverages dynamic collaboration between a large language model (LLM) agent and a reinforcement learning (RL) agent to achieve two-stage voltage control. In the day-ahead stage, the LLM agent receives coarse region-level forecasts and generates scheduling strategies for on-load tap changer (OLTC) and shunt capacitors (SCs) to regulate the overall voltage profile. Then in the intra-day stage, based on accurate node-level measurements, the RL agent refines terminal voltages by deriving reactive power generation strategies for PV inverters. On top of the LLM-RL collaboration framework, we further propose a self-evolution mechanism for the LLM agent and a pretrain-finetune pipeline for the RL agent, effectively enhancing and coordinating the policies for both agents. The proposed approach not only aligns more closely with practical operational characteristics but also effectively utilizes the inherent knowledge and reasoning capabilities of the LLM agent, significantly improving training efficiency and voltage control performance. Comprehensive comparisons and ablation studies demonstrate the effectiveness of the proposed method.

eess.SY

Scanning-free three-dimensional fluorescent dipoles imaging by polarization self-interference digital holography (pSIDH)

Polarization microscopy provides insights into the structure and orientational organization of biomolecules and their architectures in cells. The above key functional signatures, which are natively 3D, can be only detected in 2D for a single measurement in conventional polarization microscopy. It is so far a challenging task to capture simultaneously the 3D structure and molecular orientation in a single frame of far-field intensity distribution, within the timescale of rapid-happened spatial organization events of bio-complexes. We report an optical imaging method called pSIDH, to encode multidimensional sample information includes 3D structures and dipole orientations, in their far-field fluorescence-self-interference pattern. The computational reconstruction from the holographic extracted complex-valued light field provides optical-aberration-corrected 3D polarization images of the sample. In pSIDH microscope incorporating planar liquid crystal lens and high numerical aperture objective, we demonstrate scanning-free 3D volumetric polarization imaging of fluorescently-labelled sample, with simultaneously computational-improved system measuring accuracy on the 3D spatial and polarization dimensions. The pSIDH imaging on phalloidin-fluorophore labelling U2OS cells provides rapid tools of capturing simultaneous the 3D structural details and spatial-averaged molecular orientation distributions of biological complex architectures such as actin filaments.

physics.optics

Beam test result and digitization of TaichuPix-3: A Monolithic Active Pixel Sensors for CEPC vertex detector

The Circular Electron-Positron Collider (CEPC), as the next-generation electron-positron collider, is tasked with advancing not only Higgs physics but also the discovery of new physics. Achieving these goals requires high-precision measurements of particles. Taichu seires, Monolithic Active Pixel Sensor (MAPS), a key component of the vertex detector for CEPC was designed to meet the CEPC's requirements. For the geometry of vertex detector is long barrel with no endcap, and current silicon lacks a complete digitization model, precise estimation of cluster size particularly causing by particle with large incident angle is needed. Testbeam results were conducted at the Beijing Synchrotron Radiation Facility (BSRF) to evaluate cluster size dependence on different incident angles and threshold settings. Experimental results confirmed that cluster size increases with incident angle. Simulations using the Allpix$^2$ framework replicated experimental trends at small angles but exhibited discrepancies at large angles, suggesting limitations in linear electric field assumptions and sensor thickness approximations. The results from both testbeam and simulations have provided insights into the performance of the TaichuPix chip at large incident angles, offering a crucial foundation for the establishment of a digital model and addressing the estimation of cluster size in the forward region of the long barrel. Furthermore, it offers valuable references for future iterations of TaichuPix, the development of digital models, and the simulation and estimation of the vertex detector's performance.

physics.ins-det

Beam test of a baseline vertex detector prototype for CEPC

The Circular Electron Positron Collider (CEPC) has been proposed to enable more thorough and precise measurements of the properties of Higgs, W, and Z bosons, as well as to search for new physics. In response to the stringent performance requirements of the vertex detector for the CEPC, a baseline vertex detector prototype was tested and characterized for the first time using a 6 GeV electron beam at DESY II Test Beam Line 21. The baseline vertex detector prototype is designed with a cylindrical barrel structure that contains six double-sided detector modules (ladders). Each side of the ladder includes TaichuPix-3 sensors based on Monolithic Active Pixel Sensor (MAPS) technology, a flexible printed circuit, and a carbon fiber support structure. Additionally, the readout electronics and the Data Acquisition system were also examined during this beam test. The performance of the prototype was evaluated using an electron beam that passed through six ladders in a perpendicular direction. The offline data analysis indicates a spatial resolution of about 5 um, with detection efficiency exceeding 99 % and an impact parameter resolution of about 5.1 um. These promising results from this baseline vertex detector prototype mark a significant step toward realizing the optimal vertex detector for the CEPC.

physics.ins-det

Beam test of a 180 nm CMOS Pixel Sensor for the CEPC vertex detector

The proposed Circular Electron Positron Collider (CEPC) imposes new challenges for the vertex detector in terms of pixel size and material budget. A Monolithic Active Pixel Sensor (MAPS) prototype called TaichuPix, based on a column drain readout architecture, has been developed to address the need for high spatial resolution. In order to evaluate the performance of the TaichuPix-3 chips, a beam test was carried out at DESY II TB21 in December 2022. Meanwhile, the Data Acquisition (DAQ) for a muti-plane configuration was tested during the beam test. This work presents the characterization of the TaichuPix-3 chips with two different processes, including cluster size, spatial resolution, and detection efficiency. The analysis results indicate the spatial resolution better than 5 $μm$ and the detection efficiency exceeds 99.5 % for both TaichuPix-3 chips with the two different processes.

physics.ins-det

An unified material interpolation for topology optimization of multi-materials

Topology optimization is one of the engineering tools for finding efficient design. For the material interpolation scheme, it is usual to employ the SIMP (Solid Isotropic Material with Penalization) or the homogenization based interpolation function for the parameterization of the material properties with respect to the design variables assigned to each finite element. For topology optimization with single material design, i.e., solid or void, the parameterization with 1 for solid and 0 for void becomes relatively straight forward using a polynomial function. For the case of multiple materials, some issues of the equality modeling of each material and \textcolor{red}{the clear 0, 1 result of each element for the topology optimization} issues become serious because of the curse of the dimension. To relieve these issues, this research proposes a new mapping based interpolation function for multi-material topology optimization. Unlike the polynomial based interpolation, this new interpolation is formulated by the ratio of the $p$-norm of the design variables to the 1-norm of the design variable multiplied by the design variable for a specific material. With this alternative mapping based interpolation function, each material are equally modeled and \textcolor{red}{ the clear 0, 1 result of each material for the multi-material topology optimization model} can be improved. This paper solves several topology optimization problems to prove the validity of the present interpolation function.

cs.CE

One-dimensional energy spectra in three-dimensional incompressible homogeneous isotropic turbulence

The paper investigates the detailed features of one-dimensional energy spectra in three-dimensional isotropic turbulence, based on the exact solution of Karman-Howarth equation. Particular interest will be paid on the degree to which spectral scaling can lead the spectral data to be collapsed. The theory appears to be consistent with the wealth of experimental data (G.Comte-Bellot and S. Corrsin, 1971) at least in low Taylor microscale Reynolds number.

physics.flu-dyn