Searcharxiv⌕ Search

arXiv subjects

Yu Luo

Publications and source records attributed to Yu Luo.

At least 55 records · Page 3Linked to original sources

ML-ARIS: Multilayer Underwater Acoustic Reconfigurable Intelligent Surface with High-Resolution Reflection Control

This article introduces a multilayered acoustic reconfigurable intelligent surface (ML-ARIS) architecture designed for the next generation of underwater communications. ML-ARIS incorporates multiple layers of piezoelectric material in each acoustic reflector, with the load impedance of each layer independently adjustable via a control circuit. This design increases the flexibility in generating reflected signals with desired amplitudes and orthogonal phases, enabling passive synthetic reflection using a single acoustic reflector. Such a feature enables precise beam steering, enhancing sound levels in targeted directions while minimizing interference in surrounding environments. Extensive simulations and tank experiments were conducted to verify the feasibility of ML-ARIS. The experimental results indicate that implementing synthetic reflection with a multilayer structure is indeed practical in real-world scenarios, making it possible to use a single reflection unit to generate reflected waves with high-resolution amplitudes and phases.

eess.AS↗

High-fidelity Multi-view Normal Integration with Scale-encoded Neural Surface Representation

Previous multi-view normal integration methods typically sample a single ray per pixel, without considering the spatial area covered by each pixel, which varies with camera intrinsics and the camera-to-object distance. Consequently, when the target object is captured at different distances, the normals at corresponding pixels may differ across views. This multi-view surface normal inconsistency results in the blurring of high-frequency details in the reconstructed surface. To address this issue, we propose a scale-encoded neural surface representation that incorporates the pixel coverage area into the neural representation. By associating each 3D point with a spatial scale and calculating its normal from a hybrid grid-based encoding, our method effectively represents multi-scale surface normals captured at varying distances. Furthermore, to enable scale-aware surface reconstruction, we introduce a mesh extraction module that assigns an optimal local scale to each vertex based on the training observations. Experimental results demonstrate that our approach consistently yields high-fidelity surface reconstruction from normals observed at varying distances, outperforming existing multi-view normal integration methods.

cs.CV↗

ARIADNE: A Perception-Reasoning Synergy Framework for Trustworthy Coronary Angiography Analysis

Conventional pixel-wise loss functions fail to enforce topological constraints in coronary vessel segmentation, producing fragmented vascular trees despite high pixel-level accuracy. We present ARIADNE, a two-stage framework coupling preference-aligned perception with RL-based diagnostic reasoning for topologically coherent stenosis detection. The perception module employs DPO to fine-tune the Sa2VA vision-language foundation model using Betti number constraints as preference signals, aligning the policy toward geometrically complete vessel structures rather than pixel-wise overlap metrics. The reasoning module formulates stenosis localization as a Markov Decision Process with an explicit rejection mechanism that autonomously defers ambiguous anatomical candidates such as bifurcations and vessel crossings, shifting from coverage maximization to reliability optimization. On 1,400 clinical angiograms, ARIADNE achieves state-of-the-art centerline Dice of 0.838, reduces false positives by 41% compared to geometric baselines. External validation on multi-center benchmarks ARCADE and XCAD confirms generalization across acquisition protocols. This represents the first application of DPO for topological alignment in medical imaging, demonstrating that preference-based learning over structural constraints mitigates topological violations while maintaining diagnostic sensitivity in interventional cardiology workflows.

cs.CV↗

A Robust Geometric Distortion Solution for Main Survey Camera of CSST

The advancement in sensitivity and field of view of next-generation wide-field survey telescopes requires astrometric measurements with high precision, even in the presence of significant geometric distortions. To address this challenge, we develop a Weighted Polynomial Distortion Correction in 2-Phase (WPDC-2P) method. This approach enhances stellar cross-matching, incorporates distance-based weighting into the traditional polynomial fitting, and employs a look-up table to absorb the remaining distortion residuals. Validated on simulated data from the Main Survey Camera of the \emph{Chinese Space Station Survey Telescope} (CSST), incorporating geometric distortions up to approximately $200$ pixels, the method achieves astrometric standard deviation ranging from 0.013 to 0.107 pixels (0.03 pixels for the $g$-1 detector) across all 18 detectors. Under extreme crowding conditions (e.g., globular cluster NGC 2298), the astrometric precision for the $g$-1 detector reaches 0.05-pixel level within the central region ($r_d < 4000$), despite a centroiding precision of $\sim$0.04 pixels. When applied to the Beijing-Arizona Sky Survey data, for which the standard pipeline delivers an astrometric uncertainty of $\sim$20 mas, our method reduces the positional scatter to $ σ_{Δα}=5.494$ mas (0.01 pixels) and $ σ_{Δδ}=9.981$ mas (0.02 pixels) using only a weighted 3rd-order polynomial correction. The method has been integrated into the CSST data processing pipeline and is prepared for further refinement using on-orbit calibration data.

astro-ph.IM↗

Probing mesoscopic nonlocal screening in van der Waals heterostructures with polaritons

Predictive optical modelling of van der Waals (vdW) heterostructures is critical for meta-optics, near-field photonics and quantum technologies. At their buried interfaces, charge transfer and spatially extended screening challenge local descriptions based on layer-by-layer stacking of fixed permittivity tensors. However, such nonlocal corrections have been established mainly for plasmonic systems at ångström-nanometre scales and are often assumed negligible on optical-wavelength scales. Here we challenge this view by uncovering a mesoscopic nonlocal screening regime, extending up to ~140 nm, at buried charge-transfer interfaces in transition-metal dichalcogenide/α-molybdenum trioxide (TMDC/α-MoO3) phonon-polaritonic heterostructures. Using phonon polaritons as an ultrasensitive probe, we quantify charge transfer from polariton-wavelength shifts and find a thickness-independent saturated response as α-MoO3 is thinned. Rather than merely complicating optical modelling, this nonlocal saturation turns a design-level correction into an opportunity by yielding a transferable cross-material metric. Across more than 120 devices, this metric scales linearly with the work-function difference between the TMDC and α-MoO3. We further identify a lattice-mismatch-set energy threshold for charge transfer, revising Anderson-type band alignment for vdW interfaces.

physics.optics↗

$\textbf{Re}^{2}$: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving

Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning performance of large language models (LLMs) by increasing test-time compute. However, even after extensive RLVR training, such models still tend to generate unnecessary and low-quality steps in their chain-of-thought (CoT), leading to inefficient overthinking and lower answer quality. We show that when the initial direction or quality of the CoT is suboptimal, the model often fails to reach the correct answer, even after generating several times more tokens than when the initial CoT is well-initialized. To this end, we introduce Reinforcement Learning with Re-solving (Re$^2$), in which LLMs learn to flexibly abandon unproductive reasoning paths and restart the solution process when necessary, rather than always committing to a final answer. Re$^2$ applies pure reinforcement learning without any preliminary supervised fine-tuning, successfully amplifying the rare redo behavior in vanilla models from only 0.5% to over 30%. This leads to substantial performance gains over standard RLVR under the same training compute budget, and also demonstrates notable improvements in test-time performance as the number of samples increases.

cs.AI↗

Uncertainty-Aware Concept and Motion Segmentation for Semi-Supervised Angiography Videos

Segmentation of the main coronary artery from X-ray coronary angiography (XCA) sequences is crucial for the diagnosis of coronary artery diseases. However, this task is challenging due to issues such as blurred boundaries, inconsistent radiation contrast, complex motion patterns, and a lack of annotated images for training. Although Semi-Supervised Learning (SSL) can alleviate the annotation burden, conventional methods struggle with complicated temporal dynamics and unreliable uncertainty quantification. To address these challenges, we propose SAM3-based Teacher-student framework with Motion-Aware consistency and Progressive Confidence Regularization (SMART), a semi-supervised vessel segmentation approach for X-ray angiography videos. First, our method utilizes SAM3's unique promptable concept segmentation design and innovates a SAM3-based teacher-student framework to maximize the performance potential of both the teacher and the student. Second, we enhance segmentation by integrating the vessel mask warping technique and motion consistency loss to model complex vessel dynamics. To address the issue of unreliable teacher predictions caused by blurred boundaries and minimal contrast, we further propose a progressive confidence-aware consistency regularization to mitigate the risk of unreliable outputs. Extensive experiments on three datasets of XCA sequences from different institutions demonstrate that SMART achieves state-of-the-art performance while requiring significantly fewer annotations, making it particularly valuable for real-world clinical applications where labeled data is scarce. Our code is available at: https://github.com/qimingfan10/SMART.

cs.CV↗

RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesis

Large Language Models have evolved from single-round generators into long-horizon agents, capable of complex text synthesis scenarios. However, current evaluation frameworks lack the ability to assess the actual synthesis operations, such as outlining, drafting, and editing. Consequently, they fail to evaluate the actual and detailed capabilities of LLMs. To bridge this gap, we introduce RAVEL, an agentic framework that enables the LLM testers to autonomously plan and execute typical synthesis operations, including outlining, drafting, reviewing, and refining. Complementing this framework, we present C3EBench, a comprehensive benchmark comprising 1,258 samples derived from professional human writings. We utilize a "reverse-engineering" pipeline to isolate specific capabilities across four tasks: Cloze, Edit, Expand, and End-to-End. Through our analysis of 14 LLMs, we uncover that most LLMs struggle with tasks that demand contextual understanding under limited or under-specified instructions. By augmenting RAVEL with SOTA LLMs as operators, we find that such agentic text synthesis is dominated by the LLM's reasoning capability rather than raw generative capacity. Furthermore, we find that a strong reasoner can guide a weaker generator to yield higher-quality results, whereas the inverse does not hold. Our code and data are available at this link: https://github.com/ZhuoerFeng/RAVEL-Reasoning-Agents-Text-Eval.

cs.CL↗

RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models

Large language model alignment via reinforcement learning depends critically on reward function quality. However, static, domain-specific reward models are often costly to train and exhibit poor generalization in out-of-distribution scenarios encountered during RL iterations. We present RLAR (Reinforcement Learning from Agent Rewards), an agent-driven framework that dynamically assigns tailored reward functions to individual queries. Specifically, RLAR transforms reward acquisition into a dynamic tool synthesis and invocation task. It leverages LLM agents to autonomously retrieve optimal reward models from the Internet and synthesize programmatic verifiers through code generation. This allows the reward system to self-evolve with the shifting data distributions during training. Experimental results demonstrate that RLAR yields consistent performance gains ranging from 10 to 60 across mathematics, coding, translation, and dialogue tasks. On RewardBench-V2, RLAR significantly outperforms static baselines and approaches the performance upper bound, demonstrating superior generalization through dynamic reward orchestration. The data and code are available on this link: https://github.com/ZhuoerFeng/RLAR.

cs.CL↗

When to repeat a biomarker test? Decomposing sources of variation from conditionally repeated measurements

Repeating an imperfect biomarker test based on an initial result can introduce bias and influence misclassification risk. For example, in some blood donation settings, blood donors' hemoglobin is remeasured when the initial measurement falls below a minimum threshold for donor eligibility. This paper explores methods that use data resulting from processes with conditionally repeated biomarker measurement to decompose the variation in observed measurements of a continuous biomarker into population variability and variability arising from the measurement procedure. We present two frequentist approaches with analytical solutions, but these approaches perform poorly in a dataset of conditionally repeated blood donor hemoglobin measurements where normality assumptions are not met. We then develop a Bayesian hierarchical framework that allows for different distributional assumptions, which we apply to the blood donor hemoglobin dataset. Using a Bayesian hierarchical model that assumes normally distributed population hemoglobin and heavy tailed $t$-distributed measurement variation, we found that the total measurement variation accounted for 22\% of the total variance among females and 25\% among males, with population standard deviations of $1.07\, \rm g/dL$ for female donors and $1.28\, \rm g/dL$ for male donors. Our Bayesian framework can use data resulting from any clinical process with conditionally repeated biomarker measurements to estimate individuals' misclassification risk after one or more noisy continuous measurements and inform evidence-based conditional retesting decision rules.

stat.AP↗

FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching

Iterative generative policies, such as diffusion models and flow matching, offer superior expressivity for continuous control but complicate Maximum Entropy Reinforcement Learning because their action log-densities are not directly accessible. To address this, we propose Field Least-Energy Actor-Critic (FLAC), a likelihood-free framework that regulates policy stochasticity by penalizing the kinetic energy of the velocity field. Our key insight is to formulate policy optimization as a Generalized Schrödinger Bridge (GSB) problem relative to a high-entropy reference process (e.g., uniform). Under this view, the maximum-entropy principle emerges naturally as staying close to a high-entropy reference while optimizing return, without requiring explicit action densities. In this framework, kinetic energy serves as a physically grounded proxy for divergence from the reference: minimizing path-space energy bounds the deviation of the induced terminal action distribution. Building on this view, we derive an energy-regularized policy iteration scheme and a practical off-policy algorithm that automatically tunes the kinetic energy via a Lagrangian dual mechanism. Empirically, FLAC achieves superior or comparable performance on high-dimensional benchmarks relative to strong baselines, while avoiding explicit density estimation.

cs.LG↗

Two-parameter bipartite entanglement measure

Entanglement concurrence is an important bipartite entanglement measure that has found wide applications in quantum technologies. In this work, inspired by unified entropy, we introduce a two-parameter family of entanglement measures, referred to as the unified $(q,s)$-concurrence. Both the standard entanglement concurrence and the recently proposed $q$-concurrence emerge as special cases within this family. By combining the positive partial transposition and realignment criteria, we derive an analytical lower bound for this measure for arbitrary bipartite mixed states, revealing a connection to strong separability criteria. Explicit expressions are obtained for the unified $(q,s)$-concurrence in the cases of isotropic and Werner states under the constraint $q>1$ and $qs\geq 1$. Furthermore, we explore the monogamy properties of the unified $(q,s)$-concurrence for $q\geq 2$, $0\leq s\leq 1$ and $1\leq qs\leq 3$, in qubit systems. In addition, we derive an entanglement polygon inequality for the unified $(q,s)$-concurrence with $q\geq 1$ and $qs\geq 1$, which manifests the relationship among all the marginal entanglements in any multipartite qudit system.

quant-ph↗

Multipartite entanglement measures based on the thermodynamic framework

In this work, we introduce a unified method to characterize and measure multipartite entanglement using the framework of thermodynamics. A family of the new entanglement measures is proposed: \textit{ergotropic-gap concentratable entanglement}. Furthermore, we establish that ergotropic-gap concentratable entanglement constitutes a well-defined entanglement measure within a specific parameter regime, satisfying key properties including continuity, majorization monotonicity and monogamy. We demonstrate the utility of this measure by showing it effectively distinguishes between multi-qubit Greenberger-Horne-Zeilinger states and W states. It also proves effective in detecting entanglement in specific classes of four-partite star quantum network states.

quant-ph↗

Non-parametric Bayesian inference via loss functions under model misspecification

In the usual Bayesian setting, a full probabilistic model is required to link the data and parameters, and the form of this model and the inference and prediction mechanisms are specified via de Finetti's representation. In general, such a formulation is not robust to model misspecification of its component parts. An alternative approach is to draw inference based on loss functions, where the quantity of interest is defined as a minimizer of some expected loss, and to construct posterior distributions based on the loss-based formulation; this strategy underpins the construction of the Gibbs posterior. We develop a Bayesian non-parametric approach; specifically, we generalize the Bayesian bootstrap, and specify a Dirichlet process model for the distribution of the observables. We implement this using direct prior-to-posterior calculations, but also using predictive sampling. We also study the assessment of posterior validity for non-standard Bayesian calculations. We show that the developed non-standard Bayesian updating procedures yield valid posterior distributions in terms of consistency and asymptotic normality under model misspecification. Simulation studies show that the proposed methods can recover the true value of the parameter under misspecification.

stat.ME↗

Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning

On-policy reinforcement learning (RL), particularly Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), has become the dominant paradigm for fine-tuning large language models (LLMs). While policy ratio clipping stabilizes training, this heuristic hard constraint incurs a fundamental cost: it indiscriminately truncates gradients from high-return yet high-divergence actions, suppressing rare but highly informative "eureka moments" in complex reasoning. Moreover, once data becomes slightly stale, hard clipping renders it unusable, leading to severe sample inefficiency. In this work, we revisit the trust-region objective in policy optimization and show that explicitly constraining the \emph{variance (second central moment) of the policy ratio} provides a principled and smooth relaxation of hard clipping. This distributional constraint stabilizes policy updates while preserving gradient signals from valuable trajectories. Building on this insight, we propose $R^2VPO$ (Ratio-Variance Regularized Policy Optimization), a novel primal-dual framework that supports stable on-policy learning and enables principled off-policy data reuse by dynamically reweighting stale samples rather than discarding them. We extensively evaluate $R^2VPO$ on fine-tuning state-of-the-art LLMs, including DeepSeek-Distill-Qwen-1.5B and the openPangu-Embedded series (1B and 7B), across challenging mathematical reasoning benchmarks. Experimental results show that $R^2VPO$ consistently achieves superior asymptotic performance, with average relative gains of up to 17% over strong clipping-based baselines, while requiring approximately 50% fewer rollouts to reach convergence. These findings establish ratio-variance control as a promising direction for improving both stability and data efficiency in RL-based LLM alignment.

cs.LG↗

Towards Task-Oriented Flying: Framework, Infrastructure, and Principles

Deploying robot learning methods to aerial robots in unstructured environments remains both challenging and promising. While recent advances in deep reinforcement learning (DRL) have enabled end-to-end flight control, the field still lacks systematic design guidelines and a unified infrastructure to support reproducible training and real-world deployment. We present a task-oriented framework for end-to-end DRL in quadrotors that integrates design principles for complex task specification and reveals the interdependencies among simulated task definition, training design principles, and physical deployment. Our framework involves software infrastructure, hardware platforms, and open-source firmware to support a full-stack learning infrastructure and workflow. Extensive empirical results demonstrate robust flight and sim-to-real generalization under real-world disturbances. By reducing the entry barrier for deploying learning-based controllers on aerial robots, our work lays a practical foundation for advancing autonomous flight in dynamic and unstructured environments.

cs.RO↗

Family of two-parameter multipartite entanglement measures

Multipartite entanglement is regarded as a crucial physical resource in quantum network communication. However, due to the intrinsic complexity of quantum many-body systems, identifying a multipartite entanglement measure that is both efficiently computable and capable of accurately characterizing entanglement remains a challenging problem. To address these issues, we propose a family of two-parameter multipartite entanglement measures for mixed states, termed unified-entropy concentratable entanglements. Many well-known multipartite entanglement measures are recovered as special cases of this family of measures, such as the entanglement of formation and the concentratable entanglements introduced in [Phys. Rev. Lett. 127, 140501 (2021)]. We demonstrate that the unified-entropy concentratable entanglements constitutes a well-defined entanglement monotones, and establish several desirable properties it satisfies, such as subadditivity and continuity. We further investigate the ordering relations of unified-entropy concentratable entanglements and discuss how these quantities can be efficiently estimated on near-term quantum devices. As an application, we demonstrate that the unified-entropy concentratable entanglements can effectively distinguish between multi-qubit Greenberger-Horne-Zeilinger (GHZ) states and W states. The ordering relations of these entanglement measures are further validated using four-partite star quantum network states and four-qubit Dicke states. Moreover, we find that the unified-entropy concentratable entanglements exhibit greater sensitivity than the original concentratable entanglements in detecting certain four-partite star quantum network states.

quant-ph↗

Weight-based measure of quantum memory as a universal and operational benchmark

Quantum memory plays a critical role in quantum communication, sensing, and computation. However, studies on quantum memory under a unified benchmarking framework remain scarce. In this paper, we propose a weight-based quantifier as a benchmarking method to evaluate the performance advantage of quantum memory in nonlocal exclusion tasks. We establish a general lower bound for the weight-based measure of quantum memory. Moreover, this measure provides fundamental theoretical bounds for transforming a general channel into an ideal quantum memory. Finally, we present explicit calculations of the weight-based quantifier for various channels, including unitary channels, depolarizing channels, maximal replacement channels, stochastic damping channels, and erasure channels.

quant-ph↗