SearcharxivSearch

arXiv subjects

Yubo Zhang

Publications and source records attributed to Yubo Zhang.

At least 19 recordsLinked to original sources

HALO: A Physics-Aware LLM Agent Framework for Nanophotonic Design

Language models have recently been applied to nanophotonic design, but it remains unclear whether they can reliably translate optical objectives into simulation-ready designs, execute electromagnetic analysis, and revise decisions from numerical feedback. We introduce HALO, a physics-aware framework that couples language-model planners with typed design specifications, electromagnetic simulation, diagnostic evaluation, and optional reuse of prior failure trajectories in an iterative design loop. We further introduce HALO-Bench, a 52-task benchmark spanning lab-derived, paper-derived, and open-ended nanophotonic design tasks under a shared evaluation protocol. We compare three planner configurations: a Fixed Structured Workflow, an Autonomous Structured Agent using the same simulation interface, and an Autonomous Coding Agent that directly writes and executes simulation code. The Fixed Structured Workflow is the most token-efficient and exhibits no observed code- or path-level failures, while autonomous coding can achieve higher task success with stronger models at the cost of additional operational failures. We also study reuse of prior failed trajectories. On targeted multi-round tasks, retrieved failure feedback reduces both iterations to first success and total token use. These results clarify the tradeoffs between explicit interfaces, autonomous execution, and reusable design experience in scientific agents.

physics.optics

CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications

Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.

cs.CV

Learning-to-Transition for Large-scale and High-Order MIMO Detection

High-order multiple-input multiple-output (MIMO) detection requires efficient search over a large discrete symbol space while producing reliable soft information for channel decoding. This paper develops a learning-to-transition (L2T) framework that formulates MIMO detection as a stochastic sequence of complete-vector transitions. At each transition, a channel-coupled Transformer updates both the instance embedding and the sampling policy, while a blockwise autoregressive factorization captures inter-stream dependence with moderate sequential complexity. For hard-output detection, a transition network is applied recursively and trained through a residual-to-BER curriculum, which first learns the MIMO search geometry from the exact residual metric and then aligns the policy with transmitted-bit accuracy. For soft-output reception, the well-trained hard policy is cloned at the parameter level into every layer of an untied soft-input soft-output iterative detection and decoding (IDD) receiver. This tied-to-untied transfer preserves the learned zero-prior search dynamics while enabling layer- and round-specific specialization under decoder feedback. Within each IDD round, decoder priors tilt candidate generation according to Bayes' rule, and likelihood-weighted terminal hypotheses produce posterior and extrinsic log-likelihood ratios for LDPC decoding. A multi-stage training strategy further stabilizes the hard-to-soft transfer by progressively exposing the receiver to synthetic and in-loop decoder-generated priors.

cs.IT

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this gap with dense token supervision, yet applying it throughout training creates a different failure mode: the teacher is a biased, low-variance surrogate for the reward objective, so persistent imitation can oppose reward-improving updates after the policy becomes capable of producing successful trajectories. We introduce I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent. I-SDPO makes one routing decision per input instance and shares it across that instance's rollout group: all-incorrect groups use a privileged self-distillation objective, whereas any-success groups remain intact for GRPO. This design uses imitation only where group-relative rewards are uninformative. A local analysis characterizes when teacher and reward directions align and shows that a non-vanishing biased distillation weight induces an optimization bias floor. The routing rule automatically reduces the expected distillation rate as success probability rises, withdrawing teacher influence without a hand-designed schedule. On SciKnowEval, I-SDPO obtains the best result in all four scientific domains and improves average mean@16 accuracy from 56.67% with GRPO to 70.31%, with a maximum domain gain of 18.24 points.

cs.LG

Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems

We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-output (MU-MISO) systems. The proposed scheme integrates the ICL-Transformer backbone with the pilot encoder-decoder network (EDN) and the beamformer EDN. A crucial feature of our ICL network is that it can handle multiple channel models without retraining, enabled by the construction of model-specific context datasets. To improve convergence and robustness, we introduce three key innovations: (a) a curriculum learning (CL) strategy that smoothly transitions from supervised LMMSE-labeled imitation to unsupervised sum-rate maximization, (b) a self-evolving mechanism that dynamically expands and refines the context datasets for all channel models during CL-based training, and (c) a mismatch-aware extension that incorporates several mismatches into the general ICL framework and bypasses explicit channel calibrations. Ablation studies validate the effectiveness of the in-context architecture and enhanced training strategies. Simulation results over diverse communication environments show that the proposed scheme is able to rapidly adapt to both seen and unseen channel models without gradient-based parameter updates, and can mitigate the mismatch issues via intelligent context constructions. Furthermore, our scheme consistently outperforms the existing beamforming schemes under pilot-based settings, including the WMMSE benchmark and the recent Transformer-based methods.

cs.LG

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document layout elements, geometric distortions (\eg, paper warping and bending, perspective variations), and reading order within diverse layout structures. Existing approaches typically rely on fragmented multi-stage pipelines or computationally heavy generative Transformer architectures, leading to error propagation and limited efficiency. In this paper, we present RT-DocLayout, a highly efficient end-to-end framework for document layout analysis, designed as a front-end for document parsing tasks. The proposed model unifies classification, detection, pixel-level segmentation, and reading order prediction for layout elements within a single 33M-parameter architecture. Built upon the RT-DETR, our key contribution is a unified multi-task formulation within a single query-based decoder that simultaneously classifies, regresses bounding box, generates masks, and constructs relationship to reason reading order. By jointly learning geometric and structural representations, RT-DocLayout introduces multi-task optimization that substantially improves robustness under real-world document distortions. Extensive experiments on public benchmarks demonstrate state-of-the-art performance in document layout analysis while maintaining real-time inference speed(132.1 FPS). When coupled with downstream OCR engines, RT-DocLayout significantly improves full-document reconstruction quality, providing a scalable and practical foundation for real-world document intelligence systems.

cs.CV

PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks

Vision-Language Models (VLMs) have achieved impressive results on general vision-language tasks, yet they suffer from hallucination, imprecise localization, and prohibitive computational cost when applied to dedicated OCR scenarios. This paper presents PP-OCRv6, a lightweight OCR system that combines architectural innovation with data-centric optimization. PP-OCRv6 redesigns the backbone, detection neck, and recognition neck around a unified MetaFormer-style building block with structural reparameterization, decoupling spatial token mixing from channel mixing and supporting both tasks through task-specific stride configurations. Three model tiers (medium, small, tiny) share the same block primitives, covering deployment scenarios from server to edge. On our in-house benchmarks, PP-OCRv6_medium achieves 83.2% recognition accuracy and 86.2% detection Hmean, outperforming PP-OCRv5_server by +5.1% and +4.6% respectively while surpassing Qwen3-VL-235B, GPT-5.5, and Gemini-3.1-Pro with orders of magnitude fewer parameters. The tiny tier achieves 3.9$\times$ faster inference than PP-OCRv5_mobile on Intel Xeon CPU while maintaining comparable accuracy.

cs.CV

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining errors concentrate in under-optimized regions where model behavior is unstable, data coverage is sparse, or supervision is unreliable. Rather than expanding the training corpus indiscriminately, PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to these regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score of 96.33% on OmniDocBench v1.6, demonstrates strong competitiveness against top-tier VLMs, and provides a practical post-training recipe for the PaddleOCR-VL series.

cs.CV

Nitrogen doping induced metal-insulator transition with iso-symmetric character in rutile VO2

Metal-insulator transitions (MITs) in correlated oxides offer immense potential for next-generation Mottronic devices. However, their integration into practical applications is often hindered by the coupling of MITs with symmetry-lowering structural phase transitions, which limits switching speed and endurance. In this study, we engineered an iso-symmetric MIT on average in epitaxial rutile VO2 thin films via an in-situ nitrogen doping strategy. Nitrogen incorporation effectively suppresses V-V dimerization, enabling an iso-symmetric MIT, while preserving the original crystal symmetry. Furthermore, in-operando time-resolved optical reflectivity measurements revealed a shortened switching time in nitrogen-doped films, highlighting their enhanced performance. Our findings provide critical insights into the underlying mechanisms of MITs and introduce anion doping as a powerful tool for tailoring phase transitions in strongly correlated electron systems. This approach opens new avenues for the development of high-performance electronic and photonic devices.

cond-mat.mtrl-sci

Direct and Ambient Backscatter Communications with a Dual-Function Radar Transmitter

This work considers a system where a dual-function radar transmitter (source) performs direct communication with a reader while simultaneously enabling ambient backscatter communication from a tag. The source embeds its message into a coded pulse repeatedly transmitted over a frame, whereas the tag exploits the resulting environmental reverberation (clutter) as an ambient carrier to convey its own message. By leveraging the structure induced by the radar waveforms, we develop two signaling schemes. In the pilot-free scheme, the source and tag messages are conveyed through nonlinear vector modulation; the induced subspace structure enables both joint decoding, where all unknown quantities are simultaneously estimated, and disjoint decoding, where the tag codeword is recovered first, followed by the estimation of the source codeword and the channel vectors. In the pilot-aided scheme, pilot symbols and linearly modulated data symbols are embedded within each frame, enabling both non-iterative decoding based on pilot-derived channel estimates and iterative decoding via alternating channel estimation and data detection. We establish sufficient conditions on the source and tag codebooks that guarantee noiseless identifiability of the involved messages and channels. Finally, performance is evaluated in terms of source/tag error probabilities and channel-estimation accuracy, and the resulting system-level tradeoffs are discussed.

eess.SP

CT Saturation Detection and Compensation: A Hybrid Physical Model- and Data-Driven Method

Current transformer (CT) saturation is one of the dominant causes of relay protection devices' malfunctions, which pose a threat to the safe operation of the power system. To address this problem, we propose a hybrid physical model- and data-driven method. The method firstly detects the CT saturation and then compensates it to reproduce the real waveform. Considering the multi-factor and strong nonlinearity of CT saturation, a data-driven model, namely the Fully Convolutional Network (FCN), is built to detect the operation status of CT. As for the compensation, a physical model of short-circuit current is used for its conciseness and universality. Through tactfully integrating the data model and the physical model, the proposed method is endowed with two major merits: the arduous adjustment of universal thresholds and parameters in existing methods is avoided, and the deficiency in generalization and interpretability of the data-driven method is assuaged. Simulation and experimental results verify the effectiveness of the proposed method. Furthermore, its application potential to future protection is explored.

eess.SY

A Data-Aided Power Transformer Differential Protection without Inrush Blocking Module

When a slightly faulty transformer closes without load, the current waveform presents the coexistence of inrush and fault current. At this time, the inrush blocking module will block the relay, which may delay the removal of the slight fault and lead to more serious faults. To address this problem, this paper proposes a data-aided power transformer differential protection without inrush blocking module. The key to eliminating the negative influence of inrush current is to extract the fundamental component from the non-inrush part of the current waveform, which corresponds to the unsaturation period of the transformer core. Firstly, a data-aided module, namely an Attention module embedded Fully Convolutional Network (A-FCN), is built to distinguish the inrush and non-inrush parts of the current waveform. Then, a physical model of the current waveform is built for the non-inrush part, and the fundamental component is extracted by the nonlinear least square (NLS) algorithm. The proposed method can avoid the block of differential protections when inrush current occurs, which improves the sensitivity and rapidity of the relay, especially in the case of a weak internal fault hidden in inrush current. Finally, simulation and experimental data verify the effectiveness and generalization of the proposed method.

eess.SY

Model-Free Fast Frequency Support of Wind Farms for Tracking Optimal Frequency Trajectory

The fast frequency support (FFS) towards frequency trajectory optimization provides a system view for the frequency regulation of wind farms (WFs). However, the existing frequency trajectory optimization-based FFS generally relies on the accurate governor dynamics model of synchronous generators (SGs), which aggrandizes the difficulty of controller implementation. In this paper, a proportional-integral (PI) based FFS of WFs is designed for tracking the optimal frequency trajectory, which gets rid of the dependence on the governor model. Firstly, the prototypical PI-based FFS of WFs is proposed and its feasibility for tracking the optimal frequency trajectory is analyzed and demonstrated. Then, based on the "frequency-RoCoF" form of the optimal frequency trajectory, a more practical PI controller is constructed, avoiding the time dependence of the prototypical PI controller. Besides, an adaptive gain associated with PI parameters is designed for multi-WF coordination. Finally, the validity of the proposed method is verified in both the single-WF system and the multi-WF system.

eess.SY

A System-View Optimal Additional Active Power Control of Wind Turbines for Grid Frequency Support

Additional active power control (AAPC) of wind turbines (WTs) is essential to improve the transient frequency stability of low-inertia power systems. Most of the existing research has focused on imitating the frequency response of the synchronous generator (SG), known as virtual inertia control (VIC), but are such control laws optimal for the power systems? Inspired by this question, this paper proposes an optimal AAPC of WTs to maximize the frequency nadir post a major power deficit. By decoupling the WT response and the frequency dynamics, the optimal frequency trajectory is solved based on the trajectory model, and its universality is strictly proven. Then the optimal AAPC of WTs is constructed reversely based on the average system frequency (ASF) model with the optimal frequency trajectory as the desired control results. The proposed method can significantly improve the system frequency nadir. Meanwhile, the event insensitivity makes it can be deployed based on the on-line rolling update under a hypothetic disturbance, avoiding the heavy post-event computational burden. Finally, simulation results in a two-machine power system and the IEEE 39 bus power system verify the effectiveness of the optimal AAPC of WTs.

eess.SY

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to a quadratic increase in the number of vision tokens and significantly raises computational costs. We attribute this inefficiency to substantial visual regions redundancy in document images, like background. To tackle this, we propose PaddleOCR-VL, a novel coarse-to-fine architecture that focuses on semantically relevant regions while suppressing redundant ones, thereby improving both efficiency and performance. Specifically, we introduce a lightweight Valid Region Focus Module (VRFM) which leverages localization and contextual relationship prediction capabilities to identify valid vision tokens. Subsequently, we design and train a compact yet powerful 0.9B vision-language model (PaddleOCR-VL-0.9B) to perform detailed recognition, guided by VRFM outputs to avoid direct processing of the entire large image. Extensive experiments demonstrate that PaddleOCR-VL achieves state-of-the-art performance in both page-level parsing and element-level recognition. It significantly outperforms existing solutions, exhibits strong competitiveness against top-tier VLMs, and delivers fast inference while utilizing substantially fewer vision tokens and parameters, highlighting the effectiveness of targeted coarse-to-fine parsing for accurate and efficient document understanding. The source code and models are publicly available at https://github.com/PaddlePaddle/PaddleOCR.

cs.CV

PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks

The advent of "OCR 2.0" and large-scale vision-language models (VLMs) has set new benchmarks in text recognition. However, these unified architectures often come with significant computational demands, challenges in precise text localization within complex layouts, and a propensity for textual hallucinations. Revisiting the prevailing notion that model scale is the sole path to high accuracy, this paper introduces PP-OCRv5, a meticulously optimized, lightweight OCR system with merely 5 million parameters. We demonstrate that PP-OCRv5 achieves performance competitive with many billion-parameter VLMs on standard OCR benchmarks, while offering superior localization precision and reduced hallucinations. The cornerstone of our success lies not in architectural expansion but in a data-centric investigation. We systematically dissect the role of training data by quantifying three critical dimensions: data difficulty, data accuracy, and data diversity. Our extensive experiments reveal that with a sufficient volume of high-quality, accurately labeled, and diverse data, the performance ceiling for traditional, efficient two-stage OCR pipelines is far higher than commonly assumed. This work provides compelling evidence for the viability of lightweight, specialized models in the large-model era and offers practical insights into data curation for OCR. The source code and models are publicly available at https://github.com/PaddlePaddle/PaddleOCR.

cs.CV

Limits and Trade-Offs of Shift-Invariant Meta-Optical Encoders for Image Compression

Meta-optical encoders can reduce image data before electronic readout or transmission, but engineered point-spread functions (PSFs) do not automatically outperform conventional imaging. We study scene-agnostic, shift-invariant, linear optical encoders using both a fixed total-variation (TV) reconstruction backend and a learned YOLOv8 detection backend. Under a measurement-budget definition of compression ratio that counts all sensed samples across all channels, we compare lens imaging with spatial binning, positive random multi-channel PSFs, signed random kernels, and orthogonal multi-channel kernels. In the low-noise regime, lens-binning gives the highest reconstruction fidelity and strongest YOLOv8 detection metrics at the same compression ratio. Multi-channel encoders, however, degrade more slowly under measurement noise because the measurements are distributed across complementary channels. These results show that, for scene-agnostic incoherent imaging, engineered convolutional PSFs should be justified primarily by robustness, multiplexing, or downstream system constraints, rather than by an expectation that generic wavefront coding will outperform lens-based binning.

physics.optics

Elastic, Quasielastic, and Superelastic Electron Scattering from Thermal Lattice Distortions in Perfect Crystals

In standard treatments of electron transport, momentum relaxation in a perfect, defect-free crystal is linked with phonon creation or annihilation. In this work, we reconsider this problem for a finite, isolated crystal, retaining the lattice center-of-mass (recoil) degree of freedom and enforcing conservation of total mechanical momentum together with discrete crystal pseudomomentum. Starting from the density-density form of the electron-lattice interaction, we show that an electron in the interior of a perfect crystal admits elastic momentum-transfer channels in which total momentum is conserved by recoil of the lattice background without phonon excitation. These elastic channels can provide the leading contribution to momentum relaxation. We further identify mixed quasi-elastic and superelastic processes in which phonon occupations change but do not account entirely for the electron's momentum transfer. The elastic channels arise within the standard microscopic Hamiltonian and do not require additional disorder or defects. The resulting framework provides a complementary microscopic perspective on momentum relaxation in clean crystals and is consistent with experimental phenomena such as weak localization, quantum oscillations, ultrasonic attenuation, and the observed separation of momentum and energy relaxation times.

cond-mat.other