Searcharxiv⌕ Search

arXiv subjects

Zhiyuan Wang

Publications and source records attributed to Zhiyuan Wang.

At least 19 recordsLinked to original sources

High-Field Brightness Limits of Alkali-Antimonide Photocathodes

Pushing the brightness limit of electron sources requires simultaneously minimizing the intrinsic emittance and maximizing the accelerating field. Alkali-antimonide photocathodes exhibit excellent properties at low fields, yet their high-field photoemission physics remains poorly understood owing to acute vacuum sensitivity. Here, we demonstrate robust photoemission from alkali-antimonide photocathodes in a radio-frequency gun at peak fields exceeding 100 MV/m, with quantum efficiency maintained above 1% for more than two weeks, enabling the first systematic measurements of their high-field brightness limits. The measured field dependence of the intrinsic emittance provides direct insight into photoemission physics at high fields. Near-threshold photoemission further reduces the intrinsic emittance, increasing the attainable brightness, and reveals constraints imposed by surface roughness. This work opens the high-field regime for advanced photocathodes, extending the brightness limits of electron sources.

physics.acc-ph↗

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.

cs.CV↗

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embodied model whose asymmetric joint attention lets action tokens read semantic, current-visual, predicted-future, and action information at every layer while the perceptual experts retain their distinct roles. Without layer-wise supervision, EWAM develops an emergent depth-wise specialization: action queries attend mainly to vision-language features in shallow layers, to predicted future frames in intermediate layers, and to action tokens themselves in deep layers. This handoff replicates across tasks and is stable across denoising steps. Checkpoint tracking and causal interventions show that it is learned and that action generation depends on it. EWAM is pretrained in two separate regimes, one on cross-embodiment robot trajectories and one on human egocentric video. In simulation and real-robot experiments, it surpasses existing VLA, WAM, and hybrid baselines. Human egocentric data improve both cross-embodiment transfer and real-robot robustness, and subtask-phase supervision improves long-horizon completion. Together, these results suggest that unified embodied learning can induce an ordered internal progression from semantic understanding, through visual foresight, to action formation.

cs.RO↗

A compact vapor-cell optical frequency reference with fractional frequency instability around $10^{-16}$

Compact optical frequency reference with high stability is essential for field applications such as navigation and geodesy, yet vapor cell systems have remained confined to fractional instabilities over $10^{-15}$. Here, we report a molecular iodine reference that reaches an instability of $7 \times 10^{-16}$ at 1000 s and operates at the $10^{-16}$ level from 200 to 2000 s, surpassing the best reported vapor cell standards by approximately a factor of three. This achievement is enabled by a monolithic, drift immune spectroscopic unit bonded to an ultra low expansion glass substrate with precision control of key parameters.The entire system occupies only 25 L.The system achieves $5 \times 10^{-15}$ instability at 1 s and reaches the $10^{-16}$ level over the 200 to 2000 s averaging-time range, representing the first medium term stability at the $10^{-16}$ level from a compact, field ready vapor-cell reference. Our work demonstrates that $10^{-16}$ instability can be engineered into portable systems, opening a path to high precision time-keeping beyond the laboratory.

physics.atom-ph↗

Gravitational wave signal and noise response of an optically levitated sensor in a Fabry-Pérot cavity

Optically levitated sensors inside a Fabry-Pérot cavity have been proposed for high-frequency gravitational-wave (GW) detection, though their configuration for gravitational wave sensitivity exhibits counterintuitive features. We provide a new detailed general relativistic derivation of the interaction between a gravitational wave and a levitated object in an optical cavity, demonstrating gauge independence of the observable response. We find a strong asymmetric dependence of the strain signal on trap position, maximized when the sensor is located near the input mirror, and provide an in-depth explanation of its origin from multiple gauge perspectives. A key new result of this work is the consequence of this asymmetry on the noise coupling: the coupling of input-mirror displacements to the strain signal can be highly suppressed relative to that of end-mirror displacements and common-mode mirror motion. These results clarify the physical origin of the gravitational wave interaction with such a sensor and establish crucial design principles for optical levitation based high-frequency GW detectors.

gr-qc↗

ForeSightGuide: An Anticipatory Framework toward Accurate and Low-Redundancy Guidance for the Visually Impaired

Electronic travel aids are pivotal for the independent mobility of the visually impaired. While Vision-Language Models (VLMs) offer rich environmental understanding, they often suffer from excessive false positives in dynamic scenarios, leading to cognitive overload. To address this, we present ForeSightGuide, an anticipatory assistive guidance framework that couples semantic scene understanding with predictive hazard assessment. Unlike reactive systems, ForeSightGuide leverages the reasoning capabilities of VLMs to anticipate obstacle motion, effectively filtering out non-threatening objects to provide concise, actionable guidance. To validate our approach, we introduce a novel dataset captured in complex, dynamic real-world traffic scenes, designed to benchmark predictive capabilities. Extensive experiments on both public benchmarks and our proposed dataset demonstrate that ForeSightGuide achieves state-of-the-art performance. Notably, it significantly mitigates information overload by reducing redundant alerts to 0.299 per guidance output while maintaining a low missed-hazard rate of 0.112, proving its efficacy for safe walking assistance.

cs.CV↗

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may return graded rewards that reflect partial success. We diagnose this mismatch on fixed responses from two representative long-context evidence-aggregation tasks. Across longer input ranges, trajectory-level OPD scores become progressively less aligned with verifier rewards, indicating teacher-verifier disagreement. Motivated by this observation, we introduce Group-Calibrated On-Policy Distillation (GC-OPD). GC-OPD separately normalizes verifier rewards and trajectory-level OPD scores within each rollout group and uses their difference as a signed teacher-verifier disagreement residual. Relative-advantage-based credit assignment (RACA) distributes this trajectory-level residual across tokens according to their relative OPD advantages while preserving the original OPD signal. Across five long-context benchmarks, post-training with GC-OPD raises the five-benchmark averages of the official Qwen3-4B and Qwen3-8B checkpoints from 29.08 to 40.47 and from 35.12 to 44.65, respectively. Vanilla OPD reaches 39.31 and 43.56 under the same setup. Controlled ablations show that the signed residual is more effective than either an additional OPD-derived term or direct group-normalized verifier reward addition, while RACA further improves over uniform token allocation. Together, these results demonstrate that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance. Code is available at https://github.com/SolereZhang/GC-OPD.

cs.LG↗

PULSE: Agentic Investigation with Passive Sensing for Proactive Affective Intervention in Cancer Survivorship

Cancer survivors face elevated rates of depression, anxiety, and emotional distress, yet self-report may be unavailable at some moments when support is relevant, a challenge we term the diary paradox. We present PULSE, a system for agentic sensing investigation: LLM agents equipped with eight purpose-built tools query smartphone sensing data, compare current behavior with personal baselines, and retrieve outcome-labeled historical cases. Rather than receiving only a fixed feature summary, agents choose which modalities and time windows to inspect. We evaluate PULSE through a two-by-two evaluation design crossing system architecture (structured single-pass vs. multi-turn agentic) with concurrent input modality (no current diary vs. sensing plus current diary) on 50 cancer survivors. The agentic multimodal condition achieves balanced accuracy of 0.743 for emotion-regulation desire; the agentic no-current-diary condition achieves 0.713 for self-reported intervention availability. This is a system-level comparison because the architecture conditions also differ in tool-mediated information access. The results provide a retrospective benchmark for interactive sensing investigation and motivate prospective evaluation at diary non-response moments.

cs.HC↗

Evolving Parallel Algorithm Portfolios via Potential-Aware Instance Generation with LLMs

The Automatic Construction of Portfolios via Large Language Models (LLM-ACP) suffers from poor generalization in practical few-shot scenarios when solving complex combinatorial optimization problems. Instance and algorithm co-evolution frameworks address this by expanding the training dataset with generated hard instances on which the current algorithm portfolio underperforms, thereby enhancing generalization. However, this paradigm faces two critical limitations: evaluating instance hardness relies on high-quality reference solutions, and single-mode generation patterns limit instance diversity. To overcome these limitations, we introduce the Potential-aware Instance and Algorithm Co-evolution (PIAC) framework. Our core contribution is twofold. First, we propose potential gain, a novel metric that eliminates the need for reference solutions. This metric estimates generalization gain by perturbing the generated algorithms and assessing their improvement potential on generated problem instances. Second, PIAC leverages LLMs to synthesize diverse instance mutators, exploring a broader region of the problem-instance space and thereby enhancing the portfolio's generalization capabilities. Given that perturbation spaces vary across different algorithms, we instantiate our framework on Greedy Constructive, Ant Colony Optimization, and Guided Local Search algorithmic backbones. Comprehensive evaluations on the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP) across six distinct data distributions demonstrate that PIAC consistently outperforms state-of-the-art LLM-ACP baselines, notably achieving a 19.76% relative improvement for TSP Greedy Constructive portfolios.

cs.AI↗

Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagation and can eventually match the performance of gradient-based reinforcement learning (RL). However, resource-constrained settings typically offer only a few GPUs, so the high GPU-hour requirements of ES translate into prohibitively long training times. To address this, we introduce Cooperative Parameter-subspace Evolution Strategy (CoPES), a cooperative coevolutionary method that decomposes the full parameter space into lower-dimensional subspaces and searches over them cooperatively to improve optimization efficiency. We post-train a Qwen3.5-4B tool-using agent for the math task and evaluate it on five benchmarks of varying difficulty. Under the GPU-hour budget of full-parameter GRPO's best validation checkpoint, CoPES recovers 92% of GRPO's validation-accuracy gain, versus 67% for standard ES, while its theoretical GPU memory requirement is less than one-eighth that of full-parameter GRPO. It consistently outperforms standard ES and LoRA-based GRPO on all evaluated pass@k metrics across the five benchmarks. Additional experiments further show the advantage of CoPES on the question-answering task. These results demonstrate an improved trade-off between memory requirements and training time for agentic LLM post-training under resource constraints. The code is open-sourced in https://github.com/MetaronWang/CoPES

cs.AI↗

On $R$-parastatistics I: Foundation

Parastatistics is an exotic type of exchange statistics beyond fermions and bosons. Paraparticles transform in higher dimensional representations of the exchange symmetry group, analogous to non-Abelian anyons, yet consistently defined in any dimension. Although paraparticles have long been proposed, they were widely believed to be physically equivalent to fermions or bosons. Nevertheless, a recent paper proposed a different theory, called $R$-parastatistics, and demonstrated that nontrivial $R$-paraparticles can emerge as quasiparticles in condensed matter systems, and are observably distinct from both fermions and bosons. This paper develops the theoretical foundation and several extensions of $R$-parastatistics, with particular emphasis on its observable consequences. Central to this paper is a general theory of local observables extending the basic family introduced before. First, we define local observables that distinguish particle types. Second, we formulate local observables at special point defects that probe the internal indices of $R$-paraparticles, crucial for observing $R$-parastatistics and for the proposed applications in quantum information. Third, we introduce local observables that create or annihilate particle-antiparticle pairs, important for building a relativistic quantum field theory for $R$-paraparticles. We further introduce generalized hidden symmetries that act on internal indices of $R$-paraparticles while preserving the local observable algebra, providing a basis for proving local indistinguishability and for connecting to a categorical description of $R$-paraparticles. This work sets a solid theoretical foundation for understanding the fundamental physical properties of $R$-paraparticles and pave the way for finding them in nature.

quant-ph↗

Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. To address this issue, we introduce Neo-Classic, an evaluation benchmark that combines a constructionist Out-of-Sample (OOS) dataset with a suite of reverse understanding probes. Unlike traditional benchmarks that rely on verification or generation over historical corpora, Neo-Classic comprises strictly metrical poetry authored by contemporary experts, reducing the possibility of direct retrieval. We evaluate state-of-the-art models, including Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2, across five behavioral probes designed to test hierarchical constraint satisfaction. Our results reveal two primary limitations. First, a performance gap of 20 to 50 percent emerges when models transition from historical to contemporary texts. Second, models exhibit substantial difficulties in discourse-level ordering tasks, with standard accuracy remaining low (0 to 13 percent). Although expert-level guidance improves the performance of reasoning-enhanced models to 36 percent, a notable gap with human experts persists. These findings suggest that while current LLMs capture local formal patterns, they struggle with global hierarchical planning required for robust Linguistic-Aesthetic Reasoning.

cs.CL↗

MEGO: Learning Mixture-of-Experts for General-Purpose Binary Optimization

Discrete optimization is ubiquitous in science and engineering. The vast array of existing discrete optimization problems, coupled with the continuous emergence of new ones, necessitates off-the-shelf optimizers capable of generating high-quality solutions for a large variety of optimization problems. This article introduces MEGO, a novel general-purpose neural optimizer for binary optimization under the black-box setting, intended for broad applicability across diverse binary optimization problem classes with minimal problem-specific customization. MEGO comprises a mixture-of-experts trained without domain knowledge. When presented with a new problem instance to solve, it employs a routing policy to dynamically activate the most relevant expert models to generate high-quality solutions. The strong generalization capability of MEGO is demonstrated on six problem classes from different disciplines, including classic problems and real-world applications. Trained solely on classic problems, MEGO effectively generalizes to unseen and complex real-world problem classes, significantly outperforming widely-used general-purpose optimizers in both solution quality and efficiency. Furthermore, MEGO provides a computational approach for quantifying similarity between optimization problems and classifying them, which is fundamentally different from the conventional analysis-based problem classification.

cs.NE↗

Optimal displacement detection of arbitrarily-shaped levitated dielectric objects using optical radiation

Optically-levitated dielectric objects are promising for precision force, acceleration, torque, and rotation sensing due to their extreme environmental decoupling. While many levitated optomechanics experiments employ spherical objects, for some applications non-spherical geometries offer advantages. For example, rod-shaped or dumbbell shaped particles have been demonstrated for torque and rotation sensing and high aspect ratio plate-like particles can exhibit reduced photon recoil heating and may be useful for high-frequency gravitational wave detection or as high bandwidth accelerometers. To achieve optimal sensitivity, cooling, and quantum control in these systems, it is beneficial to achieve optimal displacement detection using scattered light. We describe and numerically implement a method based on Fisher information that is applicable to suspended particles of arbitrary geometry. We demonstrate the agreement between our method and prior methods employed for spherical particles, both in the Rayleigh and Lorentz-Mie regimes. As practical examples, we analyze the optical detection limits of an optically-levitated high-aspect-ratio disc-like dielectric object and a rod-shaped object for configurations recently realized in experimental work.

physics.optics↗

Bridging the Post-discharge Gap: A Traceable Multi-agent Framework for Safe and Continuous Care

Post-discharge clinical follow-up is critical for maintaining continuity of care and mitigating long-term health risks. However, traditional follow-up paradigms suffer from shortage of health workforce, fragmented patient histories, and information silos across clinical departments. While large language models have demonstrated potential in medical question-answering, their deployment in continuous care is hindered by hallucination risks and a fundamental inability to reason over longitudinal, patient-specific constraints. Here we present Healink, a memory-enhanced multi-agent framework to support AI-assisted post-discharge follow-up by generating prescription-grounded, traceable responses that improved completeness and perceived clinical utility in retrospective and physician-blinded evaluations. The architecture seamlessly integrates a triage routing mechanism, a unified memory enhancement module utilizing a robust relational database for optimal latency, and a strict constraint-based retrieval-augmented generation engine. By vectorizing historical clinical records and employing weighted similarity functions across diverse phenotypic and intervention dimensions, Healink ensures precise inter-patient and intra-patient case matching while actively preventing cross-departmental drug conflicts. We evaluated Healink on a dataset comprising 400 continuous and 85 highly complex real-world follow-up cases, alongside the webMedQA benchmark. In a rigorous single-blind evaluation conducted by clinical experts, the framework outperformed human physician baselines in both authoritativeness and clinical safety. By generating a traceable, white-box evidence chain, Healink provides a scalable, safe, and highly effective paradigm for intelligent patient management, ultimately enhancing societal healthcare outcomes.

cs.MA↗

A persistent-homology-Gaussian prior for solving infinite-dimensional Bayesian inverse scattering problems

Bayesian inference methods have been developed to address inverse problems in function spaces where the unknown parameters are of infinite dimension. However, conventional Gaussian priors remain inadequate for reconstructing discontinuous or sharply varying target functions encountered in practical applications like obstacle reconstruction. Although hybrid priors have emerged as a promising solution, significant challenges remain in developing theoretically rigorous and computationally tractable frameworks in engineering applications. To address these issues, we propose a persistent-homology-Gaussian (PHG) prior for solving the acoustic obstacle scattering inverse problem in the infinite-dimensional Bayesian setting, which combines a weighted persistence-based regularization term with a periodic Gaussian reference measure through a Gibbs tilt. Then, the complex boundary is represented by a log-radial function on the unit circle, so that the reconstruction from far-field data is formulated as a function-space inverse problem. The well-posedness of the resulting posterior measure is established in the Hellinger, total variation, and Wasserstein-\(p\) metrics. Furthermore, the convergence of finite-dimensional posterior approximations is obtained, and posterior sampling is performed by a preconditioned Crank--Nicolson (pCN) method. Numerical experiments show that the proposed PHG prior yields accurate and stable reconstructions under more extensive noisy conditions, providing explicit control of multiscale topological features and better performance compared to other conventional priors.

math.NA↗

Dual Dimensionality for Local and Global Attention

Decoder-only Transformers compute attention over the KV cache of preceding tokens. Keys (and Values) are typically represented with the same dimensionality, regardless of its distance from the prediction target. In natural language, however, the next word is most strongly influenced by the immediately preceding tokens. We hypothesize that local and distant tokens impose asymmetric demands on representational capacity: local tokens are more critical for predicting immediate outputs and thus require richer representations, whereas distant tokens primarily serve as long-range memory, for which lower-dimensional representations may suffice. We formalize this idea as Distance-Adaptive Representation (DAR), implemented in a controlled setting that preserves full-dimensional representations within a local context window while assigning reduced-dimensional representations (e.g. 1/4 of the original dimensionality) to tokens beyond that window. Across multiple pretraining scales (70M to 410M parameters), as well as continued supervised fine-tuning on a 1B-scale model, this approach closely matches the performance of full-dimensional baselines. In contrast, uniformly reducing dimensionality across all token positions leads to worse performance. These results challenge the common assumption that key and value dimensionality should be uniform across token positions. Our findings suggest a new direction for designing attention architectures that adaptively allocate representational capacity across sequences, enabling further reductions in KV cache during inference.

cs.CL↗

Ultralow shot noise limited giant passive resonant gyroscope for Earth rotation measurement

Optical gyroscopes directly measure the Earth's rotation and are promising instruments for real-time geophysical observations and Earth orientation parameter (EOP) determination requiring both high precision and high temporal resolution. Large-scale ring laser gyroscopes (RLGs) currently reach rotational resolutions around $10^{-11}\,\mathrm{(rad/s)/\sqrt{Hz}}$, but their quantum noise limits make it challenging to meet the requirements of future high-temporal-resolution EOP measurements. Passive resonant gyroscopes (PRGs), on the other hand, offer a potentially lower photon shot noise limit and more flexible power scaling, even if their demonstrated rotational resolutions are still about two orders of magnitude below those of leading RLGs. Here we demonstrate a $64\,\mathrm{m^{2}}$ giant passive resonant gyroscope HUST-2, and develop with an extremely low shot noise level. We experimentally obtain a shot noise limited of $5.7(1)\times10^{-13}\,\mathrm{(rad/s)/\sqrt{Hz}}$ at $1\,\mathrm{mW}$ incident optical power, following the characteristic $1/\sqrt{P}$ scaling. Through systematic suppression of dominant technical noise sources, HUST-2 further achieves a measured rotational resolution of $3\times10^{-11}\,\mathrm{(rad/s)/\sqrt{Hz}}$, bringing PRGs into the performance regime of leading large-scale RLGs for the first time. The gap between the present demonstrated rotational resolution and the shot noise limit indicates nearly two orders of magnitude further improvement potential. Reaching this limit would enable high-precision length-of-day (LOD) measurements with $10$-$100\,\mathrm{s}$ temporal resolution and lays the foundation for future large-scale gyroscope networks dedicated to real-time EOP determination.

physics.optics↗