SearcharxivSearch

arXiv subjects

Zhenyu Guo

Publications and source records attributed to Zhenyu Guo.

At least 19 recordsLinked to original sources

HyQuant: Hybrid-Precision Quantization for LLM Attention

Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .

cs.AI

Multi-quantum-channel mediated tunable single-photon skyrmions from metasurfaces

Quantum optical skyrmions, as topologically robust quantum information carriers, hold transformative potential for resilient high-dimensional quantum information networks. However, their practical exploitation was still restricted to a single quantum channel, which precludes the multiplexing essential for practical high-capacity quantum networks. Here, we utilize a metasurface to achieve multi-channel quantum state distribution of the polarization-entangled photon pairs, inducing a two-photon bunching effect in both the spin and spatial dimensions with compact flat optics. At the spatial bunching port, controlled manipulation of the spin-orbit interaction enables the generation of a tunable single-photon skyrmion pair. In contrast to any prior skyrmion generation, the single-photon skyrmions are mediated and topologically controlled by quantum measurement in multiple channels. Concurrently, during the amplitude and phase modulation process, both the skyrmion localization and the texture helicity can be precisely customized. The proposed tunable single-photon skyromions offer multidimensional controllability and topological stability provide a viable path toward noise-resilient high-dimensional quantum information processing.

physics.optics

Alignment-Free Nanometric Optical Metrology Enabled by Structured Light

Advances in the semiconductor industry are driven by the development of increasingly compact devices featuring intricate etched geometries, the characterization of which essentially requires ultraprecise, label-free, and real-time metrology. However, non-destructive and alignment-free optical metrology of sub-wavelength structures with nanometric resolution remains a major challenge. Here, we demonstrate a novel single-shot, label-free, and alignment-free optical metrology approach for determining the 1D position of sub-wavelength nanostructures, achieving lambda/110 (7.2 nm) precision. The high precision benefits from utilizing structured illuminations of Laguerre-Gaussian (LG) or Hermite-Gaussian (HG) beams, and the AI analyzing method can retrieve the information when such structured light interacts with sub-wavelength objects. Instead of relying on phase singularities in superoscillatory microscopy, our approach leverages spatially distributed phase jumps in HG and LG beams interacting with the nanostructures, providing an alignment-robust solution to the challenges in optical metrology. Such an alignment-free, non-destructive, and high-precision metrology technique enables real-time machine vision, semiconductor inspection, and advanced manufacturing.

physics.optics

High-speed electrically driven liquid-crystal compact optical skyrmion encoder

Optical skyrmions possess topological polarization textures that can maintain topological robustness under external perturbations, making them promising carriers for disturbance-resistant optical information transmission. However, existing optical skyrmion generation schemes mostly rely on static optical elements or fixed nanostructures, making high-speed dynamic switching of the topological state difficult. Here, we propose a high-speed switchable optical skyrmion generator based on a patterned liquid-crystal spin-orbit device. The device employs the in-plane orientation of liquid crystals to imprint a fixed Pancharatnam-Berry geometric phase, while an applied voltage rapidly tunes the liquid-crystal retardance, enabling reversible switching between skyrmion and non-skyrmion states. Experimental results show that the device exhibits millisecond electrical response, with bidirectional response times of 1.76 ms and 0.72 ms, corresponding to an ideal cycling rate of approximately 403 Hz, making it the fastest switchable optical skyrmion generator to date. Furthermore, by exploiting this rapid topological refreshing capability, we demonstrate image encoding and decoding, providing a new liquid-crystal device platform for high-speed, refreshable, and disturbance-resistant topological optical information transmission.

physics.optics

Can non-orthogonal bases form stable skyrmionic beams?

Skyrmions, topologically stable spin textures, have recently garnered significant attention in optics promising robust high-density information transition and nontrivial light-matter interaction. It was believed that the optical skyrmionic beams should be constructed by superposition of two orthogonal spatial modes with orthogonal polarizations to obtain topologically stable propagation. Here, we surprisingly find that propagation-stable skyrmionic beams can still be formed by superpositions of neither orthogonal spatial modes nor orthogonal polarizations. We theoretically present the mechanism to control the stable skyrmionics beams through the hybrid superposition of modes from the Hermite-Gaussian and Laguerre-Gaussian families and experimentally control the longitudinal on-demand dynamics of the skyrmions. This work redefines the topological stability of optical skyrmions, breaks limits and reduces the requirement for manipulating topologically structured light for practical multidimensional implementation of topologically robust information technologies.

physics.optics

Optical hopfions with arbitrary two winding numbers

Hopfions, as three-dimensional topologically nontrivial structures described by poloidal and toroidal winding numbers, hold promise as robust information carriers in spintronics, functional materials, and optical communications. Although they have been experimentally realized in various physical systems, such realizations have been restricted to low orders, with the winding numbers lacking tunability. Here, using optical fields as our platform, we outline how to make tunable hopfions in any order with any winding number. We use tailored superpositions of Laguerre-Gaussian modes in free-space as our construction, achieving effective control for arbitrary-order poloidal and toroidal winding numbers, which we demonstrate up to orders 5 and 3, respectively, for a new state-of-the-art. The resulting torus-knot structures are visualized experimentally via polarization filaments, confirming the designed topological textures. Our work reports an exotic optical topologies observed in free space, provides a systematic route hopfions of any order, with implications for topological photonics, optical communications, and analogies in magnetic and condensed-matter systems.

physics.optics

Partial coherence control delivers skyrmionic topological resilience and transitions

Optical skyrmions have recently unlocked topological quasiparticle textures of light, rising in prominence for next-generation ultra-robust information processing. However, to date, their study has been mainly confined to coherent laser fields. Here we extend skyrmions to more general light sources of partially coherent, stochastic optical fields. We define stochastic optical skyrmions and uncover a hidden regime where spatial coherence acts as a primary determinant of topological stability. While environmental randomness typically degrades fully coherent states, we demonstrate that engineered partial coherence provides a self-healing mechanism that preserves topology under extreme turbulence. Moreover, we show that the coherence structure can be actively tailored to trigger on-demand topological phase transitions, such as skyrmion-to-skyrmionium conversion and skyrmion lattice splitting. These findings redefine the boundaries of topological photonics, paving the way for resilient and high-fidelity information platforms that remain operational in general, non-ideal, real-world environments.

physics.optics

Unfolding unstable skyrmionic polarization textures

Polarization of light can form skyrmionic textures, akin to nonlinear solitons in condensed matter, yet their disparate physical context has motivated extensive debate regarding their stability. Here we show that the topological charge of such structures (skyrmion number) changes when an arbitrarily small perturbation splits coalescent phase singularities. In a superposition of two vortex beams, the skyrmion number generally only depends on the higher order topological charge $\lrr{Q_{\rm sk}=\max\lr{\ell_2,\ell_1}}$ rather than the difference of charges of the vortices in superposition $\lrr{Q_{\rm sk}=\ell_2-\ell_1}$, which only holds in the absence of perturbation. These results have significant implications for polarization structures with wavelength-scale localization and those experiencing complex aberrations.

physics.optics

Topological robustness of orbital angular momentum entanglement in stochastic channels

Orbital angular momentum (OAM) entanglement gives access to multiple qubit and high dimensional Hilbert spaces, but is unfortunately susceptible to disturbance, decaying in real-world noisy channels. Here, we show there is an underlying topology arising from OAM entanglement that is robust to such channels, which we demonstrate using atmospheric turbulence -- exemplary of stochastic or chaotic media. Using a quantum channel with various turbulence strengths, we find the OAM topological observable preserved even though the OAM itself is shown to be highly sensitive to the turbulence. We show this is true for mixed states too, with the OAM topology intact even as the purity of the state decreases due to decoherence. Our work offers a new perspective on OAM entanglement preservation, and may easily be extended to other spatial bases, degrees of freedom, as well as complex channels, whether static or dynamic.

quant-ph

UI-Venus-1.5 Technical Report

GUI agents have emerged as a powerful paradigm for automating interactions in digital environments, yet achieving both broad generality and consistently strong task performance remains challenging. In this report, we present UI-Venus-1.5, a unified, end-to-end GUI Agent designed for robust real-world applications. The proposed model family comprises two dense variants (2B and 8B) and one mixture-of-experts variant (30B-A3B) to meet various downstream application scenarios. Compared to our previous version, UI-Venus-1.5 introduces three key technical advances: (1) a comprehensive Mid-Training stage leveraging 10 billion tokens across 30+ datasets to establish foundational GUI semantics; (2) Online Reinforcement Learning with full-trajectory rollouts, aligning training objectives with long-horizon, dynamic navigation in large-scale environments; and (3) a single unified GUI Agent constructed via Model Merging, which synthesizes domain-specific models (grounding, web, and mobile) into one cohesive checkpoint. Extensive evaluations demonstrate that UI-Venus-1.5 establishes new state-of-the-art performance on benchmarks such as ScreenSpot-Pro (69.6%), VenusBench-GD (75.0%), and AndroidWorld (77.6%), significantly outperforming previous strong baselines. In addition, UI-Venus-1.5 demonstrates robust navigation capabilities across a variety of Chinese mobile apps, effectively executing user instructions in real-world scenarios. Code: https://github.com/inclusionAI/UI-Venus; Model: https://huggingface.co/collections/inclusionAI/ui-venus

cs.CV

GRAFT: Grid-Aware Load Forecasting with Multi-Source Textual Alignment and Fusion

Electric load is simultaneously affected across multiple time scales by exogenous factors such as weather and calendar rhythms, sudden events, and policies. Therefore, this paper proposes GRAFT (GRid-Aware Forecasting with Text), which modifies and improves STanHOP to better support grid-aware forecasting and multi-source textual interventions. Specifically, GRAFT strictly aligns daily-aggregated news, social media, and policy texts with half-hour load, and realizes text-guided fusion to specific time positions via cross-attention during both training and rolling forecasting. In addition, GRAFT provides a plug-and-play external-memory interface to accommodate different information sources in real-world deployment. We construct and release a unified aligned benchmark covering 2019--2021 for five Australian states (half-hour load, daily-aligned weather/calendar variables, and three categories of external texts), and conduct systematic, reproducible evaluations at three scales -- hourly, daily, and monthly -- under a unified protocol for comparison across regions, external sources, and time scales. Experimental results show that GRAFT significantly outperforms strong baselines and reaches or surpasses the state of the art across multiple regions and forecasting horizons. Moreover, the model is robust in event-driven scenarios and enables temporal localization and source-level interpretation of text-to-load effects through attention read-out. We release the benchmark, preprocessing scripts, and forecasting results to facilitate standardized empirical evaluation and reproducibility in power grid load forecasting.

cs.LG

Topological robustness of classical and quantum optical skyrmions in atmospheric turbulence

The degradation of classical and quantum structured light induced by complex media constitutes a critical barrier to its practical implementation in a range of applications, from communication and energy transport to imaging and sensing. Atmospheric turbulence is an exemplary case due to its complex phase structure and dynamic variations, driving the need to find invariances in light. Here we construct classical and quantum optical skyrmions and pass them through experimentally simulated atmospheric turbulence, revealing the embedded topological resilience of their structure. In the quantum realm, we show that while skyrmions undergo diminished entanglement, their topological characteristics maintain stable. This is paralleled classically, where the vectorial structure is scrambled by the medium yet the skyrmion remains stable by virtue of its intrinsic topological protection mechanism. Our experimental results are supported by rigorous analytical and numerical modelling, validating that the quantum-classical equivalence of the topological behaviour is due to the non-separability of the states and one-sided nature of the channel. Our work blurs the classical-quantum divide in the context of topology and opens a new path to information resilience in noisy channels, such as terrestrial and satellite-to-ground communication networks.

physics.optics

UI-Venus Technical Report: Building High-performance UI Agents with RFT

We present UI-Venus, a native UI agent that takes only screenshots as input based on a multimodal large language model. UI-Venus achieves SOTA performance on both UI grounding and navigation tasks using only several hundred thousand high-quality training samples through reinforcement finetune (RFT) based on Qwen2.5-VL. Specifically, the 7B and 72B variants of UI-Venus obtain 94.1% / 50.8% and 95.3% / 61.9% on the standard grounding benchmarks, i.e., Screenspot-V2 / Pro, surpassing the previous SOTA baselines including open-source GTA1 and closed-source UI-TARS-1.5. To show UI-Venus's summary and planing ability, we also evaluate it on the AndroidWorld, an online UI navigation arena, on which our 7B and 72B variants achieve 49.1% and 65.9% success rate, also beating existing models. To achieve this, we introduce carefully designed reward functions for both UI grounding and navigation tasks and corresponding efficient data cleaning strategies. To further boost navigation performance, we propose Self-Evolving Trajectory History Alignment & Sparse Action Enhancement that refine historical reasoning traces and balances the distribution of sparse but critical actions, leading to more coherent planning and better generalization in complex UI tasks. Our contributions include the publish of SOTA open-source UI agents, comprehensive data cleaning protocols and a novel self-evolving framework for improving navigation performance, which encourage further research and development in the community. Code is available at https://github.com/inclusionAI/UI-Venus.

cs.CV

Causal Mean Field Multi-Agent Reinforcement Learning

Scalability remains a challenge in multi-agent reinforcement learning and is currently under active research. A framework named mean-field reinforcement learning (MFRL) could alleviate the scalability problem by employing the Mean Field Theory to turn a many-agent problem into a two-agent problem. However, this framework lacks the ability to identify essential interactions under nonstationary environments. Causality contains relatively invariant mechanisms behind interactions, though environments are nonstationary. Therefore, we propose an algorithm called causal mean-field Q-learning (CMFQ) to address the scalability problem. CMFQ is ever more robust toward the change of the number of agents though inheriting the compressed representation of MFRL's action-state space. Firstly, we model the causality behind the decision-making process of MFRL into a structural causal model (SCM). Then the essential degree of each interaction is quantified via intervening on the SCM. Furthermore, we design the causality-aware compact representation for behavioral information of agents as the weighted sum of all behavioral information according to their causal effects. We test CMFQ in a mixed cooperative-competitive game and a cooperative game. The result shows that our method has excellent scalability performance in both training in environments containing a large number of agents and testing in environments containing much more agents.

cs.AI

Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention

Transformers have achieved remarkable success across diverse domains, but their monolithic architecture presents challenges in interpretability, adaptability, and scalability. This paper introduces a novel modular Transformer architecture that explicitly decouples knowledge and reasoning through a generalized cross-attention mechanism to a globally shared knowledge base with layer-specific transformations, specifically designed for effective knowledge retrieval. Critically, we provide a rigorous mathematical derivation demonstrating that the Feed-Forward Network (FFN) in a standard Transformer is a specialized case (a closure) of this generalized cross-attention, revealing its role in implicit knowledge retrieval and validating our design. This theoretical framework provides a new lens for understanding FFNs and lays the foundation for future research exploring enhanced interpretability, adaptability, and scalability, enabling richer interplay with external knowledge bases and other systems.

cs.LG

Witnessing Quantum Incompatibility Structures in High-Dimensional Multimeasurement Systems

Quantum incompatibility, referred as the phenomenon that some quantum measurements cannot be performed simultaneously, is necessary for various quantum information processing tasks, such as nonlocality and steering. When these applications come to high-dimensional multimeasurement scenarios, it is crucial and challenging to witness the incompatibility of measurements with complex structures. To address this problem, we propose a modified quantum state discrimination protocol that decomposes complex compatibility structures into pairwise ones and employs noise robustness to bound incompatibility structures. We then derive arithmetic bounds for arbitrary measurements and analytical bounds for mutually unbiased bases, and capture some quantum incompatibility structures where measurements are partly compatible and partly incompatible. Finally, we experimentally demonstrate our results and connect them with quantum steering, quantum simulability and quantum communications.

quant-ph

High-areal-capacity Na-ion battery electrode with uncompromised energy and power densities by simultaneous electrospinning-spraying fabrication

Sodium-ion batteries (SIBs) are cost-effective alternatives to lithium-ion batteries (LIBs), but their low energy density remains a challenge. Current electrode designs fail to simultaneously achieve high areal loading, high active content, and superior performance. In response, this work introduces an ideal electrode structure, featuring a continuous conductive network with active particles securely trapped in the absence of binder, fabricated using a universal technique that combines electrospinning and electrospraying (co-ESP). We found that the particle size must be larger than the network's pores for optimised performance, an aspect overlooked in previous research. The free-standing co-ESP Na2V3(PO4)3 (NVP) cathodes demonstrated state-of-the-art 296 mg cm-2 areal loading with 97.5 wt.% active content, as well as remarkable rate-performance and cycling stability. Co-ESP full cells showed uncompromised energy and power densities (231.6 Wh kg-1 and 7152.6 W kg-1), leading among reported SIBs with industry-relevant areal loadings. The structural merit is analysed using multi-scale X-ray computed tomography, providing valuable design insights.Finally, the superior performance is validated in the pouch cells, highlighting the electrode's scalability and potential for commercial application.

cond-mat.mtrl-sci

AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production

The Agent and AIGC (Artificial Intelligence Generated Content) technologies have recently made significant progress. We propose AesopAgent, an Agent-driven Evolutionary System on Story-to-Video Production. AesopAgent is a practical application of agent technology for multimodal content generation. The system integrates multiple generative capabilities within a unified framework, so that individual users can leverage these modules easily. This innovative system would convert user story proposals into scripts, images, and audio, and then integrate these multimodal contents into videos. Additionally, the animating units (e.g., Gen-2 and Sora) could make the videos more infectious. The AesopAgent system could orchestrate task workflow for video generation, ensuring that the generated video is both rich in content and coherent. This system mainly contains two layers, i.e., the Horizontal Layer and the Utility Layer. In the Horizontal Layer, we introduce a novel RAG-based evolutionary system that optimizes the whole video generation workflow and the steps within the workflow. It continuously evolves and iteratively optimizes workflow by accumulating expert experience and professional knowledge, including optimizing the LLM prompts and utilities usage. The Utility Layer provides multiple utilities, leading to consistent image generation that is visually coherent in terms of composition, characters, and style. Meanwhile, it provides audio and special effects, integrating them into expressive and logically arranged videos. Overall, our AesopAgent achieves state-of-the-art performance compared with many previous works in visual storytelling. Our AesopAgent is designed for convenient service for individual users, which is available on the following page: https://aesopai.github.io/.

cs.CV