SearcharxivSearch

arXiv subjects

Zhengyang Li

Publications and source records attributed to Zhengyang Li.

At least 19 recordsLinked to original sources

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or continuous autoregressive (D/CAR) SE, discrete or continuous non-autoregressive (D/CNAR) SE, discrete diffusion (DDiff) SE, and continuous flow matching (CFM) SE. Second, we are the first to compare their performance in a unified experimental setup and synopsis with diverse intrusive and non-intrusive metrics, enabling a fair and comprehensive evaluation. Third, we propose a fine-tuning strategy with auxiliary losses on reconstructed speech to improve both intrusive and non-intrusive metrics. Trained and evaluated on URGENT 2025 Speech Enhancement Challenge data splits, all continuous-domain paradigms excel their discrete-domain counterparts. The overall best approach turns out to be CNAR. We further show that our proposed auxiliary loss fine-tuning strategy helps to improve DNSMOS, NISQA, PESQ, and POLQA consistently in all six paradigms.

eess.AS

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an egocentric video clip, a question, and four candidate answers, from which the model must select the correct option. This technical report introduces the challenge task, benchmark resources, and two official Codabench tracks. The Source-Limited Track restricts participants to the official baseline model and a small support set, whereas the Open-Source Track permits broader choices of models and training data under rules that prohibit the manual construction of target-domain training data. In total, the challenge received more than 1,500 submissions from over 130 participants, with 19 teams participating in the Open-Source Track and 38 teams in the Source-Limited Track. We further present the official leaderboard results and summarize the winning solutions from both tracks. We hope that this report will serve as a useful technical reference for advancing cross-domain egocentric video understanding. All resources, including the challenge data, baseline implementation, and code released by the winning teams, are made publicly available.

cs.CV

Ising superconductivity and anomalous metallic states in a bulk crystal with artificial unidirectional stacking layers

The two-dimensional (2D) limit in macroscopic bulk crystals provides a powerful platform for exploring exotic quantum phases. Here, we report the synthesis of a Sr0.75ClNbS2 superconductor that achieves unidirectional, parallel AA stacking-a configuration never before realized in a bulk crystal. Unlike conventional intercalation, which merely expands the interlayer spacing, our approach employs a planar Sr-Cl network to enforce a complete stacking reorganization, driving all NbS2 layers from the native antiparallel AB stacking into a unidirectional, parallel AA arrangement. This stacking switch globally breaks inversion symmetry, transforming centrosymmetric 2H-NbS2 into a noncentrosymmetric bulk crystal with D3h point group symmetry. Crucially, this structural design reproduces, in three dimensions, the electronic environment of an isolated monolayer, thereby preventing cancellation of the local Ising fields. As a result, strong Ising spin-orbit coupling and spin-split bands persist throughout the bulk. Transport measurements reveal extreme superconducting anisotropy ({\gamma} ~ 77), an in-plane upper critical field (~ 10.65 T) that far exceeds the Pauli paramagnetic limit, and clean-limit superconductivity indicative of high crystalline quality. Moreover, magnetotransport uncovers a novel magnetic-field-induced anomalous metallic state characterized by finite dissipation yet a vanishing Hall response. Direct band-structure measurements corroborate the layer-decoupled, quasi-2D electronic nature of the system. This work establishes stacking-geometry engineering as a powerful strategy to artificially enforce a globally noncentrosymmetric, quasi-2D superconducting state in bulk crystals, paving the way for designing quantum materials with tunable crystalline symmetry and electronic band topology.

cond-mat.supr-con

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced Large Language Model (LLM) reasoning; however, it faces a fundamental optimization instability: uniform token updates precipitate entropy collapse, leading to premature convergence to suboptimal strategies, whereas excessive Shannon Entropy maximization can cause entropy explosion, driving blind exploration toward incoherent reasoning chains. To resolve this dichotomy, we introduce the Independent Combinatorial Tokens (ICT) framework, which shifts the optimization focus from scalar uncertainty to the distributional properties of token logits. By leveraging the Jensen-Shannon (JS) divergence between token logits distributions, ICT identifies tokens with distinctive distributional patterns as critical branching points for guiding effective exploration in LLM reasoning. Our theoretical analysis, grounded in both Shannon and second-order R\'enyi entropy, proves that selectively updating on these tokens regulates policy concentration: it reduces the overall distribution uncertainty measured by Shannon entropy, while controlling probability concentration captured by second-order R\'enyi entropy. This dual effect prevents over-concentrated token generation from weakening exploration and effectively stabilizes the training landscape. Empirical results demonstrate that updating only the top 10% of unique tokens on Qwen2.5 (0.5B/1.5B/7B) models yields an average pass@4 improvement of 4.58%, with a maximum gain of 14.9%, over GRPO, 20-Entropy, and STAPO baselines across seven benchmarks spanning math, commonsense, and Olympiad-level problems.

cs.AI

CASTLE2026 Team WDL Technical Report

The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ hours of multi-perspective recordings. Each four-choice question requires evidence from videos, transcripts, auxiliary photos, people, days, rooms, and temporal context. We propose an evidence-aware multimodal reasoning pipeline based on Qwen. Our system parses question hints, retrieves ASR chunks, attaches auxiliary images, samples candidate video frames, and routes questions into static visual, speech/text, temporal, and mixed types with specialized prompts. Multiple inference passes are aggregated by confidence-weighted voting and converted into the official Codabench format. In ablation, LoRA improves the score from 0.21 to 0.50, and more sampled frames further raise it to 0.58. Our final system ranks first in the CASTLE Challenge @ EgoVis 2026.

cs.CV

Breaking the Trade-off: Bulk 2D Ising Superconductivity with High Tc and Giant Interlayer Spacing via a Unique Chain Intercalation in (BaS)1/3TaS2

Two-dimensional (2D) transition metal dichalcogenides (TMDs) are promising platforms for low dimensional superconductivity. However, in conventional intercalated systems, achieving a high superconducting transition temperature (Tc) often comes at the expense of reduced interlayer spacing and weakened 2D character. Here, we overcome this long-standing compromise through a unique chain-like intercalation strategy. We report the synthesis and properties of a new polymorph, (BaS)1/3TaS2, in which a distinctive Ba-S-S-Ba chain structure is inserted between TaS2 bilayers. This unique configuration breaks the bulk c axis mirror symmetry while achieving exceptional interlayer decoupling, with an inter-bilayer spacing of 12.75 {\AA}-more than three times that of pristine 2H-TaS2. By suppressing interlayer electronic coupling, this structural evolution allows local inversion symmetry breaking within individual TaS2 layers to dominate. This prevents compensation of the Ising spin-orbit fields typical of centrosymmetric bulk phases, enabling robust 2D Ising superconductivity. Remarkably, the compound exhibits an enhanced Tc without sacrificing its large interlayer spacing, thereby breaking the conventional trade-off between large spacing/high anisotropy and high Tc. Comprehensive transport, magnetic, and thermodynamic measurements confirm its robust superconducting state. Our work establishes a versatile intercalation framework for designing bulk-like 2D Ising superconductors, providing a new route to reconcile competing material demands and expanding the scope of Ising superconductivity research.

cond-mat.supr-con

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The challenge is motivated by a central gap in current evaluation practice: perceptual quality alone is insufficient to judge whether generated dynamics are physically plausible, temporally coherent, and consistent with input conditions. Participants are required to build a metric that jointly predicts four dimensions, i.e., Video Quality, Physical Realism, Condition-Video Alignment, and Temporal Consistency. Depart from that, participants also need to localize physical anomaly timestamps for fine-grained diagnosis. The benchmark dataset contains 1,554 videos generated by seven representative world generative models, organized into three tracks (text-2D, image-to-4D, and video-to-4D) and spanning 26 categories. These categories explicitly cover physics-relevant scenarios, including dynamics, optics, and thermodynamics, together with diverse real-world and creative content. To ensure label reliability, scores and anomaly timestamps are produced through trained human annotation with an additional automated quality-control pass. Evaluation is based on both score prediction and anomaly localization, with a composite protocol that combines TimeStamp_IOU and SRCC/PLCC. This report summarizes the challenge design and provides method-level insights from submitted solutions.

cs.CV

Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder

In audiovisual automatic speech recognition (AV-ASR) systems, information fusion of visual features in a pre-trained ASR has been proven as a promising method to improve noise robustness. In this work, based on the prominent Whisper ASR, first, we propose a simple and effective visual fusion method -- use of visual features both in encoder and decoder (dual-use) -- to learn the audiovisual interactions in the encoder and to weigh modalities in the decoder. Second, we compare visual fusion methods in Whisper models of various sizes. Our proposed dual-use method shows consistent noise robustness improvement, e.g., a 35% relative improvement (WER: 4.41% vs. 6.83%) based on Whisper small, and a 57% relative improvement (WER: 4.07% vs. 9.53%) based on Whisper medium, compared to typical reference middle fusion in babble noise with a signal-to-noise ratio (SNR) of 0dB. Third, we conduct ablation studies examining the impact of various module designs and fusion options. Fine-tuned on 1929 hours of audiovisual data, our dual-use method using Whisper medium achieves 4.08% (MUSAN babble noise) and 4.43% (NoiseX babble noise) average WER across various SNRs, thereby establishing a new state-of-the-art in noisy conditions on the LRS3 AV-ASR benchmark. Our code is at https://github.com/ifnspaml/Dual-Use-AVASR

eess.AS

OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data

Recent successful video generation systems that predict and create realistic automotive driving scenes from short video inputs assign tokenization, future state prediction (world model), and video decoding to dedicated models. These approaches often utilize large models that require significant training resources, offer limited insight into design choices, and lack publicly available code and datasets. In this work, we address these deficiencies and present OpenViGA, an open video generation system for automotive driving scenes. Our contributions are: Unlike several earlier works for video generation, such as GAIA-1, we provide a deep analysis of the three components of our system by separate quantitative and qualitative evaluation: Image tokenizer, world model, video decoder. Second, we purely build upon powerful pre-trained open source models from various domains, which we fine-tune by publicly available automotive data (BDD100K) on GPU hardware at academic scale. Third, we build a coherent video generation system by streamlining interfaces of our components. Fourth, due to public availability of the underlying models and data, we allow full reproducibility. Finally, we also publish our code and models on Github. For an image size of 256x256 at 4 fps we are able to predict realistic driving scene videos frame-by-frame with only one frame of algorithmic latency.

cs.CV

Integrated diffractive full-Stokes spectro-polarimetric imaging

Spectro-polarimetric imaging provides multidimensional optical information acquisition capabilities, offering significant potential for diverse applications. Current spectro-polarimetric imaging systems typically suffer from large physical footprints, high design complexity, elevated costs, or the drawback of requiring replacement of standard components with polarization optics. To address these issues, we propose an integrated diffractive full-Stokes spectro-polarimetric imaging framework that synergistically combines end-to-end designed diffractive polarization spectral element (DPSE) with SPMSA-Net to demonstrate high-performance spectro-polarimetric imaging. The DPSE modulates scene and generates modulated images carrying phase-encoding and polarization information. The modulated images are the input of the SPMSA-NET for the reconstruction of the spectro-polarimetric data cube. The framework achieves an average improvement of 0.78 dB in PSNR and 0.012 in SSIM over existing state-of-the-art algorithms. Based on this framework, our prototype system can simultaneously capture spectral information (400-700 nm) with 10 nm spectral resolution and full-Stokes parameters (S0,S1,S2,S3). Meanwhile, the system provides high spatial resolution of 2252*2252 pixels. Experimental results demonstrate that our system achieves high-fidelity spectral imaging (over 98.9% fidelity) and precise polarization characterization, with a compact architecture (modulation component of merely 2-mm thickness).

eess.IV

A Comprehensive Review of Multi-Agent Reinforcement Learning in Video Games

Recent advancements in multi-agent reinforcement learning (MARL) have demonstrated its application potential in modern games. Beginning with foundational work and progressing to landmark achievements such as AlphaStar in StarCraft II and OpenAI Five in Dota 2, MARL has proven capable of achieving superhuman performance across diverse game environments through techniques like self-play, supervised learning, and deep reinforcement learning. With its growing impact, a comprehensive review has become increasingly important in this field. This paper aims to provide a thorough examination of MARL's application from turn-based two-agent games to real-time multi-agent video games including popular genres such as Sports games, First-Person Shooter (FPS) games, Real-Time Strategy (RTS) games and Multiplayer Online Battle Arena (MOBA) games. We further analyze critical challenges posed by MARL in video games, including nonstationary, partial observability, sparse rewards, team coordination, and scalability, and highlight successful implementations in games like Rocket League, Minecraft, Quake III Arena, StarCraft II, Dota 2, Honor of Kings, etc. This paper offers insights into MARL in video game AI systems, proposes a novel method to estimate game complexity, and suggests future research directions to advance MARL and its applications in game development, inspiring further innovation in this rapidly evolving field.

cs.LG

Multi-origin driven giant planar Hall effect in topological antiferromagnet EuAl2Si2 with tunable spin texture

In topological materials, the planar Hall effect (PHE) is often regarded as a hallmark of profound quantum phenomena-most notably the Adler-Bell-Jackiw chiral anomaly and Berry curvature-rendering it an indispensable tool for deciphering the topological essence of emergent phases. In this study, we delve into the PHE and anisotropic magnetoresistance in the recently discovered layered topological antiferromagnet EuAl2Si2. Our analysis of the robust PHE signal (~3.8 {\mu}{\Omega} cm at 2 K and 8 T) unveils a distinct interplay of mechanisms. While Berry curvature plays a minor role, the dominant contributions stem from classical orbital MR in the field-induced ferromagnetic state and field-suppressed spin fluctuations in the paramagnetic regime. These insights not only position EuAl2Si2-with its highly tunable spin texture-as an exemplary system for probing the intricate coupling between spin configurations and band topology in magnetotransport but also pave the way for designing novel materials with tailored PHE responses, highlighting significant application prospects in quantum sensing, spintronic devices, and topologically protected electronic systems.

cond-mat.str-el

Enhancing Text Comprehension for Dyslexic Readers: A 3D Semantic Visualization Approach Using Transformer Mode

Dyslexic individuals often face significant challenges with traditional reading, particularly when engaging with complex texts such as mystery novels. These texts typically demand advanced narrative tracking and information integration skills, making it difficult for dyslexic readers to fully comprehend the content. However, research indicates that while dyslexic individuals may struggle with textual processing, they often possess strong spatial imagination abilities. Leveraging this strength, this study proposes an innovative approach using Transformer models to map sentences and words into three-dimensional vector representations. This process clusters semantically similar sentences and words in spatial proximity, allowing dyslexic readers to interpret the semantic structure and narrative flow of the text through spatial perception. Experimental results demonstrate that, compared to direct text reading, this three-dimensional semantic visualization method significantly enhances dyslexic readers' comprehension of complex texts. In particular, it shows marked advantages in identifying narrative relationships and character connections. This study provides a novel pathway for improving textual comprehension among dyslexic individuals

cs.HC

Cognitive Load-Driven VR Memory Palaces: Personalizing Focus and Recall Enhancement

Cognitive load, which varies across individuals, can significantly affect focus and memory performance.This study explores the integration of Virtual Reality (VR) with memory palace techniques, aiming to optimize VR environments tailored to individual cognitive load levels to improve focus and memory. We utilized EEG devices, specifically the Oculus Quest 2, to monitor Beta wave activity in 10 participants.By modeling their cognitive load profiles through polynomial regression, we dynamically adjusted spatial variables within a VR environment using Grasshopper, creating personalized experiences. Results indicate that 8 participants showed a notable increase in Beta wave activity, demonstrating improved focus and cognitive performance in the customized VR settings.These findings underscore the potential of VR-based memory environments, driven by cognitive load considerations, and provide valuable insights for advancing VR memory research

cs.HC

Language-Driven Coordination and Learning in Multi-Agent Simulation Environments

This paper introduces LLM-MARL, a unified framework that incorporates large language models (LLMs) into multi-agent reinforcement learning (MARL) to enhance coordination, communication, and generalization in simulated game environments. The framework features three modular components of Coordinator, Communicator, and Memory, which dynamically generate subgoals, facilitate symbolic inter-agent messaging, and support episodic recall. Training combines PPO with a language-conditioned loss and LLM query gating. LLM-MARL is evaluated in Google Research Football, MAgent Battle, and StarCraft II. Results show consistent improvements over MAPPO and QMIX in win rate, coordination score, and zero-shot generalization. Ablation studies demonstrate that subgoal generation and language-based messaging each contribute significantly to performance gains. Qualitative analysis reveals emergent behaviors such as role specialization and communication-driven tactics. By bridging language modeling and policy learning, this work contributes to the design of intelligent, cooperative agents in interactive simulations. It offers a path forward for leveraging LLMs in multi-agent systems used for training, games, and human-AI collaboration.

cs.AI

An Optimization Framework for Wide-Field Small Aperture Telescope Arrays Used in Sky Surveys

For time-domain astronomy, it is crucial to frequently image celestial objects at specific depths within a predetermined cadence. To fulfill these scientific demands, scientists globally have started or planned the development of non-interferometric telescope arrays in recent years. Due to the numerous parameters involved in configuring these arrays, there is a need for an automated optimization framework that selects parameter sets to satisfy scientific needs while minimizing costs. In this paper, we introduce such a framework, which integrates optical design software, an exposure time calculator, and an optimization algorithm, to balance the observation capabilities and the cost of optical telescope arrays. Neural networks are utilized to speed up results retrieval of the system with different configurations. We use the SiTian project as a case study to demonstrate the framework's effectiveness, showing that this approach can aid scientists in selecting optimal parameter sets. The code for this framework is published in the China Virtual Observatory PaperData Repository, enabling users to optimize parameters for various non-interferometric telescope array projects.

astro-ph.IM

A compact $T1$ theorem for Calderón-Zygmund operators associated with Zygmund dilations

We develop a compact version of $T1$ theorem for singular integrals of Zygmund type on $\mathbb{R}^3$. More specifically, if a $(D_θ, δ_1, δ_{2, 3})$-Calderón-Zygmund operator $T$ associated with Zygmund dilations admits the compact full and partial kernel representations, and satisfies the weak compactness property and the cancellation condition, then $T$ can be extended to a compact operator on $L^p(w)$ whenever (i) $p \in (1, \infty)$, $w \in A_{p, \mathcal{R}}$, and $θ, δ_1, δ_{2, 3} \in (0, 1]$, or (ii) $p \in (1, \infty)$, $w \in A_{p, \mathcal{Z}}$, $θ= δ_1 = 1$, and $δ_{2, 3} \in (0, 1]$. Here $A_{p, \mathcal{R}}$ and $A_{p, \mathcal{Z}}$ respectively denote the class of of strong $A_p$ weights and the class of Zygmund $A_p$ weights. Beyond that, under similar bilinear assumptions, we prove bilinear Calderón-Zygmund operators associated with Zygmund dilations are compact from $L^{p_1}(\mathbb{R}^3) \times L^{p_2}(\mathbb{R}^3)$ to $L^p(\mathbb{R}^3)$ for all $p_1, p_2 \in (1, \infty)$, where $\frac1p = \frac{1}{p_1} + \frac{1}{p_2}$. The core of the proof is a compact dyadic representation, which asserts that under the hypotheses above, a (bilinear) Calderón-Zygmund operator associated with Zygmund dilations can be represented an average of some compact (bilinear) dyadic shifts of Zygmund nature. This further deepens our understanding of the compactness of singular integral operators.

math.CA