SearcharxivSearch

arXiv subjects

Biao Wang

Publications and source records attributed to Biao Wang.

At least 19 recordsLinked to original sources

Simple critical zeros and distinct zeros of the Riemann zeta-function in short intervals

Recently, on the non-trivial zeros of the Riemann zeta function, it is discovered by Claude and verified by Alp\"oge and Furman that more than 67.25% of the zeros are simple and on the critical line, and more than 83.62% are distinct. Later, Lamzouri gave a different and more direct proof. In this article, we will use the method of Lamzouri to give lower bounds on the number of the non-trivial zeros of the Riemann zeta function in short intervals. To prove the main result, we establish Montgomery's theorem on the pair correlation of zeros of the zeta function in short intervals by following the approach of Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, and then use Lamzouri's inequality on any finite multiset of complex numbers which is invariant under complex conjugation.

math.NT

Two averaged dynamical generalizations of Chowla's conjecture

Let $k\ge1$ be an integer and let $\lambda$ be the Liouville function. In 1965, Chowla gave a conjecture that the values of $\lambda(n+h_1),\dots, \lambda(n+h_k)$ are asymptotically unrelated for any distinct natural numbers $h_1, \dots, h_k$. In this article, motivated by the recent work of Bergelson and Richter on the dynamical generalizations of the prime number theorem, we will show a dynamical generalization of Chowla's conjecture on average. In the proof, we follow an approach of Qi and Zheng who established a variant of Bergelson and Richter's theorem over irreducible binary cubic forms. Moreover, we will use this approach to show an analogue of the dynamical Chowla's conjecture along the primes on average. In 2016, Tao proved that the two-point logarithmic Chowla's conjecture holds. Recently, Charamaras and Richter generalized Tao's theorem to bounded arithmetic functions and proposed a conjecture that generalizes Chowla's conjecture to bounded multi-variable arithmetic functions. Building on their work, we prove a dynamical generalization of Tao's theorem and a variant for the composition of the sum-of-digits function with the prime Omega function.

math.NT

Asymptotic uncorrelations between functions with squarefull kernel and functions of invariant average

In 1986, Ivi\'c and Tenenbaum introduced arithmetic functions with squarefull kernel, which are also called $s$-functions. Later, Erd\H{o}s and Ivi\'c gave an asymptotic estimate on the shifted convolution sums of $s$-functions. Recently, Bergelson and Richter studied the orbits along the prime Omega function in a uniquely ergodic topological dynamical system and established a new dynamical generalization of the prime number theorem (PNT). These orbits can be viewed as functions of invariant average under multiplications. In this paper, we show that both $s$-functions and their shifted convolutions are asymptotically uncorrelated to the orbits along the prime Omega function in a uniquely ergodic system. As a consequence, we obtain a refinement of the PNT via the local distribution of $s$-functions. Furthermore, several variants of these results are established as well.

math.NT

Surface code logical operations on a superconducting quantum processor

Fault-tolerant quantum computation requires logical operations that manipulate encoded information while preserving quantum error-correction protection. In planar surface-code architectures, code deformation and lattice surgery provide a local, measurement-based route to such operations. Here we experimentally realize key elements of patch-based surface-code logical processing on a 107-qubit superconducting quantum processor. We first implement a reusable primitive layer comprising merge and split, patch expansion and shrinkage, and deformations mediated by domain walls and twist defects. We then compose these primitives to realize logical state routing, the logical controlled-NOT gate, and the single-qubit Hadamard and phase gates, which together form a Clifford-generating set. All operations are implemented on distance-three rotated surface-code patches with multi-round syndrome extraction and neural-network decoding, without post-selection. Our results advance superconducting surface-code experiments from protected logical memory to active, patch-based fault-tolerant logical operations.

quant-ph

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: AI Flash Portrait (Track 3)

In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI Flash Portrait. Despite significant advancements in deep learning for image restoration, existing models still encounter substantial challenges in real-world low-light portrait scenarios. Specifically, they struggle to achieve an optimal balance among noise suppression, detail preservation, and faithful illumination and color reproduction. To bridge this gap, this challenge aims to establish a novel benchmark for real-world low-light portrait restoration. We comprehensively evaluate the proposed algorithms utilizing a hybrid evaluation system that integrates objective quantitative metrics with rigorous subjective assessment protocols. For this competition, we provide a dataset containing 800 groups of real-captured low-light portrait data. Each group consists of a 1K-resolution low-light input image, a 1K ground truth (GT), and a 1K person mask. This challenge has garnered widespread attention from both academia and industry, attracting over 100 participating teams and receiving more than 3,000 valid submissions. This report details the motivation behind the challenge, the dataset construction process, the evaluation metrics, and the various phases of the competition. The released dataset and baseline code for this track are publicly available from the same \href{https://github.com/zsn1434/AI_Flash-BaseLine/tree/main}{GitHub repository}, and the official challenge webpage is hosted on \href{https://www.codabench.org/competitions/12885/}{CodaBench}.

cs.CV

Temperature-activated dislocation avalanches signaling brittle-to-ductile transition in BCC micropillars

We carry out strain-controlled in-situ compression experiments of micron-sized tungsten (W) micropillars in the temperature range 300-900 K, together with simulations of three-dimensional discrete dislocation dynamics (DDD) at the same scale. Two distinct regimes are observed. At low temperatures, plastic deformation appears smooth, both temporally and spatially. Stress fluctuations are consistent with a Wiener stochastic process resulting from uncorrelated dislocation activity within the pillars. However, high-temperature stress fluctuations are highly correlated and exhibit features of self-organized criticality (SOC), with deformation located within well-defined slip bands. The high-temperature stress relaxation statistics are consistent with a thermally activated nucleation process from the surface. The nature of the transition between the two regimes is a manifestation of the brittle to ductile transition in BCC metals.

cond-mat.mtrl-sci

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visual understanding tasks. In this paper, we propose OmniVideo-R1, a novel reinforced framework that improves mixed-modality reasoning. OmniVideo-R1 empowers models to "think with omnimodal cues" by two key strategies: (1) query-intensive grounding based on self-supervised learning paradigms; and (2) modality-attentive fusion built upon contrastive learning paradigms. Extensive experiments on multiple benchmarks demonstrate that OmniVideo-R1 consistently outperforms strong baselines, highlighting its effectiveness and robust generalization capabilities.

cs.AI

Mamyshev oscillator based on gain-managed nonlinearity and chirped pulse amplification

We experimentally demonstrate a Mamyshev oscillator based on gain-managed nonlinearity and chirped pulse amplification. Different from other Mamyshev oscillators, the gain-managed nonlinear regime serves as a seed provider instead of a power amplifier in one arm of this laser. The output pulse energy over 300 nJ and pulse width of 739 fs has been achieved from the chirped pulsed amplification in another arm. This configuration provided a new approach to design a high-energy ultrafast laser.

physics.optics

RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation

Most text-to-video(T2V) diffusion models depend on pre-trained text encoders for semantic alignment, yet they often fail to maintain video quality when provided with concise prompts rather than well-designed ones. The primary issue lies in their limited textual semantics understanding. Moreover, these text encoders cannot rephrase prompts online to better align with user intentions, which limits both the scalability and usability of the models, To address these challenges, we introduce RISE-T2V, which uniquely integrates the processes of prompt rephrasing and semantic feature extraction into a single and seamless step instead of two separate steps. RISE-T2V is universal and can be applied to various pre-trained LLMs and video diffusion models(VDMs), significantly enhancing their capabilities for T2V tasks. We propose an innovative module called the Rephrasing Adapter, enabling diffusion models to utilize text hidden states during the next token prediction of the LLM as a condition for video generation. By employing a Rephrasing Adapter, the video generation model can implicitly rephrase basic prompts into more comprehensive representations that better match the user's intent. Furthermore, we leverage the powerful capabilities of LLMs to enable video generation models to accomplish a broader range of T2V tasks. Extensive experiments demonstrate that RISE-T2V is a versatile framework applicable to different video diffusion model architectures, significantly enhancing the ability of T2V models to generate high-quality videos that align with user intent. Visual results are available on the webpage at https://rise-t2v.github.io.

cs.CV

VC4VG: Optimizing Video Captions for Text-to-Video Generation

Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video captions specifically for T2V training remain underexplored. In this paper, we introduce VC4VG (Video Captioning for Video Generation), a comprehensive caption optimization framework tailored to the needs of T2V models. We begin by analyzing caption content from a T2V perspective, decomposing the essential elements required for video reconstruction into multiple dimensions, and proposing a principled caption design methodology. To support evaluation, we construct VC4VG-Bench, a new benchmark featuring fine-grained, multi-dimensional, and necessity-graded metrics aligned with T2V-specific requirements. Extensive T2V fine-tuning experiments demonstrate a strong correlation between improved caption quality and video generation performance, validating the effectiveness of our approach. We release all benchmark tools and code at https://github.com/alimama-creative/VC4VG to support further research.

cs.CV

A New Type of Adversarial Examples

Most machine learning models are vulnerable to adversarial examples, which poses security concerns on these models. Adversarial examples are crafted by applying subtle but intentionally worst-case modifications to examples from the dataset, leading the model to output a different answer from the original example. In this paper, adversarial examples are formed in an exactly opposite manner, which are significantly different from the original examples but result in the same answer. We propose a novel set of algorithms to produce such adversarial examples, including the negative iterative fast gradient sign method (NI-FGSM) and the negative iterative fast gradient method (NI-FGM), along with their momentum variants: the negative momentum iterative fast gradient sign method (NMI-FGSM) and the negative momentum iterative fast gradient method (NMI-FGM). Adversarial examples constructed by these methods could be used to perform an attack on machine learning systems in certain occasions. Moreover, our results show that the adversarial examples are not merely distributed in the neighbourhood of the examples from the dataset; instead, they are distributed extensively in the sample space.

cs.LG

Some ergodic theorems over $k$-full numbers

In 2022, Bergelson and Richter established a new dynamical generalization of the prime number theorem. Later, Loyd showed a disjoint form with the Erdős-Kac theorem. Recently, the author and his coauthors proved some ergodic theorems over squarefree numbers related to these results. In this paper, building on the previous work, we will derive the analogues of Bergelson-Richter's theorem, Erdős-Kac theorem and Loyd's theorem over $k$-full numbers for any integer $k\geq2$.

math.NT

Injection and Imaging of Achiral Microswimmers in Zebrafish

Achieving stable in vivo locomotion is essential for using magnetically actuated microswimmers for biomedical applications; however, while existing microswimmers have excellent motion control in vitro, their motion is greatly hindered inside living organisms. Moreover, previous work had only visually demonstrated in vivo motion through gradient pulling or rolling, but not swimming. This study investigated the injection and imaging of the achiral planar microswimmers (APMs) inside a live zebrafish embryo. The APMs can be actuated under a rotating magnetic field to generate a forward thrust at low Reynolds number environment. Combined with a safe injection technique and clear in vivo imaging at high resolution, it would be possible to control the swimming motion of APMs inside the zebrafish embryo. This work shows the safe injection and the clear imaging of an APM in a transparent zebrafish model, demonstrating the possibility for follow-up in-depth studies of the swimming motion of microswimmers in vivo.

physics.bio-ph

A logarithmic analogue of Alladi's formula

Let $μ(n)$ be the Möbius function. Let $P^-(n)$ denote the smallest prime factor of an integer $n$. In 1977, Alladi established the following formula related to the prime number theorem for arithmetic progressions \[ -\sum_{\substack{n\geq 2\\ P^-(n)\equiv \ell ({\rm mod}k)}}\frac{μ(n)}{n}=\frac1{φ(k)} \] for positive integers $\ell, k\ge$ with $(\ell,k)=1$, where $φ$ is Euler's totient function. In this note, we will show a logarithmic analogue of Alladi's formula in an elementary proof.

math.NT

Lattice distortions and non-sluggish diffusion in BCC refractory high entropy alloys

Refractory high-entropy alloys (RHEAs) have emerged as promising candidates for extreme high-temperature applications, for example, in next-generation turbines and nuclear reactors. In such applications, atomic diffusion critically governs essential properties including creep resistance and microstructural stability. The present study systematically investigates impurity diffusion of Co, Mn, and Zn in single phase (BCC solid solution) HfTiZrNbTa and HfTiZrNbV RHEAs applying the radiotracer technique. A neutron total scattering technique is used to evaluate the pair distribution functions and element-specific lattice distortions in these alloys. \textit{Ab initio}-based calculations give access to lattice distortions and solubilities of the impurities under investigation, including the impact of short-range order. The diffusion results are discussed in relation to calculated substitutional and interstitial solution energies, local lattice distortions, and short-range order effects. Co diffusion is found to be dominated by the interstitial mechanism, exhibiting fast diffusion. These findings reveal important structure-property relationships between local atomic environments and diffusion kinetics in BCC RHEAs, providing critical insights for designing alloys with enhanced high-temperature performance through targeted control of impurity diffusion processes.

cond-mat.mtrl-sci

CoSight: Exploring Viewer Contributions to Online Video Accessibility Through Descriptive Commenting

The rapid growth of online video content has outpaced efforts to make visual information accessible to blind and low vision (BLV) audiences. While professional Audio Description (AD) remains the gold standard, it is costly and difficult to scale across the vast volume of online media. In this work, we explore a complementary approach to broaden participation in video accessibility: engaging everyday video viewers at their watching and commenting time. We introduce CoSight, a Chrome extension that augments YouTube with lightweight, in-situ nudges to support descriptive commenting. Drawing from Fogg's Behavior Model, CoSight provides visual indicators of accessibility gaps, pop-up hints for what to describe, reminders to clarify vague comments, and related captions and comments as references. In an exploratory study with 48 sighted users, CoSight helped integrate accessibility contribution into natural viewing and commenting practices, resulting in 89% of comments including grounded visual descriptions. Follow-up interviews with four BLV viewers and four professional AD writers suggest that while such comments do not match the rigor of professional AD, they can offer complementary value by conveying visual context and emotional nuance for understanding the videos.

cs.HC

R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning

Visual single object tracking aims to continuously localize and estimate the scale of a target in subsequent video frames, given only its initial state in the first frame. This task has traditionally been framed as a template matching problem, evolving through major phases including correlation filters, two-stream networks, and one-stream networks with significant progress achieved. However, these methods typically require explicit classification and regression modeling, depend on supervised training with large-scale datasets, and are limited to the single task of tracking, lacking flexibility. In recent years, multi-modal large language models (MLLMs) have advanced rapidly. Open-source models like Qwen2.5-VL, a flagship MLLMs with strong foundational capabilities, demonstrate excellent performance in grounding tasks. This has spurred interest in applying such models directly to visual tracking. However, experiments reveal that Qwen2.5-VL struggles with template matching between image pairs (i.e., tracking tasks). Inspired by deepseek-R1, we fine-tuned Qwen2.5-VL using the group relative policy optimization (GRPO) reinforcement learning method on a small-scale dataset with a rule-based reward function. The resulting model, R1-Track, achieved notable performance on the GOT-10k benchmark. R1-Track supports flexible initialization via bounding boxes or text descriptions while retaining most of the original model's general capabilities. And we further discuss potential improvements for R1-Track. This rough technical report summarizes our findings as of May 2025.

cs.CV

Generation of 95-qubit genuine entanglement and verification of symmetry-protected topological phases

Symmetry-protected topological (SPT) phases are fundamental features of cluster states, serving as key resources for measurement-based quantum computation (MBQC). Generating large-scale cluster states and verifying their SPT phases are essential steps toward practical MBQC, which however still presents significant experimental challenges. In this work, we address these challenges by utilizing advanced superconducting hardware with optimized gate operations, enhanced readout fidelity, and error mitigation techniques. We successfully generate and verify 95-qubit one-dimensional and 72-qubit two-dimensional genuine entangled cluster states, achieving fidelities of $0.5603 \pm 0.0084$ and $0.5519 \pm 0.0054$, respectively. Leveraging these high-fidelity cluster states, we investigate SPT phases through quantum teleportation across all 95 qubits and demonstrate input-state-dependent robustness against symmetry-breaking perturbations, highlighting the practicality and intrinsic robustness of MBQC enabled by the SPT order. Our results represent a significant advancement in large-scale entanglement generation and topological phase simulation, laying the foundation for scalable and practical MBQC using superconducting quantum systems.

quant-ph