SearcharxivSearch

arXiv subjects

Yaqi Zhang

Publications and source records attributed to Yaqi Zhang.

At least 19 recordsLinked to original sources

Polarization-Conditioned Fourier-enhanced DeepONet for Electric Field Reconstruction from EFISH Measurements

Electric-field-induced second-harmonic generation (EFISH) is an established laser diagnostic for quantifying electric fields in plasmas, yet ensuring field accuracy remains challenging given the coherent, path-integrated nature of the signal. We address this via machine learning, developing a Polarization-Conditioned Fourier-enhanced Deep Operator Network (PC-FDON) -- a unified operator-learning model that reconstructs field profiles from EFISH measurements across both polarizations and various optical parameters. Its architecture incorporates three advances: (i) a Fourier-enhanced branch providing inductive bias for the Gouy phase shift and wave-vector mismatch; (ii) a polarization-conditioning branch encoding signal polarization via Feature-wise Linear Modulation (FiLM) and gated units, enabling one model to handle both polarizations; and (iii) a physics-informed loss enforcing self-consistency with the governing EFISH equation. Trained on data spanning multiple function families, polarization states, and phase-mismatch values, PC-FDON achieves promising reconstruction under noise-free, incomplete, and noisy inputs, with generalizability comparable to our previous polarization-specific model. Pointwise epistemic uncertainty estimates via Monte Carlo dropout reflect model confidence and enable out-of-distribution (OOD) detection through a location-dependent exceedance fraction metric. Validation is performed on realistic electrode configurations under both polarizations and varying Rayleigh ranges, including a simulated surface dielectric barrier discharge where the framework correctly flags OOD inputs, and experimental data showing good agreement with simulations. The architecture -- spectral inductive bias, conditional modulation, and dataset-specific uncertainty -- shows strong potential for broader application beyond plasma diagnostics.

physics.plasm-ph

On the uniqueness of solutions to quadratic BSDEs with non-convex generators and unbounded terminal conditions: the certain exponential moment case

With the terminal value $|ξ|$ admitting some given exponential moments, we propose and prove several existence and uniqueness results for the unbounded solutions of quadratic backward stochastic differential equations whose generators may be represented as a uniformly continuous (not necessarily locally Lipschitz continuous) perturbation of some convex/concave function with quadratic growth. This perturbation satisfies various feasible conditions such as boundedness, sub-linear growth or linear growth. In particular, in some cases, the first component of the unique solution can be expressed as the value function of an optimal control problem. These results improves those posed in Delbaen, Hu and Richou [AIHP, 2011] and Fan, Hu and Tang [2020, CRM] to some extent. The critical case is also tackled, which strengthens the main result of Delbaen, Hu and Richou [DCDS, 2011].

math.PR

Weighted $L^p$ solutions of scalar BSDEs with general unbounded stochastic coefficients

This paper is devoted to solving one-dimensional backward stochastic differential equations (BSDEs in short) with a general random terminal time $τ$ taking values in the extended nonnegative real numbers. The generator $g$ of BSDEs satisfies some stochastic growth/continuity conditions in the state variables $(y,z)$, featuring unbounded stochastic coefficients $μ_\cdot\in\R$ and $ν_\cdot\in\R_+$ satisfying $\int_0^τ(|μ_t|+ν^2_t) {\rm d}t<+\infty$. For any given real $p>1$, let $ρ_\cdot\geq μ_\cdot+\fracθ{2(p-1)}ν_\cdot^2$ (instead of $ρ_\cdot\geqμ_\cdot+\fracθ{2[1\wedge(p-1)]}ν_\cdot^2$ used in Zhang, Li, Hu and Fan [2026, arXiv:2603.13873v1]) be a real-valued process for some constant $θ>1$ such that $\int_0^τ|ρ_t|{\rm d}t<+\infty$. We work within a weighted $L^p$ space with the weighting factor $e^{\int_0^t ρ_r{\rm d}r}$. Within this framework, we establish several innovative results on the weighted $L^p$ solutions of BSDEs: an existence result, an existence and uniqueness result, an existence and uniqueness result of the minimal (maximal) solution, and two comparison theorems. These findings unify and improve some existing results. Some novel ideas are employed to address the challenges posed by general unbounded stochastic coefficients and general weighted spaces.

math.PR

Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning

In large urban areas, planning multi-day travel itineraries is challenging due to the abundance of Points of Interest (POIs), diverse user preferences, and constraints such as opening hours. Effective solutions must dynamically accommodate diverse traveler requirements while optimizing for satisfaction and feasibility within limited computation time. This paper addresses these challenges through introducing an innovative framework that integrates Large Language Models (LLMs) to dynamically capture user requirements with precision and flexibility, and an enhanced Greedy Randomized Adaptive Search Procedure (GRASP) algorithm as a well-suited preference-aware planner to generate feasible multi-day itineraries. The effectiveness of our integrated approach is demonstrated through extensive experiments on two real-world urban datasets from Beijing and Tianjin. Our framework significantly outperforms state-of-the-art (SOTA) methods, improving the average total itinerary score by at least 4.52% and 11.09% across 5,040 user cases with diverse preferences in the two datasets. Furthermore, through end-to-end algorithmic enhancements, it achieves notable average improvements of 17.95% and 26.07% in the computed metrics, while also delivering substantial gains in time efficiency -- realizing average performance increases of 4.64% and 25.55% within shorter computation times compared to suboptimal methods that require multiple iterations. These outcomes underscore our method's superiority in delivering both enhanced itinerary quality and computational efficiency over existing methodologies.

cs.AI

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Despite impressive advances, existing approaches are often constrained by limited control modalities, task specificity, and focus solely on body motion representations.In this paper, we present MotionGPT-2, a unified Large Motion-Language Model (LMLM) that addresses these limitations. MotionGPT-2 accommodates multiple motion-relevant tasks and supporting multimodal control conditions through pre-trained Large Language Models (LLMs). It quantizes multimodal inputs-such as text and single-frame poses-into discrete, LLM-interpretable tokens, seamlessly integrating them into the LLM's vocabulary. These tokens are then organized into unified prompts, guiding the LLM to generate motion outputs through a pretraining-then-finetuning paradigm. We also show that the proposed MotionGPT-2 is highly adaptable to the challenging 3D holistic motion generation task, enabled by the innovative motion discretization framework, Part-Aware VQVAE, which ensures fine-grained representations of body and hand movements. Extensive experiments and visualizations validate the effectiveness of our method, demonstrating the adaptability of MotionGPT-2 across motion generation, motion captioning, and generalized motion completion tasks.

cs.CV

NüshuVoice: Reviving the Voice of Endangered Nüshu with Pitch-Aware Text-to-Speech

Nüshu is an endangered phonetic script historically used by women in Jiangyong County, southern Hunan, China. While existing computational studies of Nüshu mainly focus on textual digitization and visual recognition, the acoustic reconstruction of its authentic pronunciation remains largely unexplored. Building a Nüshu text-to-speech (TTS) system is particularly challenging because available recordings are extremely limited and mostly consist of isolated syllable-level pronunciations rather than natural sentence-level utterances. In this work, we introduce NüshuVoice, the first TTS benchmark for Nüshu. We construct a sentence-level Nüshu text-to-audio dataset that aligns standardized Unicode Nüshu text, phonetic transcriptions, standard Chinese translations, and archival recordings. To synthesize speech under this extreme low-resource setting, we propose Nüshu-PitchVITS, an F0-conditioned VITS framework that leverages Nüshu's five-level pitch notation as an explicit prosodic inductive bias. Experimental results show that Nüshu-PitchVITS outperforms strong TTS baselines in spectral fidelity, pitch reconstruction, and human-rated intelligibility. We publicly release the dataset and code at: https://anonymous.4open.science/r/Nvshu-TTS-2EB6.

cs.CL

Agentic AI-Driven UAV Network Deployment: An LLM-Enhanced Exact Potential Game Approach

Unmanned aerial vehicular network (UAVN) is envisioned to provide flexible connectivity, wide-area coverage, and low-latency services in dynamic environments. From an agentic artificial intelligence (Agentic AI) perspective, UAVNs naturally operate as multi-agent systems, where UAVs act as intelligent agents that coordinate deployment and networking decisions to achieve global performance objectives. However, the strong coupling between discrete link decisions and continuous deployment parameters makes UAVN deployment optimization a mixed-integer nonconvex problem, resulting in challenges in scalability, efficiency, and solution consistency under dynamic network conditions. This paper proposes a dual spatial-scale UAVN deployment optimization framework based on exact potential games (EPGs), enhanced by Agentic AI. At the large spatial scale, a log-linear learning based EPG (L3-EPG) algorithm is developed to optimize inter-UAV link configurations, enabling sparse yet connected network topologies while reducing redundant links and interference. At the small spatial scale, an approximate gradient based EPG (AG-EPG) algorithm jointly optimizes UAV deployment, transmission power allocation, and ground user (GU) association to improve network throughput and latency. To further enhance adaptability across heterogeneous scenarios, a large language model (LLM) is incorporated as a knowledge-driven decision enhancer to automatically generate utility weights according to network characteristics, alleviating reliance on manual parameter tuning. Simulation results demonstrate that the proposed framework consistently outperforms baseline methods in terms of energy consumption, end-to-end latency, and system throughput.

cs.DC

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key bottlenecks are no longer modality coverage or latency alone, but the interaction paradigm itself. First, perception and response are still separated into alternating phases, preventing models from incorporating new inputs for timely adjustment during generation. Second, most current models remain reactive, responding only to explicit user requests instead of acting proactively in the evolving multimodal environment. We present MiniCPM-o 4.5, our latest effort towards human-like multimodal interaction, which mitigates these gaps by real-time full-duplex omni-modal interaction. It can see, listen, and speak simultaneously in real-time, while also exhibiting proactive behaviors such as issuing reminders or comments based on its continuous understanding of the live scene. The key technique behind MiniCPM-o 4.5 is Omni-Flow, a unified streaming framework that aligns omni-modal inputs and outputs along a shared temporal axis. This formulation converts conventional turn-based interaction into a full-duplex, time-aligned process, enabling simultaneous perception and response and allowing proactive behavior to arise within the same framework. With a total of 9B parameters, MiniCPM-o 4.5 approaches Gemini 2.5 Flash in vision-language capabilities, delivering state-of-the-art open-source performance at its scale. It also surpasses Qwen3-Omni-30B-A3B in omni-modal understanding and delivers better speech generation, with significantly higher computation efficiency. Driven by its efficient architecture design and inference optimization, the model can perform real-time full-duplex omni-modal interaction on edge devices with less than 12GB RAM cost.

cs.CL

Solvability of BSDEs with possibly unbounded stochastic coefficients on a general weighted $L^p$ space

This paper is devoted to solving a multidimensional backward stochastic differential equation (BSDE for short) with a general random terminal time $τ$ taking values in $[0,+\infty]$. The generator $g$ of such BSDE satisfies a stochastic monotonicity condition in the state variable $y$ and a stochastic Lipschitz condition in the state variable $z$ with possibly unbounded stochastic coefficients $μ_\cdot\in\R$ and $ν_\cdot\in\R_+$ satisfying $\int_0^τ(|μ_t|+ν^2_t) {\rm d}t<+\infty$, along with a very general growth in $y$ that is more easily verified and weaker than existing ones. Let $p>1$ be a given constant and $ρ_\cdot\geq μ_\cdot+\fracθ{2[1\wedge(p-1)]}ν_\cdot^2$ be a given real-valued process for some constant $θ>1$ such that $\int_0^τ|ρ_t|{\rm d}t<+\infty$. In a general weighted $L^p$ space with a weighted factor $e^{\int_0^t ρ_r{\rm d}r}$, we establish an existence and uniqueness result for the adapted solution of previous BSDE when the terminal value satisfies an associated weighted integrability condition, broadening the scope of the process $ρ_\cdot$ in the weighted factor and thereby unifying and strengthening some corresponding existing results obtained in \citet{DarlingandPardoux1997}, \citet{Briand2003}, \citet{LiFan2024SD} and \citet{Li2025}. Some innovative ideas are presented in order to address the general weighted space and the very general growth condition. As applications, we prove the existence of viscosity solutions for parabolic and elliptic PDEs linked with previous BSDEs under some general assumptions on their nonlinear terms, and establish a dual representation of an unbounded dynamic concave utility defined on a general weighted $L^p$ space via the weighted $L^p$ solutions of previous BSDEs.

math.PR

DecIF: Improving Instruction-Following through Meta-Decomposition

Instruction-following has emerged as a crucial capability for large language models (LLMs). However, existing approaches often rely on pre-existing documents or external resources to synthesize instruction-following data, which limits their flexibility and generalizability. In this paper, we introduce DecIF, a fully autonomous, meta-decomposition guided framework that generates diverse and high-quality instruction-following data using only LLMs. DecIF is grounded in the principle of decomposition. For instruction generation, we guide LLMs to iteratively produce various types of meta-information, which are then combined with response constraints to form well-structured and semantically rich instructions. We further utilize LLMs to detect and resolve potential inconsistencies within the generated instructions. Regarding response generation, we decompose each instruction into atomic-level evaluation criteria, enabling rigorous validation and the elimination of inaccurate instruction-response pairs. Extensive experiments across a wide range of scenarios and settings demonstrate DecIF's superior performance on instruction-following tasks. Further analysis highlights its strong flexibility, scalability, and generalizability in automatically synthesizing high-quality instruction data.

cs.CL

Weighted solutions of random time horizon BSDEs with stochastic monotonicity and general growth generators and related PDEs

This study focuses on a multidimensional backward stochastic differential equation (BSDE) with a general random terminal time $τ$ taking values in $[0,+\infty]$. The generator $g$ satisfies a stochastic monotonicity condition in the first unknown variable $y$ and a stochastic Lipschitz continuity condition in the second unknown variable $z$, and it can have a more general growth with respect to $y$ than the classical one stated in (H5) of \cite{Briand2003}. Without imposing any restriction of finite moment on the stochastic coefficients, we establish a general existence and uniqueness result for the weighted solution of such BSDE in a proper weighted $L^2$-space with a suitable weighted factor. This result is proved via some innovative ideas and delicate analytical techniques, and it unifies and strengthens some existing works on BSDEs with stochastic monotonicity generators, BSDEs with stochastic Lipschitz generators, and BSDEs with deterministic Lipschitz/monotonicity generators. Then, a continuous dependence property and a stability theorem for the weighted $L^2$-solutions are given. We also derive the nonlinear Feynman-Kac formulas for both parabolic and elliptic PDEs in our context.

math.PR

Smaller Language Models Are Better Instruction Evolvers

Instruction tuning has been widely used to unleash the complete potential of large language models. Notably, complex and diverse instructions are of significant importance as they can effectively align models with various downstream tasks. However, current approaches to constructing large-scale instructions predominantly favour powerful models such as GPT-4 or those with over 70 billion parameters, under the empirical presumption that such larger language models (LLMs) inherently possess enhanced capabilities. In this study, we question this prevalent assumption and conduct an in-depth exploration into the potential of smaller language models (SLMs) in the context of instruction evolution. Extensive experiments across three scenarios of instruction evolution reveal that smaller language models (SLMs) can synthesize more effective instructions than LLMs. Further analysis demonstrates that SLMs possess a broader output space during instruction evolution, resulting in more complex and diverse variants. We also observe that the existing metrics fail to focus on the impact of the instructions. Thus, we propose Instruction Complex-Aware IFD (IC-IFD), which introduces instruction complexity in the original IFD score to evaluate the effectiveness of instruction data more accurately. Our source code is available at: \href{https://github.com/HypherX/Evolution-Analysis}{https://github.com/HypherX/Evolution-Analysis}

cs.CL

CountMamba: Exploring Multi-directional Selective State-Space Models for Plant Counting

Plant counting is essential in every stage of agriculture, including seed breeding, germination, cultivation, fertilization, pollination yield estimation, and harvesting. Inspired by the fact that humans count objects in high-resolution images by sequential scanning, we explore the potential of handling plant counting tasks via state space models (SSMs) for generating counting results. In this paper, we propose a new counting approach named CountMamba that constructs multiple counting experts to scan from various directions simultaneously. Specifically, we design a Multi-directional State-Space Group to process the image patch sequences in multiple orders and aim to simulate different counting experts. We also design Global-Local Adaptive Fusion to adaptively aggregate global features extracted from multiple directions and local features extracted from the CNN branch in a sample-wise manner. Extensive experiments demonstrate that the proposed CountMamba performs competitively on various plant counting tasks, including maize tassels, wheat ears, and sorghum head counting.

cs.CV

Compatible Forts and Maximum Nullity of a Graph

We consider bounds on maximum nullity of a graph via transversal numbers of compatible collections of forts. Results include generalizations of theorems from symmetric to combinatorially symmetric matrices, special bases of matrix nullspaces derived from transversal sets, and examples of issues that arise when considering only minimal forts and how to avoid them. We also show an important difference between constructing symmetric and combinatorially symmetric matrices associated to a graph whose nullspaces are supported on collections of disjoint forts.

math.CO

MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion directly from textual action descriptions, they often support only a single modality of the control signal, which limits their application in the real digital human industry. This paper presents a Motion General-Purpose generaTor (MotionGPT) that can use multimodal control signals, e.g., text and single-frame poses, for generating consecutive human motions by treating multimodal signals as special input tokens in large language models (LLMs). Specifically, we first quantize multimodal control signals into discrete codes and then formulate them in a unified prompt instruction to ask the LLMs to generate the motion answer. Our MotionGPT demonstrates a unified human motion generation model with multimodal control signals by tuning a mere 0.4% of LLM parameters. To the best of our knowledge, MotionGPT is the first method to generate human motion by multimodal control signals, which we hope can shed light on this new direction. Visit our webpage at https://qiqiapink.github.io/MotionGPT/.

cs.CV

New Structures and their Applications to Variants of Zero Forcing and Propagation Time

We introduce a generalization of the concept of a chronological list of forces, called a relaxed chronology. This concept is used to introduce a new way of formulating the standard zero forcing process, which we refer to as parallel increasing path covers, or PIPs. The combinatorial properties of PIPs are utilized to identify bounds comparing standard zero forcing propagation time to positive semidefinite propagation time. A collection of paths within a set of PSD forcing trees, called a path bundle, is used to identify the PSD forcing analog of the reversal of a standard zero forcing process, as well as to draw a connection between PSD forcing and rigid-linkage forcing.

math.CO

The Spark of Symmetric Matrices Described by a Graph

We investigate the sparsity of null vectors of real symmetric matrices whose off-diagonal pattern of zero and nonzero entries is described by the adjacencies of a graph. We use the definition of the spark of a matrix, the smallest number of nonzero coordinates of any null vector, to define the spark of a graph as the smallest possible spark of a corresponding matrix. We study connections of graph spark to well-known concepts including minimum rank, forts, orthogonal representations, Parter and Fiedler vertices, and vertex connectivity.

math.CO

Gauge-Invariant Double-Copies via Recursion

We prove that all tree-level amplitudes in pure (super-)gravity can be expressed as term-wise, gauge-invariant double-copies of those of pure (super-)Yang-Mills obtained via BCFW recursion. These representations are far from unique: varying the recursive scheme leads to a wide variety of distinct, but equally valid representations of gravitational amplitudes, all realized as double-copies.

hep-th