SearcharxivSearch

arXiv subjects

Ya Gao

Publications and source records attributed to Ya Gao.

At least 19 recordsLinked to original sources

The capillary $L_p$ dual Minkowski problem for $p>q$ in higher dimensions

We prove existence and uniqueness for the capillary $L_p$ dual Minkowski problem for $p>q$ and contact angle $\theta\in (0,\frac{\pi}{2})$ in $\mathbb{R}^{n+1}$, $n\geq 3$. We reduce it to a Monge-Amp\`ere type equation with a Robin boundary condition on the unit spherical cap, by building a new auxiliary function to obtain the $C^2$ estimate for $p>q$ in $\mathbb{R}^{n+1}$ ($n\geq 3$), we then prove that there exists a unique smooth solution that solves this problem provided $\theta\in (0, \frac{\pi}{2})$.

math.AP

Secrecy Energy Efficiency for IRS-Assisted Low-Altitude Communications: A D3QN-PER Based Approach

To address the security and energy efficiency challenges in low-altitude economy (LAE) wireless communications, we develop a secure synergistic network integrating unmanned aerial vehicle (UAV) and intelligent reflecting surface (IRS), with an emphasis on maximizing secrecy energy efficiency (SEE) for downlink transmission scenarios. In particular, firstly, we establish the channel transmission models for UAV-IRS assisted LAE communications network. Then, we formulate a non-convex fractional optimization problem for SEE maximization, involving three tightly coupled variables, i.e., the beamforming, IRS phase and UAV trajectory. To tackle the fractional structure and variable coupling, Dinkelbach's method and equivalent transformations are leveraged to reformulate the objective function, which is then decoupled and decomposed into three independent subproblems via an alternating optimization strategy for iterative resolution. Slack variables and Semidefinite Relaxation (SDR) are further employed to convexify the subproblems of beamforming and IRS phase shift optimization, thereby obtaining their optimal solutions. For the UAV trajectory optimization subproblem, we propose a D3QN-PER algorithm, which integrates a Dueling Double Deep Q-Network with Prioritized Experience Replay, to tackle the slow convergence and training instability inherent in conventional Deep Q-Network (DQN). Numerical simulations validate the performance for our proposed joint optimization scheme. Comparative results demonstrate that the developed D3QN-PER-based algorithm outperforms existing state-of-the-art learning approaches which verifies its superiority in improving SEE for UAV-IRS-assisted LAE wireless communications network.

cs.IT

Full-Path Nonlinear Modeling of Microwave Power Transmission Through Ionospheric Plasma for Space Solar Power Station

Space Solar Power Station (SSPS) concepts rely on gigawatt-class microwave beams to carry orbital solar energy through the ionosphere, where the beam and the plasma form a coupled nonlinear system: the field heats electrons, the heating alters the collision frequency and plasma density, and the modified medium in turn reshapes the field. To our knowledge, this work is the first study to quantify this two-way interaction between microwave power transmission and the ionospheric plasma environment through full-path nonlinear modeling. The 340 km path from 400 km to 60 km altitude is reconstructed by 34 cascaded two-dimensional axisymmetric finite-element full-wave segments with complex-field transfer, using International Reference Ionosphere (IRI) electron-density and NRLMSISE-00 neutral-atmosphere inputs. A Shallow Neural Network (SNN) surrogate replaces the implicit electron energy balance with an explicit closure that maps altitude and local field magnitude to electron temperature and effective collision frequency, enabling stable nonlinear iteration. For 1 GW beams at 2.45 GHz and 5.8 GHz, the volume-integrated Ohmic deposition is 29.4 kW and 5.11 kW, respectively -- fractional losses of order $10^{-5}$ -- and the ratio between the two bands follows the $\omega^{-2}$ scaling of collisional absorption. The deposition concentrates near 95 km altitude, where the product of electron density and collision frequency peaks, whereas the electron-temperature perturbation (up to 3815 K) maximizes in the F region, where cooling is weakest; ponderomotive density depletion remains below 0.02\%. The ionosphere is therefore effectively transparent to the SSPS power budget but not to the beam phase: localized heating and refractive perturbation accumulate phase-front distortion relevant to phased-array beam control, rectenna phase compensation, and environmental assessment.

physics.plasm-ph

Evidence-State Rewards for Long-Context Reasoning

Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods usually reward final answers or static evidence extraction, offering little feedback on how intermediate actions change the model's evidence state. We propose Maven, a reinforcement learning framework with an editable evidence memory. Maven defines an answer-conditioned evidence-state value and rewards action-level state transitions: add actions are credited by marginal gain and hindsight contribution, link actions by evidence synergy, and drop actions by improved answer support after removing misleading evidence. These rewards are assigned to the corresponding action spans in GRPO. Across Llama and Qwen models on LongBench v2, LongReason, and RULER, Maven outperforms outcome-only RL and evidence-identification baselines, producing more sufficient evidence sets and lower distractor retention. Our results show that long-context RL benefits from optimizing stateful evidence navigation rather than one-shot evidence extraction.

cs.AI

Cascaded Rydberg antiblockade: Multi-atom excitation dynamics and entanglement

We propose a cascaded Rydberg antiblockade (RAB) regime via a Floquet modulation in four fully connected interacting atoms, which establishes a new synthetic dimension, Dicke-state lattice (DSL), in the space of collective spin excitations. By applying a global periodic driving, we synthesize an effective Hamiltonian that enables perfect state transfer across the five-site DSL with multiple programmable pathways from stepwise nearest-neighbor jumps to a single-step transition. This DSL platform further allows us to simulate a dynamic Su-Schrieffer-Heeger model, where soft quantum control is employed to achieve topologically inspired full RAB $|0000\rangle \to |1111\rangle$ with enhanced robustness against disorder. Moreover, by incorporating the shortcut to adiabaticity technique, we generate high-fidelity entangled twin-Fock and Greenberger-Horne-Zeilinger states on the four atoms within sub-microsecond timescales, outperforming the speed limits of conventional adiabatic protocols. Our work demonstrates a flexible and programmable synthetic dimension for quantum simulation and multipartite entanglement engineering in Rydberg atom arrays, paving the way for the future development of quantum information processing.

quant-ph

Global adiabatic criterion for fast topological photon transfer in Fock-state lattices

Topological state transfer in Fock-state lattices has been demonstrated with high speed using sinusoidal profiles of coupling, yet the underlying reason has remained unclear. A global adiabatic criterion (GAC) is developed to bound the infidelity by the mean and variance of the nonadiabatic factor. The GAC reveals that the key to fast transfer is not a constant energy gap but the vanishing nonadiabaticity variance. For power-law coupling profiles, the variance vanishes only for the sinusoidal shape, which is thus globally optimal. Incorporating experimental decoherence parameters, it is predicted that the optimal transfer duration for a five-photon state is 161 ns, far shorter than 600 ns used in the experiment, reducing time by over 73% while increasing transferred photons by 29%. The optimal duration follow a simple linear scaling with photon number, providing a practical guideline. Through constructing an alternative constant-gap coupling family, it is confirmed that a constant gap alone is not sufficient for fast topological photon transfer. The essential condition is uniformity of nonadiabaticity. This work offers a rigorous explanation for the observed speed and a general framework for fast topological photonics engineering.

quant-ph

Physics-informed neural networks for quantitative assessment of cancellous bone microstructure from photoacoustic signals

Artificial intelligence (AI) empowers innovative diagnostic tools for common diseases, yet its clinical application in skeletal health evaluation is constrained by unsatisfactory accuracy, owing to the inherent porous and poroelastic biophysical features of bone. To address such bottlenecks amid global population aging, this study targets skeletal health and develops a reliable AI framework for precise bone microstructural characterization. We proposed Biot-PINN, a physics-informed neural network embedded with Biot's poroelasticity theory to characterize mechanical responses and wave propagation in poroelastic bone tissues. By decoding photoacoustic signals encoding bone mineral and microstructural features, the framework enables automatic bone microstructural grading. Experimental results reveal that Biot-PINN reaches an accuracy of 97%, markedly surpassing traditional data-driven approaches and providing a robust solution for early skeletal health diagnosis.

physics.med-ph

A Physics-Constrained Learning Framework for Wave Propagation in Complex Poroelastic Multilayered Media

Wave propagation through complex poroelastic multilayered media is difficult to model and invert because pronounced heterogeneity, scattering, mode conversion and fluid-solid coupling jointly distort acoustic signals during propagation. Here we present Physics-Constrained Learning for Complex Multilayered Media (PCL-CMM), a general framework that integrates Biot's poroelastic theory with the elastic wave equation to bridge the gap between physically rigorous wave modelling and data-driven learning. PCL-CMM constructs a high-fidelity digital twin that dynamically computes an effective acoustic stiffness tensor for forward wave modelling and incorporates the resulting physical constraint as a loss term to regularize the training of deep neural networks. We demonstrate PCL-CMM on transcranial photoacoustic imaging, where skull-induced acoustic distortions severely degrade image formation. Across simulations and ex vivo experiments, PCL-CMM effectively compensates for these distortions and improves SSIM by more than 0.06 compared with purely data-driven neural networks. This work establishes a physics-constrained learning framework for acoustic wave modelling in complex poroelastic multilayered media.

physics.med-ph

Edit Knowledge, Not Just Facts via Multi-Step Reasoning over Background Stories

Enabling artificial intelligence systems, particularly large language models, to update knowledge and flexibly apply it during reasoning remains a central challenge. Existing knowledge editing approaches emphasize atomic facts, improving factual recall but often failing to integrate updated information into a coherent framework usable across contexts. In this work, we argue that knowledge update is fundamentally a reasoning problem rather than a memorization problem. Consequently, a model should be trained in situations where the new information is instrumental to solving a task, combined with pre-existing knowledge, and exercised through multi-step reasoning. Based on this insight, we propose a training strategy based on three principles. First, new knowledge is introduced as a coherent background story that contextualizes novel facts and explains their relation to existing knowledge. Second, models are trained using self-generated multi-hop questions that require multi-step reasoning involving the new information. Third, training is done using knowledge distillation, forcing a student model to internalize the teacher's reasoning behavior without access to the novel information. Experiments show that models trained with this strategy effectively leverage newly acquired knowledge during reasoning and achieve remarkable performance on challenging questions that require combining multiple new facts.

cs.AI

DCG ReID: Disentangling Collaboration and Guidance Fusion Representations for Multi-modal Vehicle Re-Identification

Multi-modal vehicle Re-Identification (ReID) aims to leverage complementary information from RGB, Near Infrared (NIR), and Thermal Infrared (TIR) modalities to retrieve the same vehicle. The challenges of multi-modal vehicle ReID arise from the uncertainty of modality quality distribution induced by inherent discrepancies across modalities, resulting in distinct conflicting fusion requirements for data with balanced and unbalanced quality distributions. Existing methods handle all multi-modal data within a single fusion model, overlooking the different needs of the two data types and making it difficult to decouple the conflict between intra-class consistency and inter-modal heterogeneity. To this end, we propose Disentangle Collaboration and Guidance Fusion Representations for Multi-modal Vehicle ReID (DCG-ReID). Specifically, to disentangle heterogeneous quality-distributed modal data without mutual interference, we first design the Dynamic Confidence-based Disentangling Weighting (DCDW) mechanism: dynamically reweighting three-modal contributions via interaction-derived modal confidence to build a disentangled fusion framework. Building on DCDW, we develop two scenario-specific fusion strategies: (1) for balanced quality distributions, Collaboration Fusion Module (CFM) mines pairwise consensus features to capture shared discriminative information and boost intra-class consistency; (2) for unbalanced distributions, Guidance Fusion Module (GFM) implements differential amplification of modal discriminative disparities to reinforce dominant modality advantages, guide auxiliary modalities to mine complementary discriminative info, and mitigate inter-modal divergence to boost multi-modal joint decision performance. Extensive experiments on three multi-modal ReID benchmarks (WMVeID863, MSVR310, RGBNT100) validate the effectiveness of our method. Code will be released upon acceptance.

cs.CV

Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models

Large Vision-Language Models (LVLMs) typically process visual inputs as a prefix to the language decoder. As the model autoregressively generates text, this initial visual information inevitably undergoes "dilution" leading the model to over-rely on language priors and hallucinate objects. Existing interventions attempt to correct this by contrasting logits or iteratively refining outputs, but they incur prohibitive latency costs. We propose Residual-Update Directed DEcoding Regulation (RUDDER), a framework that counters visual dilution by creating a persistent visual anchor. We extract a robust evidence direction (CARD) directly from the model's prefill residual updates, and inject it into the decoding process. This injection is modulated by an adaptive gate, the Beta Gate, which acts as a trust mechanism and ensures the visual reminder is applied only when necessary. Experiments on LLaVA-1.5 (7B/13B), Idefics2, InstructBLIP, and Qwen2.5-VL demonstrate that RUDDER consistently mitigates hallucination (with greedy decoding, RUDDER reduces CHAIR_S by an average of 24.4% and CHAIR_i by 23.6% relative) and scales effectively across architectures, all while maintaining >96.0% throughput.

cs.CV

The $L_p$ dual Minkowski problem for capillary hypersurfaces

In this paper, we consider the $L_p$ dual Minkowski problem for capillary hypersurfaces for $p>q$ and $q\leq 1$, which aims to find a capillary convex body with a prescribed capillary $(p,q)$-th dual curvature measure in the Euclidean half-space. We reduce it to a Monge-Amp\`ere type equation with a Robin boundary condition on the unit spherical cap, we prove that there exists a unique smooth solution that solves this problem provided $\theta\in (0,\frac{\pi}{2})$.

math.DG

Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization

As the general capabilities of artificial intelligence (AI) agents continue to evolve, their ability to learn to master multiple complex tasks through experience remains a key challenge. Current LLM agents, particularly those based on proprietary language models, typically rely on prompts to incorporate knowledge about the target tasks. This approach does not allow the agent to internalize this information and instead relies on ever-expanding prompts to sustain its functionality in diverse scenarios. This resembles a system of notes used by a person affected by anterograde amnesia, the inability to form new memories. In this paper, we propose a novel method to train AI agents to incorporate knowledge and skills for multiple tasks without the need for either cumbersome note systems or prior high-quality demonstration data. Our approach employs an iterative process where the agent collects new experiences, receives corrective feedback from humans in the form of hints, and integrates this feedback into its weights via a context distillation training procedure. We demonstrate the efficacy of our approach by implementing it in a Llama-3-based agent that, after only a few rounds of feedback, outperforms advanced models GPT-4o and DeepSeek-V3 in tasksets requiring correct sequencing of information retrieval, tool use, and question answering.

cs.LG

Query-Guided Self-Supervised Summarization of Nursing Notes

Nursing notes, an important part of Electronic Health Records (EHRs), track a patient's health during a care episode. Summarizing key information in nursing notes can help clinicians quickly understand patients' conditions. However, existing summarization methods in the clinical setting, especially abstractive methods, have overlooked nursing notes and require reference summaries for training. We introduce QGSumm, a novel query-guided self-supervised domain adaptation approach for abstractive nursing note summarization. The method uses patient-related clinical queries for guidance, and hence does not need reference summaries for training. Through automatic experiments and manual evaluation by an expert clinician, we study our approach and other state-of-the-art Large Language Models (LLMs) for nursing note summarization. Our experiments show: 1) GPT-4 is competitive in maintaining information in the original nursing notes, 2) QGSumm can generate high-quality summaries with a good balance between recall of the original content and hallucination rate lower than other top methods. Ultimately, our work offers a new perspective on conditional text summarization, tailored to clinical applications.

cs.CL

Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding

Understanding of video creativity and content often varies among individuals, with differences in focal points and cognitive levels across different ages, experiences, and genders. There is currently a lack of research in this area, and most existing benchmarks suffer from several drawbacks: 1) a limited number of modalities and answers with restrictive length; 2) the content and scenarios within the videos are excessively monotonous, transmitting allegories and emotions that are overly simplistic. To bridge the gap to real-world applications, we introduce a large-scale Subjective Response Indicators for Advertisement Videos dataset, namely SRI-ADV. Specifically, we collected real changes in Electroencephalographic (EEG) and eye-tracking regions from different demographics while they viewed identical video content. Utilizing this multi-modal dataset, we developed tasks and protocols to analyze and evaluate the extent of cognitive understanding of video content among different users. Along with the dataset, we designed a Hypergraph Multi-modal Large Language Model (HMLLM) to explore the associations among different demographics, video elements, EEG, and eye-tracking indicators. HMLLM could bridge semantic gaps across rich modalities and integrate information beyond different modalities to perform logical reasoning. Extensive experimental evaluations on SRI-ADV and other additional video-based generative performance benchmarks demonstrate the effectiveness of our method. The codes and dataset will be released at https://github.com/mininglamp-MLLM/HMLLM.

cs.CV

Asymptotic convergence for a class of fully nonlinear inverse curvature flows in a cone

For a given smooth convex cone in the Euclidean $(n+1)$-space $\mathbb{R}^{n+1}$ which is centered at the origin, we investigate the evolution of strictly mean convex hypersurfaces, which are star-shaped with respect to the center of the cone and which meet the cone perpendicularly, along an inverse curvature flow with the speed equal to $\left(f(r)H\right)^{-1}$, where $f$ is a positive function of the radial distance parameter $r$ and $H$ is the mean curvature of the evolving hypersurfaces. The evolution of those hypersurfaces inside the cone yields a fully nonlinear parabolic Neumann problem. Under suitable constraints on the first and the second derivatives of the radial function $f$, we can prove the long-time existence of this flow, and moreover the evolving hypersurfaces converge smoothly to a piece of the round sphere.

math.DG

Knowledge-augmented Graph Neural Networks with Concept-aware Attention for Adverse Drug Event Detection

Adverse drug events (ADEs) are an important aspect of drug safety. Various texts such as biomedical literature, drug reviews, and user posts on social media and medical forums contain a wealth of information about ADEs. Recent studies have applied word embedding and deep learning -based natural language processing to automate ADE detection from text. However, they did not explore incorporating explicit medical knowledge about drugs and adverse reactions or the corresponding feature learning. This paper adopts the heterogenous text graph which describes relationships between documents, words and concepts, augments it with medical knowledge from the Unified Medical Language System, and proposes a concept-aware attention mechanism which learns features differently for the different types of nodes in the graph. We further utilize contextualized embeddings from pretrained language models and convolutional graph neural networks for effective feature representation and relational learning. Experiments on four public datasets show that our model achieves performance competitive to the recent advances and the concept-aware attention consistently outperforms other attention mechanisms.

cs.CL

SynerMix: Synergistic Mixup Solution for Enhanced Intra-Class Cohesion and Inter-Class Separability in Image Classification

To address the issues of MixUp and its variants (e.g., Manifold MixUp) in image classification tasks-namely, their neglect of mixing within the same class (intra-class mixup) and their inadequacy in enhancing intra-class cohesion through their mixing operations-we propose a novel mixup method named SynerMix-Intra and, building upon this, introduce a synergistic mixup solution named SynerMix. SynerMix-Intra specifically targets intra-class mixup to bolster intra-class cohesion, a feature not addressed by current mixup methods. For each mini-batch, it leverages feature representations of unaugmented original images from each class to generate a synthesized feature representation through random linear interpolation. All synthesized representations are then fed into the classification and loss layers to calculate an average classification loss that significantly enhances intra-class cohesion. Furthermore, SynerMix combines SynerMix-Intra with an existing mixup approach (e.g., MixUp, Manifold MixUp), which primarily focuses on inter-class mixup and has the benefit of enhancing inter-class separability. In doing so, it integrates both inter- and intra-class mixup in a balanced way while concurrently improving intra-class cohesion and inter-class separability. Experimental results on six datasets show that SynerMix achieves a 0.1% to 3.43% higher accuracy than the best of either MixUp or SynerMix-Intra alone, averaging a 1.16% gain. It also surpasses the top-performer of either Manifold MixUp or SynerMix-Intra by 0.12% to 5.16%, with an average gain of 1.11%. Given that SynerMix is model-agnostic, it holds significant potential for application in other domains where mixup methods have shown promise, such as speech and text classification. Our code is publicly available at: https://github.com/wxitxy/synermix.git.

cs.CV