SearcharxivSearch

arXiv subjects

Yantong Xie

Publications and source records attributed to Yantong Xie.

7 recordsLinked to original sources

Probabilistic Modeling of Intentions in Socially Intelligent LLM Agents

We present a probabilistic intent modeling framework for large language model (LLM) agents in multi-turn social dialogue. The framework maintains a belief distribution over a partner's latent intentions, initialized from contextual priors and dynamically updated through likelihood estimation after each utterance. The evolving distribution provides additional contextual grounding for the policy, enabling adaptive dialogue strategies under uncertainty. Preliminary experiments in the SOTOPIA environment show consistent improvements: the proposed framework increases the Overall score by 9.0% on SOTOPIA-All and 4.1% on SOTOPIA-Hard compared with the Qwen2.5-7B baseline, and slightly surpasses an oracle agent that directly observes partner intentions. These early results suggest that probabilistic intent modeling can contribute to the development of socially intelligent LLM agents.

cs.AI

Cross-LoRA: A Data-Free LoRA Transfer Framework across Heterogeneous LLMs

Traditional parameter-efficient fine-tuning (PEFT) methods such as LoRA are tightly coupled with the base model architecture, which constrains their applicability across heterogeneous pretrained large language models (LLMs). To address this limitation, we introduce Cross-LoRA, a data-free framework for transferring LoRA modules between diverse base models without requiring additional training data. Cross-LoRA consists of two key components: (a) LoRA-Align, which performs subspace alignment between source and target base models through rank-truncated singular value decomposition (SVD) and Frobenius-optimal linear transformation, ensuring compatibility under dimension mismatch; and (b) LoRA-Shift, which applies the aligned subspaces to project source LoRA weight updates into the target model parameter space. Both components are data-free, training-free, and enable lightweight adaptation on a commodity GPU in 20 minutes. Experiments on ARCs, OBOA and HellaSwag show that Cross-LoRA achieves relative gains of up to 5.26% over base models. Across other commonsense reasoning benchmarks, Cross-LoRA maintains performance comparable to that of directly trained LoRA adapters.

cs.LG

A synchronization-capturing multi-scale solver to the noisy integrate-and-fire neuron networks

The noisy leaky integrate-and-fire (NLIF) model describes the voltage configurations of neuron networks with an interacting many-particles system at a microscopic level. When simulating neuron networks of large sizes, computing a coarse-grained mean-field Fokker-Planck equation solving the voltage densities of the networks at a macroscopic level practically serves as a feasible alternative in its high efficiency and credible accuracy. However, the macroscopic model fails to yield valid results of the networks when simulating considerably synchronous networks with active firing events. In this paper, we propose a multi-scale solver for the NLIF networks, which inherits the low cost of the macroscopic solver and the high reliability of the microscopic solver. For each temporal step, the multi-scale solver uses the macroscopic solver when the firing rate of the simulated network is low, while it switches to the microscopic solver when the firing rate tends to blow up. Moreover, the macroscopic and microscopic solvers are integrated with a high-precision switching algorithm to ensure the accuracy of the multi-scale solver. The validity of the multi-scale solver is analyzed from two perspectives: firstly, we provide practically sufficient conditions that guarantee the mean-field approximation of the macroscopic model and present rigorous numerical analysis on simulation errors when coupling the two solvers; secondly, the numerical performance of the multi-scale solver is validated through simulating several large neuron networks, including networks with either instantaneous or periodic input currents which prompt active firing events over a period of time.

math.NA

Frozen Gaussian Sampling: A Mesh-free Monte Carlo Method For Approximating Semiclassical Schrödinger Equations

In this paper, we develop a Monte Carlo algorithm named the Frozen Gaussian Sampling (FGS) to solve the semiclassical Schrödinger equation based on the frozen Gaussian approximation. Due to the highly oscillatory structure of the wave function, traditional mesh-based algorithms suffer from "the curse of dimensionality", which gives rise to more severe computational burden when the semiclassical parameter \(\ep\) is small. The Frozen Gaussian sampling outperforms the existing algorithms in that it is mesh-free in computing the physical observables and is suitable for high dimensional problems. In this work, we provide detailed procedures to implement the FGS for both Gaussian and WKB initial data cases, where the sampling strategies on the phase space balance the need of variance reduction and sampling convenience. Moreover, we rigorously prove that, to reach a certain accuracy, the number of samples needed for the FGS is independent of the scaling parameter \(\ep\). Furthermore, the complexity of the FGS algorithm is of a sublinear scaling with respect to the microscopic degrees of freedom and, in particular, is insensitive to the dimension number. The performance of the FGS is validated through several typical numerical experiments, including simulating scattering by the barrier potential, formation of the caustics and computing the high-dimensional physical observables without mesh.

math.NA

Investigating the integrate and fire model as the limit of a random discharge model: a stochastic analysis perspective

In the mean field integrate-and-fire model, the dynamics of a typical neuron within a large network is modeled as a diffusion-jump stochastic process whose jump takes place once the voltage reaches a threshold. In this work, the main goal is to establish the convergence relationship between the regularized process and the original one where in the regularized process, the jump mechanism is replaced by a Poisson dynamic, and jump intensity within the classically forbidden domain goes to infinity as the regularization parameter vanishes. On the macroscopic level, the Fokker-Planck equation for the process with random discharges (i.e. Poisson jumps) are defined on the whole space, while the equation for the limit process is on the half space. However, with the iteration scheme, the difficulty due to the domain differences has been greatly mitigated and the convergence for the stochastic process and the firing rates can be established. Moreover, we find a polynomial-order convergence for the distribution by a re-normalization argument in probability theory. Finally, by numerical experiments, we quantitatively explore the rate and the asymptotic behavior of the convergence for both linear and nonlinear models.

math.PR

Investigating the integrate and fire model as the limit of a random discharge model: a stochastic analysis perspective

In the mean field integrate-and-fire model, the dynamics of a typical neuron within a large network is modeled as a diffusion-jump stochastic process whose jump takes place once the voltage reaches a threshold. In this work, the main goal is to establish the convergence relationship between the regularized process and the original one where in the regularized process, the jump mechanism is replaced by a Poisson dynamic, and jump intensity within the classically forbidden domain goes to infinity as the regularization parameter vanishes. On the macroscopic level, the Fokker-Planck equation for the process with random discharges (i.e. Poisson jumps) are defined on the whole space, while the equation for the limit process is on the half space. However, with the iteration scheme, the difficulty due to the domain differences has been greatly mitigated and the convergence for the stochastic process and the firing rates can be established. Moreover, we find a polynomial-order convergence for the distribution by a re-normalization argument in probability theory. Finally, by numerical experiments, we quantitatively explore the rate and the asymptotic behavior of the convergence for both linear and nonlinear models.

math.PR

A structure preserving numerical scheme for Fokker-Planck equations of neuron networks: numerical analysis and exploration

In this work, we are concerned with the Fokker-Planck equations associated with the Nonlinear Noisy Leaky Integrate-and-Fire model for neuron networks. Due to the jump mechanism at the microscopic level, such Fokker-Planck equations are endowed with an unconventional structure: transporting the boundary flux to a specific interior point. While the equations exhibit diversified solutions from various numerical observations, the properties of solutions are not yet completely understood, and by far there has been no rigorous numerical analysis work concerning such models. We propose a conservative and conditionally positivity preserving scheme for these Fokker-Planck equations, and we show that in the linear case, the semi-discrete scheme satisfies the discrete relative entropy estimate, which essentially matches the only known long time asymptotic solution property. We also provide extensive numerical tests to verify the scheme properties, and carry out several sets of numerical experiments, including finite-time blowup, convergence to equilibrium and capturing time-period solutions of the variant models.

math.NA