SearcharxivSearch

arXiv subjects

Fengyi Song

Publications and source records attributed to Fengyi Song.

9 recordsLinked to original sources

Entropy-based Code Adversarial Translation for Real-world Repository Migration

LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnable application because long-horizon translation challenges LLM-based agents' ability to maintain repository-level migration objectives. In this work, we propose Entropy-based Code Adversarial Translation (ECAT), a multi-agent framework for automated Android-to-HarmonyOS repository migration. ECAT formulates repository migration as adversarial entropy minimization through a generator-discriminator architecture. The discriminator measures migration quality using a unified metric called Code Entropy and produces text gradients that specify both file-level generation directives and the skills needed to execute them. Guided by these optimization signals, the generator iteratively updates the repository, and each update is accepted only if it reduces Code Entropy. Repeated generator--discriminator interactions progressively drive the migration from an initial template toward a functionally complete HarmonyOS repository. Successful low-entropy trajectories are further distilled into a self-evolving memory tree, enabling transferable migration knowledge across repositories. We also introduce A2H-RepoBench, the first real-world benchmark for Android-to-HarmonyOS repository migration, covering applications from tens of thousands to hundreds of thousands of lines of code. Evaluated by node alignment and an agent-based functional judge, ECAT achieves 74.7% overall migration quality and consistently outperforms existing agent-based methods across repositories of different scales.

cs.AI

Elementary magnons and interacting multi-magnon quasiparticles in the effective spin-$\frac{1}{2}$ kagome-staircase magnet Co$_{3}$V$_{2}$O$_{8}$

The excitation spectrum of an anisotropic magnet provides a direct link between its microscopic Hamiltonian and interaction-driven quasiparticles. Here we use high-resolution time-domain terahertz spectroscopy to map the magnetic excitations of the three-dimensional kagome-staircase compound Co$_{3}$V$_{2}$O$_{8}$ as functions of temperature and magnetic field. At low energies, polarization-resolved spectra identify magnetic-dipole-active one-magnon modes and track their evolution across the ferromagnetic and spin-density-wave phases. Combining their field dependence with previously reported inelastic-neutron-scattering dispersions, we determine an effective spin-$\frac{1}{2}$ Hamiltonian with strongly anisotropic exchange that quantitatively reproduces the one-magnon spectrum. This model provides a noninteracting benchmark for the high-energy response, where we observe sharp branches with field slopes that are two to four times those of the one-magnon modes, together with anticrossings between branches of different magnon numbers. Their sharpness, polarization dependence, and departure from the calculated multi-magnon continua identify them as interacting multi-magnon quasiparticles that can be stabilized by strong exchange anisotropy.

cond-mat.str-el

Elliptical Regularized Hotelling Tests for High-Dimensional Change-Point Detection

We propose an elliptical regularized Hotelling (ERHT) procedure for detecting location changes in high-dimensional sequences with heavy-tailed, cross-sectionally dependent observations. ERHT contrasts spatial medians on adjacent segments using a ridge-regularized inverse of the pooled centered spatial-sign covariance matrix, thereby combining robustness to radial variation with dependence-aware weighting. We establish Gaussian-process limits for the single- and multiple-change scans and joint convergence over a finite set of regularization parameters. These results provide asymptotically exact calibration of a Cauchy-aggregated adaptive test through the joint Gaussian limit, together with guarantees for local power and single-change localization. We further embed the ERHT score in wild binary segmentation and prove consistency for estimating the number and locations of multiple changes. Simulations show that ERHT is generally well calibrated and delivers competitive power under heavy-tailed distributions, particularly when cross-sectional dependence is substantial. An analysis of the Fama--French 49 industry portfolios reveals persistent evidence of location instability and identifies four structural breaks.

stat.ME

Difference-Based High-Dimensional Long-Run Covariance Matrix Estimation for Mean-shift Time Series

We consider estimation of high-dimensional long-run covariance matrices for time series with nonconstant means, a setting in which conventional estimators can be severely biased. To address this difficulty, we propose a difference-based initial estimator that is robust to a broad class of mean variations, and combine it with hard thresholding, soft thresholding, and tapering to obtain sparse long-run covariance estimators for high-dimensional data. We derive convergence rates for the resulting estimators under general temporal dependence and time-varying mean structures, showing explicitly how the rates depend on covariance sparsity, mean variation, dimension, and sample size. Numerical experiments show that the proposed methods perform favorably in high dimensions, especially when the mean evolves over time.

stat.ME

Adaptive Test Procedure for High Dimensional Regression Coefficient

We develop a unified $L$-statistic testing framework for high-dimensional regression coefficients that adapts to unknown sparsity. The proposed statistics rank coordinate-wise evidence measures and aggregate the top $k$ signals, bridging classical max-type and sum-type tests. We establish joint weak convergence of the extreme-value component and standardized $L$-statistics under mild conditions, yielding an asymptotic independence that justifies combining multiple $k$'s. An adaptive omnibus test is constructed via a Cauchy combination over a dyadic grid of $k$, and a wild bootstrap calibration is provided with theoretical guarantees. Simulations demonstrate accurate size and strong power across sparse and dense alternatives, including non-Gaussian designs.

stat.AP

Change-Points Detection and Support Recovery for Spatially Indexed Functional Data

Large volumes of spatiotemporal data, characterized by high spatial and temporal variability, may experience structural changes over time. Unlike traditional change-point problems, each sequence in this context consists of function-valued curves observed at multiple spatial locations, with typically only a small subset of locations affected. This paper addresses two key issues: detecting the global change-point and identifying the spatial support set, within a unified framework tailored to spatially indexed functional data. By leveraging a weakly separable cross-covariance structure -- an extension beyond the restrictive assumption of space-time separability -- we incorporate functional principal component analysis into the change-detection methodology, while preserving common temporal features across locations. A kernel-based test statistic is further developed to integrate spatial clustering pattern into the detection process, and its local variant, combined with the estimated change-point, is employed to identify the subset of locations contributing to the mean shifts. To control the false discovery rate in multiple testing, we introduce a functional symmetrized data aggregation approach that does not rely on pointwise p-values and effectively pools spatial information. We establish the asymptotic validity of the proposed change detection and support recovery method under mild regularity conditions. The efficacy of our approach is demonstrated through simulations, with its practical usefulness illustrated in an application to China's precipitation data.

stat.ME

AutoAssign+: Automatic Shared Embedding Assignment in Streaming Recommendation

In the domain of streaming recommender systems, conventional methods for addressing new user IDs or item IDs typically involve assigning initial ID embeddings randomly. However, this practice results in two practical challenges: (i) Items or users with limited interactive data may yield suboptimal prediction performance. (ii) Embedding new IDs or low-frequency IDs necessitates consistently expanding the embedding table, leading to unnecessary memory consumption. In light of these concerns, we introduce a reinforcement learning-driven framework, namely AutoAssign+, that facilitates Automatic Shared Embedding Assignment Plus. To be specific, AutoAssign+ utilizes an Identity Agent as an actor network, which plays a dual role: (i) Representing low-frequency IDs field-wise with a small set of shared embeddings to enhance the embedding initialization, and (ii) Dynamically determining which ID features should be retained or eliminated in the embedding table. The policy of the agent is optimized with the guidance of a critic network. To evaluate the effectiveness of our approach, we perform extensive experiments on three commonly used benchmark datasets. Our experiment results demonstrate that AutoAssign+ is capable of significantly enhancing recommendation performance by mitigating the cold-start problem. Furthermore, our framework yields a reduction in memory usage of approximately 20-30%, verifying its practical effectiveness and efficiency for streaming recommender systems.

cs.IR

Rank Based Tests for High Dimensional White Noise

The development of high-dimensional white noise test is important in both statistical theories and applications, where the dimension of the time series can be comparable to or exceed the length of the time series. This paper proposes several distribution-free tests using the rank based statistics for testing the high-dimensional white noise, which are robust to the heavy tails and do not quire the finite-order moment assumptions for the sample distributions. Three families of rank based tests are analyzed in this paper, including the simple linear rank statistics, non-degenerate U-statistics and degenerate U-statistics. The asymptotic null distributions and rate optimality are established for each family of these tests. Among these tests, the test based on degenerate U-statistics can also detect the non-linear and non-monotone relationships in the autocorrelations. Moreover, this is the first result on the asymptotic distributions of rank correlation statistics which allowing for the cross-sectional dependence in high dimensional data.

math.ST

Stabilizing Training of Generative Adversarial Nets via Langevin Stein Variational Gradient Descent

Generative adversarial networks (GANs), famous for the capability of learning complex underlying data distribution, are however known to be tricky in the training process, which would probably result in mode collapse or performance deterioration. Current approaches of dealing with GANs' issues almost utilize some practical training techniques for the purpose of regularization, which on the other hand undermines the convergence and theoretical soundness of GAN. In this paper, we propose to stabilize GAN training via a novel particle-based variational inference -- Langevin Stein variational gradient descent (LSVGD), which not only inherits the flexibility and efficiency of original SVGD but aims to address its instability issues by incorporating an extra disturbance into the update dynamics. We further demonstrate that by properly adjusting the noise variance, LSVGD simulates a Langevin process whose stationary distribution is exactly the target distribution. We also show that LSVGD dynamics has an implicit regularization which is able to enhance particles' spread-out and diversity. At last we present an efficient way of applying particle-based variational inference on a general GAN training procedure no matter what loss function is adopted. Experimental results on one synthetic dataset and three popular benchmark datasets -- Cifar-10, Tiny-ImageNet and CelebA validate that LSVGD can remarkably improve the performance and stability of various GAN models.

cs.LG