SearcharxivSearch

arXiv subjects

Yiling Xie

Publications and source records attributed to Yiling Xie.

14 recordsLinked to original sources

The Noise Premium in Adversarial Training for Kernel Regression

Adversarial training can improve the robustness of predictive models to bounded perturbations, often at the cost of statistical efficiency. We study this trade-off in kernel regression over a reproducing kernel Hilbert space (RKHS). It is shown that, under squared loss, adversarial training in RKHS introduces a term involving the product of the function norm with the mean absolute value of the response noise, which we call the \textit{noise premium}. Our analysis shows that the noise premium makes the prediction error of adversarial training converge strictly more slowly than the nonparametric minimax benchmark even after balancing approximation and estimation errors. Moreover, for a fixed perturbation budget, once the budget exceeds a certain threshold, the solution to adversarial training collapses to the zero function. To mitigate these effects of the noise premium, we propose noise-debiased adversarial training. The resulting noise-debiased estimator can attain the minimax optimal rate up to a logarithmic factor for the prediction error, raises the collapse threshold, and admits an explicit bound on the increase in adversarial loss. Numerical experiments on synthetic and real data support the theoretical findings and validate the effectiveness of the proposed noise-debiased method.

stat.ML

Utility-Aware Multimodal Contrastive Learning for Product Image Generation

Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can output images that closely align with text prompts. Yet existing generative AI models do not directly optimize marketplace performance. This is a critical gap, since semantic alignment alone does not guarantee that an image will sell. To address this limitation, we propose a \textit{utility-aware multimodal contrastive learning} framework that incorporates consumer demand into a novel Utility-Aware InfoNCE loss. Optimizing this utility-aware objective guides generation toward images that are both semantically coherent and demand-enhancing. This effect arises directly from a shift in the learned image-text representation space toward demand-driven visual cues, which we also validate through the theoretical bound of the proposed objective. In downstream applications on Amazon and Airbnb, product images generated and edited by our method outperform state-of-the-art models in increasing demand and preserving fidelity, while maintaining text-image consistency. Notably, our utility-aware framework preserves inverse U-shaped demand patterns for attributes such as aesthetics and uniqueness, improving demand-based performance while preserving fidelity and semantic consistency. Human-subject experiments further validate its commercial effectiveness. As generative AI technology continues to evolve, our utility-aware component can be flexibly embedded into emerging generative models to improve direct commercial use.

cs.AI

Adversarially Perturbed Precision Matrix Estimation

Precision matrix estimation is a fundamental topic in multivariate statistics and modern machine learning. This paper proposes an adversarially perturbed precision matrix estimation framework, motivated by recent developments in adversarial training. The proposed framework is versatile for the precision matrix problem since, by adapting to different perturbation geometries, the proposed framework can not only recover the existing distributionally robust method but also achieve high-dimensional model selection consistency under the scale-adaptive incoherence condition, which can be viewed as a relaxation of the classic incoherence condition in the heteroscedastic settings. Additionally, the proposed perturbed precision matrix estimation framework is asymptotically equivalent to the regularized precision matrix estimation, and the asymptotic normality can be established accordingly, where the asymptotic bias introduced by perturbation is highlighted. Numerical experiments demonstrate the desirable practical performance of the proposed adversarially perturbed approach.

stat.ME

High-dimensional (Group) Adversarial Training in Linear Regression

Adversarial training can achieve robustness against adversarial perturbations and has been widely used in machine learning models. This paper delivers a non-asymptotic consistency analysis of the adversarial training procedure under $\ell_\infty$-perturbation in high-dimensional linear regression. It will be shown that the associated convergence rate of prediction error can achieve the minimax rate up to a logarithmic factor in the high-dimensional linear regression on the class of sparse parameters. Additionally, the group adversarial training procedure is analyzed. Compared with classic adversarial training, it will be proved that the group adversarial training procedure enjoys a better prediction error upper bound under certain group-sparsity patterns.

math.ST

Asymptotic Behavior of Adversarial Training Estimator under $\ell_\infty$-Perturbation

Adversarial training has been proposed to protect machine learning models against adversarial attacks. This paper focuses on adversarial training under $\ell_\infty$-perturbation, which has recently attracted much research attention. The asymptotic behavior of the adversarial training estimator is investigated in the generalized linear model. The results imply that the asymptotic distribution of the adversarial training estimator under $\ell_\infty$-perturbation could put a positive probability mass at $0$ when the true parameter is $0$, providing a theoretical guarantee of the associated sparsity-recovery ability. Alternatively, a two-step procedure is proposed -- adaptive adversarial training, which could further improve the performance of adversarial training under $\ell_\infty$-perturbation. Specifically, the proposed procedure could achieve asymptotic variable-selection consistency and unbiasedness. Numerical experiments are conducted to show the sparsity-recovery ability of adversarial training under $\ell_\infty$-perturbation and to compare the empirical performance between classic adversarial training and adaptive adversarial training.

math.ST

The strong decay of $Y(4230)\to J/\psi f_0(980)$ in light cone sum rules

In this work, we assign the tetraquark state for $Y(4230)$ resonance, and investigate the mass and decay constant of $Y(4230)$ in the framework of SVZ sum rules through a different calculation technique. Then we calculate the strong coupling $g_{Y J/\psi f_0}$ by considering soft-meson approximation techniques, within the framework of light cone sum rules. And we use strong coupling $g_{Y J/\psi f_0}$ to obtain the width of the decay $Y(4230)\to J/\psi f_0(980)$. Our prediction for the mass is in agreement with the experimental measurement, and that for the decay width of $Y(4230)\to J/\psi f_0(980)$ is within the upper limit.

hep-ph

The role of triangle singularity in the decay process $D^0 \to \pi^+ \pi^- f_0(980),\ f_0 \to \pi^+ \pi^-$

We study the process $D^0 \to \pi^+ \pi^- f_0(980),\ f_0 \to \pi^+ \pi^-$ by introducing the triangle mechanism, in which $f_0(980)$ is considered to be dynamically generated from the meson-meson interaction. For the total contribution of this process, the contribution of the triangular loop formed by $K^{*} \bar{K} K$ particles could generate a triangular singularity of about 1418 MeV. We calculate the differential decay width of this process and show a narrow peak of about 980 MeV in the $\pi^+ \pi^-$ invariant mass distribution, which comes from $f_0$ decay. For the $M_{inv}(\pi f_0)$ invariant mass distribution, we obtain a finite peak at 1418 MeV, which is consistent with the triangle singularity.

hep-ph

Adjusted Wasserstein Distributionally Robust Estimator in Statistical Learning

We propose an adjusted Wasserstein distributionally robust estimator -- based on a nonlinear transformation of the Wasserstein distributionally robust (WDRO) estimator in statistical learning. The classic WDRO estimator is asymptotically biased, while our adjusted WDRO estimator is asymptotically unbiased, resulting in a smaller asymptotic mean squared error. Further, under certain conditions, our proposed adjustment technique provides a general principle to de-bias asymptotically biased estimators. Specifically, we will investigate how the adjusted WDRO estimator is developed in the generalized linear model, including logistic regression, linear regression, and Poisson regression. Numerical experiments demonstrate the favorable practical performance of the adjusted estimator over the classic one.

stat.ML

Investigate the strong coupling of $g_{X J/\psi\phi}$ in $X(4500) \to J/\psi \phi$ by using the three-point sum rules and the light-cone sum rules

We assign $X(4500)$ as a D-wave tetraquark state and study the decay of $X(4500)$ $\to$ $J/\psi \phi$. The mass and the decay constant of $X(4500)$ are calculated by using the SVZ sum rules. For the decay width of $X(4500)$ $\to$ $J/\psi \phi$, we present the calculation within the framework of both the three-point sum rules and the light-cone sum rules. The strong coupling $g_{X J/\psi \phi}$ is obtained by considering the soft-meson approximation when we use the light-cone sum rules calculation. Both calculations show that the decay of $X(4500)$ $\to$ $J/\psi\phi$ close to the total width of $X(4500)$ if we assign $X$(4500) as a D-wave tetraquark. In this paper, only the hidden-charm decay channel is considered. With these results and those from the open-charm decay channel, we are able to give a more rational conclusion when comparing with the total width of $X(4500)$. In the future, experiments will be more helpful in determining whether or not this structure of $X(4500)$ is appropriate.

hep-ph

Improved Rate of First Order Algorithms for Entropic Optimal Transport

This paper improves the state-of-the-art rate of a first-order algorithm for solving entropy regularized optimal transport. The resulting rate for approximating the optimal transport (OT) has been improved from $\widetilde{{O}}({n^{2.5}}/{\epsilon})$ to $\widetilde{{O}}({n^2}/{\epsilon})$, where $n$ is the problem size and $\epsilon$ is the accuracy level. In particular, we propose an accelerated primal-dual stochastic mirror descent algorithm with variance reduction. Such special design helps us improve the rate compared to other accelerated primal-dual algorithms. We further propose a batch version of our stochastic algorithm, which improves the computational performance through parallel computing. To compare, we prove that the computational complexity of the Stochastic Sinkhorn algorithm is $\widetilde{{O}}({n^2}/{\epsilon^2})$, which is slower than our accelerated primal-dual stochastic mirror algorithm. Experiments are done using synthetic and real data, and the results match our theoretical rates. Our algorithm may inspire more research to develop accelerated primal-dual algorithms that have rate $\widetilde{{O}}({n^2}/{\epsilon})$ for solving OT.

math.OC

Solving a Special Type of Optimal Transport Problem by a Modified Hungarian Algorithm

Computing the empirical Wasserstein distance in the Wasserstein-distance-based independence test is an optimal transport (OT) problem with a special structure. This observation inspires us to study a special type of OT problem and propose a modified Hungarian algorithm to solve it exactly. For the OT problem involving two marginals with $m$ and $n$ atoms ($m\geq n$), respectively, the computational complexity of the proposed algorithm is $O(m^2n)$. Computing the empirical Wasserstein distance in the independence test requires solving this special type of OT problem, where $m=n^2$. The associated computational complexity of the proposed algorithm is $O(n^5)$, while the order of applying the classic Hungarian algorithm is $O(n^6)$. In addition to the aforementioned special type of OT problem, it is shown that the modified Hungarian algorithm could be adopted to solve a wider range of OT problems. Broader applications of the proposed algorithm are discussed -- solving the one-to-many assignment problem and the many-to-many assignment problem. We conduct numerical experiments to validate our theoretical results. The experiment results demonstrate that the proposed modified Hungarian algorithm compares favorably with the Hungarian algorithm, the well-known Sinkhorn algorithm, and the network simplex algorithm.

math.OC

The strong coupling $g_{X J/\psi\phi}$ of $X(4700) \to J/\psi \phi$ in the light-cone sum rules

We assign the scalar tetraquark and the D-wave tetraquark state for $X(4700)$ and calculate the width of the decay $X(4700)$ $\to J/\psi \phi$ within the framework of light-cone sum rules. The strong coupling $g_{X J/\psi \phi}$ is obtained by considering the technique of soft-meson approximation. We also investigate the mass and the decay constant of $X(4700)$ in the framework of SVZ sum rules. Our prediction for the mass is in agreement with the experimental measurement, and that for the decay width of $X(4700)$ $\to J/\psi \phi$ support the possibility that $X(4700)$ could be a scalar tetraquark state if $X(4700)$ $\to J/\psi \phi$ is the predominant decay channel, or a D-wave tetraquark state if $X(4700)$ $\to J/\psi \phi$ is not the predominant one and there exist other decays.

hep-ph

An Accelerated Stochastic Variance-Reduced Algorithm for Entropic Wasserstein Barycenters

Fixed-support Wasserstein barycenters average probability distributions while accounting for the geometry of the support. We study the entropically regularized Wasserstein barycenter problem with a fixed regularization parameter and propose an accelerated stochastic variance-reduced primal-dual algorithm. The proposed algorithm uses a semi-dual finite-sum structure in which each stochastic gradient requires only one softmax over the barycenter support. The resulting finite-sum components have dimension-free smoothness bounds, which lead to a complexity result showing that the method improves the support-size dependence of deterministic accelerated gradient by a square-root factor while preserving accelerated dependence on the target accuracy. Experiments on synthetic data, DOTmark images, shape aggregation, and digit-averaging instances are consistent with the theoretical dependence on support size and accuracy and show lower arithmetic costs than the tested first-order baselines.

stat.ML

Triangle mechanism in the decay process $J/\psi \to K^- K^+ a_1(1260)$

The role of triangle mechanism in the decay process $J/\psi \to K^- K^+ a_1(1260)$ is probed. In this mechanism, a close-up resonance with mass $1823$ MeV and width $122$ MeV decays into $K^* \phi, K^* \to K \pi$ and then $K^* \bar{K}$ fuses into the $a_1(1260)$ resonance. We find that this mechanism leads to a triangle singularity around $M_{\rm inv}(K^- a_1(1260))\approx 1920$ MeV, where the axial-vector meson $a_1(1260)$ is considered as a dynamically generated resonance. With the help of the triangle mechanism we find sizable branching ratios $\text{Br}(J/\psi \to K^- K^+ a_1(1260),a_1 \to \pi \rho)=1.210 \times 10^{-5}$ and $\text{Br}(J/\psi \to K^- K^+ a_1(1260))=3.501 \times 10^{-5}$. Such a effect from triangle mechanism of the decay process could be investigated by such as BESIII, LHCb and Belle-II experiments. This potential investigation can help us obtain the information of the axial-vector meson $a_1(1260)$.

hep-ph