SearcharxivSearch

arXiv subjects

Xiaojun Chen

Publications and source records attributed to Xiaojun Chen.

At least 19 recordsLinked to original sources

A sequential regularized piecewise affine algorithm for nonconvex nonsmooth multicomposite optimization in RNN training

This paper focuses on a class of nonconvex nonsmooth multicomposite optimization problems for training RNNs (Recurrent Neural Networks). We first establish easily verifiable conditions under which an approximate first-order d-stationary point of the problem is guaranteed to be an approximate second-order d-stationary point. Subsequently, we propose a sequential regularized piecewise affine algorithm, SRPA, that minimizes a regularized piecewise affine function at each iteration, where the idea of active-set strategy is incorporated to reduce the computational cost per iteration. Leveraging the aforementioned conditions for second-order d-stationarity, we establish the global convergence of SRPA to second-order d-stationary points and the complexity bound of $ \mathcal{O}(ε^{-2}) $ for obtaining an $ ε$-approximate second-order d-stationary point. Finally, numerical experiments for training RNNs on synthetic and real-world datasets demonstrate the promising performance of SRPA compared with state-of-the-art algorithms.

math.OC

A sequential smoothing majorant stochastic approximation method for nonconvex nonconcave minimax problems

We propose a sequential smoothing majorant stochastic approximation (SMSA) method for nonconvex-nonconcave minimax optimization problems. To overcome the nonconcavity of the inner maximization problem, we introduce a new majorant stochastic approximation with an $O\left(β_N^2\right)$ accuracy bound, where $β_N$ is the sample coverage radius. The accuracy bound significantly improves upon the $O\left(β_N\right)$ approximation bound for the standard stochastic approximation. Moreover, we establish nonasymptotic bounds for both global optimal values and minimizer sets, and prove consistency for the Clarke stationary points, as $β_N \downarrow 0$ almost surely. We show that the generated sequence by the SMSA method is bounded, and that the returned point is an approximate Clarke stationary point of the majorant stochastic approximation models. Numerical experiments on a synthetic toy example and robust logistic regression on two UCI datasets demonstrate improved approximation fidelity and lower mean robust test losses relative to the standard sampled approximation.

math.OC

Differential Stochastic Variational Inequalities with Parametric Optimization

The differential stochastic variational inequality with parametric convex optimization (DSVI-O) is an ordinary differential equation whose right-hand side involves a stochastic variational inequality and solutions of several dynamic and random parametric convex optimization problems. We consider that the distribution of the random variable is time-dependent and assume that the involved functions are continuous and the expectation is well-defined. We show that the DSVI-O has a weak solution with integrable and measurable solutions of the parametric optimization problems. Moreover, we propose a discrete scheme of DSVI-O by using a time-stepping approximation and the sample average approximation and prove the convergence of the discrete scheme. We illustrate our theoretical results of DSVI-O with applications in an embodied intelligence system for the elderly health by synthetic health care data generated by Multimodal Large Language Models.

math.OC

An Adaptive Projected-Gradient Algorithm for Sample-Average Approximations of Stochastic Multi-Objective Optimization

We consider stochastic multi-objective optimization over a nonempty closed convex set, where every objective is an expectation and only sample-gradient information is available. We develop a line-search-free and function-value-free adaptive projected-gradient algorithm for the sample-average approximation (SAA) problem. Each iteration computes a feasible regularized multi-gradient step and updates the regularization parameter from the projected step length. A normal-cone-based certificate yields descent estimates and an explicit complexity bound for the Pareto-stationarity residual of the SAA problem. The consistency of SAA gradients then transfers vanishing SAA residuals to Pareto stationarity for the population problem, while an additional concentration argument gives a finite-sample residual bound on compact sets. Experiments on synthetic problems, classification, portfolio selection, multi-task learning, and robot control illustrate the practical performance of our algorithm.

math.OC

Biquantization of the necklace Lie bialgebra

For the double of a quiver, the works of Ginzburg, Bocklandt-Le Bruyn and Schedler show that its closed paths, called the necklaces, have a natural Lie bialgebra structure. Schedler also constructed,in [Int. Math. Res. Notices, 2005 (12), 725-760], a Hopf algebra that quantizes this Lie bialgebra. In this paper, we pursue one more step in this direction by constructing its biquantization, in the sense of Turaev [Ann. Sci. École Norm. Sup. (4) 24 (1991), no. 6, 635-704].

math.RA

Stability of Differential Stochastic Variational Inequalities with History-Dependent Responses and Transfer Learning

In this paper, we propose and study a class of differential stochastic variational inequalities (DSVIs), in which an ordinary differential equation (ODE) is coupled with history-dependent stochastic variational inequalities (SVI). This framework models closed-loop stochastic systems with time-varying random equilibria and includes optimization-constrained ODEs as special cases. Under appropriate technical conditions, we establish uniqueness, measurability, and Lipschitz continuity with respect to the state of the second-stage response, and consequently the existence and uniqueness of the induced state trajectory. Moreover, we construct a sample average approximation (SAA) based on independent sample paths and prove uniform convergence of the approximate trajectories. For transfer between related stochastic environments, we derive a local $1/2$-Hölder estimate for parametric variational inequalities with moving feasible sets and a quantitative trajectory-stability bound in terms of the initial-state difference and the Wasserstein distance between exogenous path laws. Numerical experiments illustrate the SAA convergence and transfer-stability results. We further apply the framework to an elderly-health monitoring system. Similarity-weighted reuse of precomputed responses achieves an accuracy close to the full-recomputation benchmark of 0.97, while reducing the online batch runtime from 86 seconds to less than one second. Perturbation and delayed-update experiments additionally characterize robustness to sensor noise and the trade-off between response freshness, predictive accuracy, and computational cost. These results provide theoretical and computational support for efficient transfer learning in history-dependent DSVI systems.

math.OC

A Support-Set Algorithm for Optimization Problems with Nonnegative and Orthogonal Constraints

In this paper, we investigate optimization problems with nonnegative and orthogonal constraints, where any feasible matrix of size $n \times p$ exhibits a sparsity pattern such that each row accommodates at most one nonzero entry. Our analysis demonstrates that, by fixing the support set, the global solution of the minimization subproblem for the proximal linearization of the objective function can be computed in closed form with at most $n$ nonzero entries. Exploiting this structural property offers a powerful avenue for dramatically enhancing computational efficiency. Guided by this insight, we propose a support-set algorithm preserving strictly the feasibility of iterates. A central ingredient is a strategically devised update scheme for support sets that adjusts the placement of nonzero entries. We establish the convergence of the support-set algorithm to a first-order stationary point, and show that its iteration complexity required to reach an $ε$-approximate first-order stationary point is $O (ε^{-2})$. Numerical results are strongly in favor of our algorithm in real-world applications, including nonnegative PCA, clustering, and community detection.

math.OC

EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spider, BIRD, and Spider~2.0 evaluate schema generalization, large-scale databases, and realistic workflows, but largely overlook enterprise scenarios where SQL generation depends on private business knowledge, such as internal metrics, reporting conventions, and organizational rules. We introduce EntSQL, an enterprise-oriented Text-to-SQL benchmark for evaluating long-context grounding over proprietary business documents. EntSQL contains 1,066 aligned Chinese-English semantic examples across five business domains, with most examples requiring domain knowledge beyond the question and schema and involving complex SQL structures. On English inputs, the best evaluated system reaches only 15.9\% when long-form documents are provided, highlighting the difficulty of grounding SQL generation in enterprise knowledge.

cs.CL

Robust Least Squares Problems with Binary Uncertain Data

We propose a Binary Robust Least Squares (BRLS) model that encompasses key robust least squares formulations, such as those involving uncertain binary labels and adversarial noise constrained within a hypercube. {To develop algorithms with theoretical guarantees for the BRLS problem, we exploit the structure of the inner binary maximization problem with a convex quadratic objective function. Refined guarantees are obtained when the noise correlations are sign-structured, in which case the inner problem admits sharper submodular or supermodular oracles. For the supermodular linear BRLS problem, we establish a link between saddle points of its continuous relaxation and global minimax points of BRLS, and propose a projected-gradient algorithm to find an $ε$-global minimax point in $O(ε^{-2})$ iterations. For the supermodular nonlinear BRLS problem, we develop a Moreau-envelope-based framework that finds an $ε$-stationary point in expectation within $O(ε^{-4})$ iterations. For the linear submodular case and the linear general case, we utilize a double-greedy algorithm and a semidefinite relaxation as the respective subsolvers; the latter attains an approximation ratio below $2/π$. Coupled with the projected-gradient framework, these oracles yield approximate minimax guarantees within $O(ε^{-2})$ iterations. Numerical experiments on health status prediction with candidate label-corruption sets, synthetic linear BRLS, and thresholded phase retrieval with missing binary labels illustrate the behavior and robustness gains of the BRLS model under structured noise compared with classical least-squares-based baselines.

math.OC

Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing

Recent advancements in diffusion-based image editing pose a significant threat to the authenticity of digital visual content. Traditional embedding-based watermarking methods often introduce perceptible perturbations to maintain robustness, inevitably compromising visual fidelity. Meanwhile, existing zero-watermarking approaches, typically relying on global image features, struggle to withstand sophisticated manipulations. In this work, we uncover a key observation: while individual image patches undergo substantial alterations during AI-based editing, the relational distance between patch pairs remains relatively invariant. Leveraging this property, we propose Relational Zero-Watermarking (Rel-Zero), a novel framework that requires no modification to the original image but derives a unique zero-watermark from these editing-invariant patch relations. By grounding the watermark in intrinsic structural consistency rather than absolute appearance, Rel-Zero provides a non-invasive yet resilient mechanism for content authentication. Extensive experiments demonstrate that Rel-Zero achieves substantially improved robustness across diverse editing models and manipulations compared to prior zero-watermarking approaches.

cs.CV

An extra gradient Anderson-accelerated algorithm for pseudomonotone variational inequalities

This paper proposes an extra gradient Anderson-accelerated algorithm for solving pseudomonotone variational inequalities, which uses the extra gradient scheme with line search to guarantee the global convergence and Anderson acceleration to have fast convergent rate. We prove that the sequence generated by the proposed algorithm from any initial point converges to a solution of the pseudomonotone variational inequality problem without assuming the Lipschitz continuity and contractive condition, which are used for convergence analysis of the extra gradient method and Anderson-accelerated method, respectively in existing literatures. Numerical experiments, particular emphasis on Harker-Pang problem, fractional programming, nonlinear complementarity problem and PDE problem with free boundary, are conducted to validate the effectiveness and good performance of the proposed algorithm comparing with the extra gradient method and Anderson-accelerated method.

math.OC

An Anderson-accelerated stochastic extragradient method for stochastic variational inequalities

In this paper, we propose an Anderson-accelerated stochastic extragradient algorithm for solving a class of stochastic variational inequalities, by incorporating Anderson acceleration into the stochastic extragradient method under a stochastic approximation framework. A key challenge in our setting is that the pseudomonotonicity assumption is only imposed on the expectation of the stochastic operator, rather than on the individual stochastic operator itself and the sample averages utilized in the algorithm. We prove that, despite the lack of pseudomonotonicity in the sampled operators, the sequence generated by the proposed algorithm converges almost surely to a solution of the stochastic variational inequality problem. Additionally, we establish the sublinear convergence rate of the proposed algorithm in terms of the mean residual function, along with its optimal oracle complexity. Finally, we validate the effectiveness of the proposed algorithm through numerical experiments.

math.OC

Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures

Existing ViT backdoor attacks based on backbone-overwriting full-tuning are computationally expensive and inflict performance degradation. This has forced adversaries towards the Visual Parameter-Efficient Fine-Tuning (PEFT) paradigm, dominated by adapter-based (e.g., LoRA) and prompt-based (e.g., VPT) approaches. While adapter security has seen initial study, the risks of the burgeoning prompt-based ecosystem remain critically unexplored. We fill this critical gap, exposing how the evolution of VPT towards dynamic and context-aware architectures can facilitate a far more dangerous and emergent threat. This vulnerability arises even though these dynamic modules unlock superior benign performance. We propose VIPER, an attack framework built on a lightweight, dynamic Visual Prompt Generator (VPG) that demonstrates this vulnerability. Critically, this dynamic architecture enables Functional Fusion: an emergent phenomenon where malicious logic and benign task utility are tightly fused into the same sparse, high-magnitude parameter core. This fusion creates a formidable ``hostage" dilemma, as pruning the attack necessarily destroys the benign performance. Comprehensive evaluations show VIPER effectively addresses the attacker's trilemma: VIPER not only achieves state-of-the-art performance on clean data, but also maintains near-100% ASR even under 90% VPG-module pruning (where LoRA attacks collapse), while adding only an imperceptible 0.06ms (1.16%) of inference latency. VIPER's results, driven by Functional Fusion, expose a new, paradigm-level risk in dynamic prompt architectures.

cs.CR

Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model

Text-to-SQL converts natural language questions into executable SQL queries, enabling non-technical users to access relational databases for analytics and intelligent data services. In real-world scenarios, performance is often constrained by low-resource settings, where high-quality annotated \texttt{ } pairs are scarce, particularly for domain-specific databases. Additional challenges include opaque schema definitions, abbreviations, and implicit business logic that are not explicitly encoded in the schema. Existing data synthesis and prompting techniques improve coverage but often fail to produce task-specific, semantically grounded examples aligned with database constraints. To address these challenges, we propose a knowledge-aware Text-to-SQL framework that constructs task-specific knowledge base including schema semantics, abbreviations, business logic, and query patterns, and injects them into both training and inference. This framework generates diverse, contextually grounded synthetic training data and enhances inference through targeted knowledge retrieval. Experiments on seven benchmarks, covering both general and domain-specific datasets, demonstrate that our approach substantially improves the performance of open-source and closed-source large language models in Text-to-SQL tasks, especially in low-resource domain-specific settings, enhancing generalization, robustness, and adaptability.

cs.CL

Algebraic $K$-theory, cohomotopy $K$-groups, and Koszul duality

Let $A$ be an augmented differential graded algebra over a field $k$ of characteristic zero, and let $A^!=\mathbf{R}\mathrm{Hom}_A(k,k)$ be its Koszul dual algebra. Blumberg and Mandell showed that, under some finiteness conditions of $A$, the derived Koszul duality provides an equivalence between the $K$-theory $K(\mathrm{thick}_A(k))$ of the triangulated thick subcategory generated by $k$ and the $K$-theory $K(A^!)$ of the derived category of perfect $A^!$-modules. Combining this equivalence with the Jones-Goodwillie Chern character and the Jones-McCleary isomorphism, we obtain that the $K$-groups $K_n(\mathrm{thick}_A(k))$ are a concrete candidate for Loday's conjectural contravariant $K$-groups.

math.KT

ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples

The rapid advancement of deep learning has turned models into highly valuable assets due to their reliance on massive data and costly training processes. However, these models are increasingly vulnerable to leakage and theft, highlighting the critical need for robust intellectual property protection. Model watermarking has emerged as an effective solution, with black-box watermarking gaining significant attention for its practicality and flexibility. Nonetheless, existing black-box methods often fail to better balance covertness (hiding the watermark to prevent detection and forgery) and robustness (ensuring the watermark resists removal)-two essential properties for real-world copyright verification. In this paper, we propose ComMark, a novel black-box model watermarking framework that leverages frequency-domain transformations to generate compressed, covert, and attack-resistant watermark samples by filtering out high-frequency information. To further enhance watermark robustness, our method incorporates simulated attack scenarios and a similarity loss during training. Comprehensive evaluations across diverse datasets and architectures demonstrate that ComMark achieves state-of-the-art performance in both covertness and robustness. Furthermore, we extend its applicability beyond image recognition to tasks including speech recognition, sentiment analysis, image generation, image captioning, and video recognition, underscoring its versatility and broad applicability.

cs.CR

High-Throughput and Scalable Secure Inference Protocols for Deep Learning with Packed Secret Sharing

Most existing secure neural network inference protocols based on secure multi-party computation (MPC) typically support at most four participants, demonstrating severely limited scalability. Liu et al. (USENIX Security'24) presented the first relatively practical approach by utilizing Shamir secret sharing with Mersenne prime fields. However, when processing deeper neural networks such as VGG16, their protocols incur substantial communication overhead, resulting in particularly significant latency in wide-area network (WAN) environments. In this paper, we propose a high-throughput and scalable MPC protocol for neural network inference against semi-honest adversaries in the honest-majority setting. The core of our approach lies in leveraging packed Shamir secret sharing (PSS) to enable parallel computation and reduce communication complexity. The main contributions are three-fold: i) We present a communication-efficient protocol for vector-matrix multiplication, based on our newly defined notion of vector-matrix multiplication-friendly random share tuples. ii) We design the filter packing approach that enables parallel convolution. iii) We further extend all non-linear protocols based on Shamir secret sharing to the PSS-based protocols for achieving parallel non-linear operations. Extensive experiments across various datasets and neural networks demonstrate the superiority of our approach in WAN. Compared to Liu et al. (USENIX Security'24), our scheme reduces the communication upto 5.85x, 11.17x, and 6.83x in offline, online and total communication overhead, respectively. In addition, our scheme is upto 1.59x, 2.61x, and 1.75x faster in offline, online and total running time, respectively.

cs.CR

An Adaptive Smoothing Algorithm for Non-Lipschitz Optimization on Manifolds with Complexity Guarantees

We study a class of optimization problems on Riemannian manifolds, where the objective function consists of a smooth term and quasi-norm type penalties with exponent $p \in (0, 1]$. The essential difficulty lies in the fact that the objective function may not be locally Lipschitz continuous, which places this type of problems beyond the reach of existing Riemannian techniques. To overcome this obstacle, this paper constructs a general smoothing framework and establishes fundamental properties for developing efficient algorithms. In particular, we propose a smoothing Riemannian gradient algorithm equipped with a smoothing-aware AdaGrad-type stepsize rule. Its global convergence is demonstrated together with an iteration complexity of $O (ε^{p - 4})$, which includes the best available iteration complexity of $O (ε^{- 3})$ for Lipschitz problems with $p = 1$ as a special case. To the best of our knowledge, this is the first complexity result for non-Lipschitz optimization on Riemannian manifolds. Preliminary numerical experiments corroborate the practical efficiency of the proposed approach in real-world applications arsing from machine learning and data science.

math.OC