Searcharxiv⌕ Search

arXiv subjects

Yi-Shuai Niu

Publications and source records attributed to Yi-Shuai Niu.

At least 19 recordsLinked to original sources

The SMG-Yau-Yau Filter and Lossless Data Assimilation: Resolving Infinite Lie Algebras via Statistical Fiber Theory

Classical continuous-time non-linear filtering fails in generic non-linear state spaces due to an infinite Lie algebraic derivative explosion ($\dim(\mathcal{E}) = \infty$), leading to filter divergence. To resolve this four-decade crisis, we introduce the SMG-Yau-Yau filter and lossless data assimilation framework built on Statistical Fiber Theory over an Orlicz manifold $\mathcal{M}$. By equipping $\mathcal{M}$ with a Riemannian submersion and an Ehresmann connection, the unconstrained Duncan-Mortensen-Zakai score velocity field is orthogonally decomposed into Statistically Verifiable Directions ($\text{SVD}χ_f$) and Structural Internal Directions ($\text{SID}_f$). System non-linearities and unclosed Lie commutators are orthogonally quarantined in $\text{SID}_f$, protecting macroscopic base parameters from spatial derivative pollution while preserving total score variance energy. We unify the asymptotics through a Dual-Axis Collapse Mechanism: proving our model identically recovers classical Yau-Yau dynamics when $\dim(\mathcal{E}) < \infty$, while large-sample limits ($N \to \infty$) induce a thermodynamic quench that flattens infinite-dimensional geometry under $\dim(\mathcal{E}) = \infty$. Finally, we formulate the Active Acausal Tension (AAT) functional to monitor accumulated model misspecification online. Exceeding a topological capacity threshold triggers Gauge Symmetry Breaking (GSB), which dynamically expands base coordinates ($d \to d+1$) to ensure non-asymptotic stability and convergence to the exact state density.

stat.ME↗

Information Geometry (IG) Lives at Edge or Boundary of SMG (statistically meaningful geometry): - the First Edge Theorem and Applications

Statistically Meaningful Geometry (SMG) is a differential-geometric and information-theoretic framework that lifts over-parameterized models into infinite-dimensional non-parametric Orlicz statistical fiber bundles with an Ehresmann connection, decoupling unobservable vertical gauge noise from horizontal statistically verifiable directions. We prove the First Edge Theorem: Amari's information geometry (IG) and conventional statistics (CS) are not autonomous statistical universes but degenerate boundary layers of the larger gauge-active SMG space. Taking the structural identifiability radius $R\to\infty$ breaks gauge symmetry, collapses vertical fibers, and forces the total space onto the IG manifold; subsequent local asymptotic normality as $N\to\infty$ flattens the remaining curvature, yielding \[ \mathrm{SMG}\xrightarrow{\;R\to\infty\;}\mathrm{IG}\xrightarrow{\;N\to\infty\;}\mathrm{CS}. \] The same fiber-bundle machinery transforms model non-identifiability from a singular collapse into a structured gauge space. Applications include resolving the deep-learning generalization paradox, constructing gauge-invariant gradient descent and holonomy-matched preference alignment for generative AI, and solving weak identification in structural econometrics via intrinsic horizontal geodesic search.

math.GM↗

The Second Edge Theorem: The Asymptotic Collapse of Sample-Dependent Information Geometry to the Canonical Flat Canvas of Conventional Statistics in Large Sample Limits

This paper establishes the global proof of the Second Edge Theorem: as sample size tends to infinity, sample-dependent information-geometric manifolds---formed by the parameter space, sample-scaled Fisher metric, and dual alpha-connections---undergo metric-topological collapse onto the flat tangent space of Conventional Statistics at the true parameter. We first prove a universal tensor valence scaling law, under which tensor fields of valence one through four degenerate at rates determined by sample size. Score fluctuations stabilize, the Fisher metric freezes to its true-parameter value, affine connections dissolve, and Riemann curvature is annihilated. We then bridge geometric collapse with statistical decision theory, showing that Cheeger--Gromov flattening and Le Cam risk condensation are dual projections of the same asymptotic phase transition. A Fisher-compatible Ehresmann connection extends these results to over-parameterized and singular models, yielding horizontal leaf-space collapse and uniform local asymptotic normality. Unifying the First and Second Edge Theorems yields a nested dual-edge hierarchy: Conventional Statistics is the boundary of Information Geometry, which is itself the boundary of Statistical Mechanics and Geometry. Thus, Conventional Statistics is not a heuristic approximation but the unique zero-curvature thermodynamic attractor of regular parametric information manifolds. This redefines modern statistics as a dynamic non-equilibrium field theory of finite-sample fluctuations, phase transitions, and gauge-invariant interactions.

math.GM↗

On the Convergence Analysis of DCA

Difference-of-Convex (DC) programming, which seeks to minimize a function expressed as the difference of two convex functions, arises in a wide range of applications in machine learning, signal processing, and operations research. A classical and widely used algorithm for solving DC programs is the Difference-of-Convex Algorithm (DCA). In this paper, we revisit DCA from a distinctly DC-specific perspective. We first separate well-definedness from asymptotic convergence and introduce an additional assumption ensuring the solvability of the DCA subproblems, which clarifies why the choice of DC decomposition matters. We then develop a Lyapunov-descent-regularity framework in which the descent estimate is read directly from the convex subproblems and the regularity estimate is verified from DCA optimality conditions. This yields global convergence of the iterates $\{x^k\}$ for both standard and convex-constrained DC programs under either the classical Lojasiewicz subgradient inequality or the broader Kurdyka-Lojasiewicz (KL) property. We further explain how stronger regularity regimes, such as the Polyak-Lojasiewicz (PL) condition, fit into the same framework and sharpen the resulting convergence rates. Consequently, we obtain finite-time, linear, and sublinear rates for objective values and iterates in a way that cleanly separates well-definedness, DCA-specific structure, and KL/PL regularity, and that is readily transferable to DCA-type variants.

math.OC↗

Statistically Meaningful Geometry (SMG) Beyond the Euclidean Paradigm, with Application to Generative AI

Conventional uniform convergence bounds and empirical risk minimization break down in massive over-parameterized models, such as large language transformers and biological sequence networks. With near-infinite unconstrained internal degrees of freedom, their optimization landscapes develop flat vertical gauge valleys, rendering classical generalization metrics vacuous and inducing severe pathologies, specifically generative hallucination and catastrophic forgetting. We introduce the Statistically Meaningful Geometry (SMG) framework, an information-geometric paradigm lifting deterministic parametric models into infinite-dimensional non-parametric Orlicz statistical manifolds. Modeling the total state space as a differential fiber bundle ($\mathcal{M}, \mathcal{B}, π, \mathcal{V}, \mathcal{H}, ω$), we establish a Two-Fold Inference Paradigm. We formalize an Ehresmann connection 1-form $ω$ as a dynamic geometric filter that strips away vertical gauge noise (Structural Internal Directions, or SID) and isolates learning trajectories along the strictly non-degenerate horizontal distribution (Statistical Variational Directions, or SVD$χ$). We prove that under connection-filtered pre-training, out-of-distribution predictive variance is strictly upper-bounded by the finite diameter of the identifiable quotient base manifold $\mathcal{B}$, establishing a hard geometric containment of generative hallucinations. By projecting downstream updates onto the orthogonal complement of the historical horizontal carriage, we formalize the SMG Sequential Adaptation Flow, proving the total non-asymptotic elimination of catastrophic forgetting. SMG replaces empirical fine-tuning heuristics with coordinate-free topological constraints, bridging advanced differential geometry with structural reliability in AI.

cs.LG↗

Statistically Meaningful Geometry and Gauge Symmetry Breaking: A Geometric Foundation for Scientific Discovery and Intelligence Emergence

The rapid scaling of over-parameterized machine learning architectures, particularly LLMs, raises a profound crisis: do these systems exhibit genuine intelligence, or are they merely sophisticated statistical pattern matchers? Classical flat Euclidean statistics cannot differentiate continuous interpolation from the autonomous discovery of novel causal laws. To resolve this, we introduce Statistically Meaningful Geometry (SMG), a framework modeling over-parameterized learning systems as infinite-dimensional non-parametric Orlicz fiber bundles. We prove that under persistent out-of-distribution (OOD) stimuli governed by unmodeled causal mechanisms, continuous optimization fails. Unmodeled variance is rejected by the visible horizontal base manifold, leaking into the unobservable vertical fiber space and generating an accumulation of Active Acausal Tension. Driven by the statistical manifold's non-linear curvature, this tension inevitably strikes a conjugate focal boundary ($T_{\text{crit}} = π^2 / K_{\text{max}}$), triggering localized volumetric collapse and a catastrophic matrix singularity ($[G_f]^{-1} \to \infty$). We demonstrate this geometric breakdown acts as the strict non-equilibrium trigger for a Gauge Symmetry Break (GSB). The system purges hidden tension from unobservable gauge redundancies, spontaneously crystallizing a new, mathematically independent horizontal coordinate axis. This non-parametric phase transition registers as a discrete $+1.0$ integer step-jump in observable Structural G-Entropy. By decoupling parameter charts and subjecting emergent axes to a Minimal Energy Path Criterion and a Causal Invariance Filter, we distinguish genuine discovery from malignant hallucinations. Ultimately, SMG provides a parameter-free, falsifiable dashboard to mathematically certify true intelligence, transforming AI for Science into an engine of autonomous paradigm shifts.

cs.LG↗

RA-DCA: A Randomized Active-Set DCA for Directional Stationarity in Max-Structured DC Programs

We study nonsmooth difference-of-convex programs whose subtracted convex term is a finite maximum of smooth convex functions. In this setting, standard DCA iterations may converge to critical points that are not directionally stationary, whereas exact active-vertex screening can be expensive when active sets are large or combinatorial. We propose RA-DCA, a vertex-first randomized active-set DCA that projects active gradients onto sampled directions, checks a sampled vertex residual, and uses a small linear program only as a low-residual convex-combination fallback. The method preserves the descent structure of DCA and reduces the randomized screening layer to matrix multiplications. Under the stated regularity, numerical active-set consistency, and random-embedding assumptions, every accumulation point generated by the safeguarded method is directionally stationary with probability one. MATLAB experiments first test the theorem on degenerate max-affine, max-quadratic, and sparse support-function models, where the safeguard avoids nonstationary critical points and closely tracks a full active-vertex scan. Block top-k tests then show that the same screening idea remains useful when exact aggregate enumeration is combinatorial. Trimmed-regression, complementarity, and QUBO diagnostics separate cases where active-set selection helps from cases dominated by multistart search, the DC split, or other problem-specific features.

math.OC↗

Yau's Affine-Normal Descent for Large-Scale Unrestricted Higher-Moment Portfolio Optimization

Unrestricted mean-variance-skewness-kurtosis portfolio optimization can capture asymmetry and tail risk, but sample-moment formulations become computationally impractical when the asset universe is large: they produce dense nonconvex quartic objectives with prohibitive coskewness and cokurtosis tensors and anisotropic, ill-conditioned level sets. We develop a structure-exploiting algorithm based on Yau's affine-normal descent that follows affine-normal directions of the current level set while working directly with the return matrix. The method avoids explicit higher-order tensors and exploits the quartic structure for exact sample oracles, derivative evaluation, and exact line search. We also provide theory for the reduced simplex formulation, including regularity and convexity conditions that separate data-map geometry from investor preference coefficients. Computational results show a clear implementation split: a direct configuration is effective on the standard small benchmark, whereas a preconditioned conjugate-gradient configuration with stall recovery becomes the preferred large-scale implementation by the upper end of the hundreds and remains competitive as the asset universe moves into the thousands. On a 5-minute A-share panel with 5,440 stocks, the method makes direct full-universe comparisons with exact mean-variance portfolios feasible and shows on the baseline split that the incremental value of higher moments is strongest at moderate return targets.

q-fin.PM↗

Polylab: A MATLAB Toolbox for Multivariate Polynomial Modeling

Polylab is a MATLAB toolbox for multivariate polynomial scalars and polynomial matrices with a unified symbolic-numeric interface across CPU and GPU-oriented backends. The software exposes three aligned classes: MPOLY for CPU execution, MPOLY_GPU as a legacy GPU baseline, and MPOLY_HP as an improved GPU-oriented implementation. Across these backends, Polylab supports polynomial construction, algebraic manipulation, simplification, matrix operations, differentiation, Jacobian and Hessian construction, LaTeX export, CPU-side LaTeX reconstruction, backend conversion, and interoperability with YALMIP and SOSTOOLS. Versions 3.0 and 3.1 add two practically important extensions: explicit variable identity and naming for safe mixed-variable expression handling, and affine-normal direction computation via automatic differentiation, MF-logDet-Exact, and MF-logDet-Stochastic. The toolbox has already been used successfully in prior research applications, and Polylab Version 3.1 adds a new geometry-oriented computational layer on top of a mature polynomial modeling core. This article documents the architecture and user-facing interface of the software, organizes its functionality by workflow, presents representative MATLAB sessions with actual outputs, and reports reproducible benchmarks. The results show that MPOLY is the right default for lightweight interactive workloads, whereas MPOLY-HP becomes advantageous for reduction-heavy simplification and medium-to-large affine-normal computation; the stochastic log-determinant variant becomes attractive in larger sparse regimes under approximation-oriented parameter choices.

cs.MS↗

Continuous-Time Dynamics of the Difference-of-Convex Algorithm

We study the continuous-time structure of the difference-of-convex algorithm (DCA) for smooth DC decompositions with a strongly convex component. In dual coordinates, classical DCA is exactly the full-step explicit Euler discretization of a nonlinear autonomous system. This viewpoint motivates a damped DCA scheme, which is also a Bregman-regularized DCA variant, and whose vanishing-step limit yields a Hessian-Riemannian gradient flow generated by the convex part of the decomposition. For the damped scheme we prove monotone descent, asymptotic criticality, Kurdyka-Lojasiewicz convergence under boundedness, and a global linear rate under a metric DC-PL inequality. For the limiting flow we establish an exact energy identity, asymptotic criticality of bounded trajectories, explicit global rates under metric relative error bounds, finite-length and single-point convergence under a Kurdyka-Lojasiewicz hypothesis, and local exponential convergence near nondegenerate local minima. The analysis also reveals a global-local tradeoff: the half-relaxed scheme gives the best provable global guarantee in our framework, while the full-step scheme is locally fastest near a nondegenerate minimum. Finally, we show that different DC decompositions of the same objective induce different continuous dynamics through the metric generated by the convex component, providing a geometric criterion for decomposition quality and linking DCA with Bregman geometry.

math.OC↗

Yau's Affine Normal Descent: Algorithmic Framework and Convergence Analysis

We propose Yau's Affine Normal Descent (YAND), a geometric framework for smooth unconstrained optimization in which search directions are defined by the equi-affine normal of level-set hypersurfaces. The resulting directions are invariant under volume-preserving affine transformations and intrinsically adapt to anisotropic curvature. Using the analytic representation of the affine normal from affine differential geometry, we establish its equivalence with the classical slice-centroid construction under convexity. For strictly convex quadratic objectives, affine-normal directions are collinear with Newton directions, implying one-step convergence under exact line search. For general smooth (possibly nonconvex) objectives, we characterize precisely when affine-normal directions yield strict descent and develop a line-search-based YAND. We establish global convergence under standard smoothness assumptions, linear convergence under strong convexity and Polyak-Lojasiewicz conditions, and quadratic local convergence near nondegenerate minimizers. We further show that affine-normal directions are robust under affine scalings, remaining insensitive to arbitrarily ill-conditioned transformations. Numerical experiments illustrate the geometric behavior of the method and its robustness under strong anisotropic scaling.

math.OC↗

Scalable Mean-Variance Portfolio Optimization via Subspace Embeddings and GPU-Friendly Nesterov-Accelerated Projected Gradient

We develop a sketch-based factor reduction and a Nesterov-accelerated projected gradient algorithm (NPGA) with GPU acceleration, yielding a doubly accelerated solver for large-scale constrained mean-variance portfolio optimization. Starting from the sample covariance factor $L$, the method combines randomized subspace embedding, spectral truncation, and ridge stabilization to construct an effective factor $L_{eff}$. It then solves the resulting constrained problem with a structured projection computed by scalar dual search and GPU-friendly matrix-vector kernels, yielding one computational pipeline for the baseline, sketched, and Sketch-Truncate-Ridge (STR)-regularized models. We also establish approximation, conditioning, and stability guarantees for the sketching and STR models, including explicit $O(\varepsilon)$ bounds for the covariance approximation, the optimal value error, and the solution perturbation under $(\varepsilon,δ)$-subspace embeddings. Experiments on synthetic and real equity-return data show that the method preserves objective accuracy while reducing runtime substantially. On a 5440-asset real-data benchmark with 48374 training periods, NPGA-GPU solves the unreduced full model in 2.80 seconds versus 64.84 seconds for Gurobi, while the optimized compressed GPU variants remain in the low-single-digit-second regime. These results show that the full dense model is already practical on modern GPUs and that, after compression, the remaining bottleneck is projection rather than matrix-vector multiplication.

math.OC↗

Affine Normal Directions via Log-Determinant Geometry: Scalable Computation under Sparse Polynomial Structure

Affine normal directions provide intrinsic affine-invariant descent directions derived from the geometry of level sets. Their practical use, however, has long been hindered by the need to evaluate third-order derivatives and invert tangent Hessians, which becomes computationally prohibitive in high dimensions. In this paper, we show that affine normal computation admits an exact reduction to second-order structure: the classical third-order contraction term is precisely the gradient of the log-determinant of the tangent Hessian. This identity replaces explicit third-order tensor contraction by a matrix-free formulation based on tangent linear solves, Hessian-vector products, and log-determinant gradient evaluation. Building on this reduction, we develop exact and stochastic matrix-free procedures for affine normal evaluation. For sparse polynomial objectives, the algebraic closure of derivatives further yields efficient sparse kernels for gradients, Hessian-vector products, and directional third-order contractions, leading to scalable implementations whose cost is governed by the sparsity structure of the polynomial representation. We establish end-to-end complexity bounds showing near-linear scaling with respect to the relevant sparsity scale under fixed stochastic and Krylov budgets. Numerical experiments confirm that the proposed MF-LogDet formulation reproduces the original autodifferentiation-based affine normal direction to near machine precision, delivers substantial runtime improvements in moderate and high dimensions, and exhibits empirical near-linear scaling in both dimension and sparsity. These results provide a practical computational route for affine normal evaluation and reveal a new connection between affine differential geometry, log-determinant curvature, and large-scale structured optimization.

math.OC↗

An Improved Yau-Yau Algorithm for High Dimensional Nonlinear Filtering Problems

Nonlinear state estimation under noisy observations is rapidly intractable as system dimension increases. We introduce an improved Yau-Yau filtering framework that breaks the curse of dimensionality and extends real-time nonlinear filtering to systems with up to thousands of state dimensions, achieving high-accuracy estimates in just a few seconds with rigorous theoretical error guarantees. This new approach integrates quasi-Monte Carlo low-discrepancy sampling, a novel offline-online update, high-order multi-scale kernel approximations, fully log-domain likelihood computation, and a local resampling-restart mechanism, all realized with CPU/GPU-parallel computation. Theoretical analysis guarantees local truncation error \(O(Δt^2 + D^*(n))\) and global error \(O(Δt + D^*(n)/Δt)\), where \(Δt\) is the time step and \(D^*(n)\) the star-discrepancy. Numerical experiments, spanning large-scale nonlinear cubic sensors up to 1000 dimensions, highly nonlinear small-scale problems, and linear Gaussian benchmarks, demonstrate sub-quadratic runtime scaling, sub-linear error growth, and excellent performance that surpasses the extended and unscented Kalman filters (EKF, UKF) and the particle filter (PF) under strong nonlinearity, while matching or exceeding the optimal Kalman-Bucy filter in linear regimes. By breaking the curse of dimensionality, our method enables accurate, real-time, high-dimensional nonlinear filtering, opening broad opportunities for applications in science and engineering.

math.OC↗

Understand the Effectiveness of Shortcuts through the Lens of DCA

Difference-of-Convex Algorithm (DCA) is a well-known nonconvex optimization algorithm for minimizing a nonconvex function that can be expressed as the difference of two convex ones. Many famous existing optimization algorithms, such as SGD and proximal point methods, can be viewed as special DCAs with specific DC decompositions, making it a powerful framework for optimization. On the other hand, shortcuts are a key architectural feature in modern deep neural networks, facilitating both training and optimization. We showed that the shortcut neural network gradient can be obtained by applying DCA to vanilla neural networks, networks without shortcut connections. Therefore, from the perspective of DCA, we can better understand the effectiveness of networks with shortcuts. Moreover, we proposed a new architecture called NegNet that does not fit the previous interpretation but performs on par with ResNet and can be included in the DCA framework.

cs.LG↗

On Difference-of-SOS and Difference-of-Convex-SOS Decompositions for Polynomials

In this article, we are interested in developing polynomial decomposition techniques based on sums-of-squares (SOS), namely the difference-of-sums-of-squares (D-SOS) and the difference-of-convex-sums-of-squares (DC-SOS). In particular, the DC-SOS decomposition is very useful for difference-of-convex (DC) programming formulation of polynomial optimization problems. First, we introduce the cone of convex-sums-of-squares (CSOS) polynomials and discuss its relationship to the sums-of-squares (SOS) polynomials, the non-negative polynomials and the SOS-convex polynomials. Then, we propose the set of D-SOS and DC-SOS polynomials, and prove that any polynomial can be formulated as D-SOS and DC-SOS. The problem of finding D-SOS and DC-SOS decompositions can be formulated as a semi-definite program and solved for any desired precision in polynomial time using interior point methods. Some algebraic properties of CSOS, D-SOS and DC-SOS are established. Second, we focus on establishing several practical algorithms for exact D-SOS and DC-SOS polynomial decompositions without solving any SDP. The numerical performance of the proposed D-SOS and DC-SOS decomposition algorithms and their parallel versions, tested on a dataset of 1750 randomly generated polynomials, is reported.

math.OC↗

A Boosted-DCA with Power-Sum-DC Decomposition for Linearly Constrained Polynomial Programs

This paper proposes a novel Difference-of-Convex (DC) decomposition for polynomials using a power-sum representation, achieved by solving a sparse linear system. We introduce the Boosted DCA with Exact Line Search (BDCAe) for addressing linearly constrained polynomial programs within the DC framework. Notably, we demonstrate that the exact line search equates to determining the roots of a univariate polynomial in an interval, with coefficients being computed explicitly based on the power-sum DC decompositions. The subsequential convergence of BDCAe to critical points is proven, and its convergence rate under the Kurdyka-Lojasiewicz property is established. To efficiently tackle the convex subproblems, we integrate the Fast Dual Proximal Gradient (FDPG) method by exploiting the separable block structure of the power-sum DC decompositions. We validate our approach through numerical experiments on the Mean-Variance-Skewness-Kurtosis (MVSK) portfolio optimization model and box-constrained polynomial optimization problems. Comparative analysis of BDCAe against DCA, BDCA with Armijo line search, UDCA, and UBDCA with projective DC decomposition, alongside standard nonlinear optimization solvers FMINCON and FILTERSD, substantiates the efficiency of our proposed approach.

math.OC↗

Accelerated DC Algorithms for the Asymmetric Eigenvalue Complementarity Problem

We are interested in solving the Asymmetric Eigenvalue Complementarity Problem (AEiCP) by accelerated Difference-of-Convex (DC) algorithms. Two novel hybrid accelerated DCA: the Hybrid DCA with Line search and Inertial force (HDCA-LI) and the Hybrid DCA with Nesterov's extrapolation and Inertial force (HDCA-NI), are established. We proposed three DC programming formulations of AEiCP based on Difference-of-Convex-Sums-of-Squares (DC-SOS) decomposition techniques, and applied the classical DCA and 6 accelerated variants (BDCA with exact and inexact line search, ADCA, InDCA, HDCA-LI and HDCA-NI) to the three DC formulations. Numerical simulations of 7 DCA-type methods against state-of-the-art optimization solvers IPOPT, KNITRO and FILTERSD, are reported.

math.OC↗