SearcharxivSearch

arXiv subjects

Ze-Yu Li

Publications and source records attributed to Ze-Yu Li.

3 recordsLinked to original sources

On Explicit Super-Expressive Approximation for Neural Networks

In this work, we investigate the fixed-architecture neural network approximation with explicit parameter bounds and elementary activations. While prior work demonstrated super-expressive approximation using fixed-size networks, they lack quantitative and non-asymptotic characterizations of parameter magnitude with respect to the approximation error. We resolve this issue by introducing the Chinese Remainder Theorem as a constructive encoding mechanism. For Lipschitz continuous functions on $[0,1]^D$, we construct a width-$\max\{D,4\}$, depth-$5$ network with explicit parameter-error trade-offs. For H\"older-smooth functions in $C^{r,\gamma}_A\left([0,1]^D\right)$, our fixed network of width $\max\{2D,\ D+5N+1\}$ and depth $r + 9$ achieves the parameter magnitude $\mathcal{P}$ bounded by $\log_2 \mathcal{P}=\mathcal{O}\bigl(\varepsilon^{-2D/(r+\gamma)}\log(1/\varepsilon)\bigr)$. This is the dual result compared to those in the parameter-bounded and architecture-unbounded paradigm.

cs.LG

On Expressivity of Height in Neural Networks

In this work, beyond width and depth, we augment a neural network with a new dimension called height by intra-linking neurons in the same layer to create an intra-layer hierarchy, which gives rise to the notion of height. We call a neural network characterized by width, depth, and height a 3D network. To put a 3D network in perspective, we theoretically and empirically investigate the expressivity of height. We show via bound estimation and explicit construction that given the same number of neurons and parameters, a 3D ReLU network of width $W$, depth $K$, and height $H$ has greater expressive power than a 2D network of width $H\times W$ and depth $K$, \textit{i.e.}, $\mathcal{O}((2^H-1)W)^K)$ vs $\mathcal{O}((HW)^K)$, in terms of generating more pieces in a piecewise linear function. Next, through approximation rate analysis, we show that by introducing intra-layer links into networks, a ReLU network of width $\mathcal{O}(W)$ and depth $\mathcal{O}(K)$ can approximate polynomials in $[0,1]^d$ with error $\mathcal{O}\left(2^{-2WK}\right)$, which improves $\mathcal{O}\left(W^{-K}\right)$ and $\mathcal{O}\left(2^{-K}\right)$ for fixed width networks. Lastly, numerical experiments on 5 synthetic datasets, 15 tabular datasets, and 3 image benchmarks verify that 3D networks can deliver competitive regression and classification performance.

cs.LG

Three-dimensional potential energy surface for fission of $^{236}$U within covariant density functional theory

We have calculated the three-dimensional potential energy surface (PES) for the fission of compound nucleus $^{236}$U using the covariant density functional theory with constraints on the axial quadrupole and octupole deformations $(\beta_2, \beta_3)$ as well as the nucleon number in the neck $q_N$. By considering the additonal degree of freedom $q_N$, coexistence of the elongated and compact fission modes is predicted for $0.9\lesssim \beta_3 \lesssim 1.3$. Remarkably, the PES becomes very shallow across a large range of quadrupole and octupole deformations for small $q_N$, and consequently, the scission line in $(\beta_2, \beta_3)$ plane will extend to a shallow band, which leads to a fluctuation for the estimated total kinetic energies by several to ten MeV and for the fragment masses by several to about ten nucleons.

nucl-th