SearcharxivSearch

arXiv subjects

Fuzhou Gong

Publications and source records attributed to Fuzhou Gong.

16 recordsLinked to original sources

Interpreting FCDNNs via RG on Exponential Family

We consider establishing the interpretability theory of deep learning through constructing a corresponding relationship between the renormalization group (RG) method in statistical physics and the training process of deep neural networks (DNNs). We have proved the constructed relationship using the one-dimensional Ising model as the input data. In this paper we generalize our results to the case of continuous input data, which is a necessary preparation for applying the corresponding framework to real-world data. To be representative, we consider a class of data distribution in the exponential family. We prove that when the parameters of fully connected (FC) DNNs achieve their optimal value after training, the characteristic parameters of the feature layer output of DNNs are equal to the fixed points of the characteristic parameters of input data under RG method for continuous fields. This conclusion shows that the training process of DNNs is equivalent to RG calculation on this kind of data and therefore the network can extract main features from the input data just like RG. Also, the equivalence further validates the correspondence framework we have established, providing an explanation for the outstanding performance of DNNs on real-world data.

stat.ML

Log-Sobolev Inequality for Wolff Dynamics and Application to the Condensation of Eigen Microstate in the 1D Ising Model

The Wolff dynamics is a non-local Markov chain widely used for simulating the Ising model due to its effectiveness in reducing critical slowing down compared to the Glauber dynamics. Despite extensive algorithmic and numerical studies, a rigorous probabilistic understanding remains limited. In this paper, we take a first step toward addressing this gap. For the one-dimensional (1D) Ising model, we first derive the transition probabilities of the Wolff dynamics and show that, at the critical point, it converges to the two fully aligned configurations and subsequently oscillates between them. This behavior is absent in the Glauber dynamics. Second, we establish a log-Sobolev inequality with an explicit constant for the Wolff dynamics in the entire subcritical regime and derive quantitative bounds on its ergodic averages. As a by-product, at infinite temperature, the obtained constant coincides with the classical log-Sobolev constant of the random walk on the hypercube. Finally, we apply these results to analyze the spectrum of the sample covariance matrix generated by the Wolff dynamics, which was used by Chen et al. to study condensation of eigen microstate. We prove that the spectral behavior agrees with their simulations in the 1D Ising model, thereby providing theoretical support for their findings.

math.PR

The log-Sobolev inequality and correlation functions for the renormalization of 1D Ising model

The renormalization group (RG) method is an important tool for studying critical phenomena. In this paper, we employ stochastic analysis techniques to investigate the stochastic partial differential equation (SPDE) derived by regularizing and continuousizing the discrete stochastic equation, which is a variant of stochastic quantization equation of the one dimensional (1D) Ising model. Firstly, we give the regularity estimates for the solution to SPDE. Secondly, we prove the Clark-Ocone-Haussmann formula and derive the log-Sobolev inequality up to the terminal time $T$, as well as obtain a priori form of the renormalization relation. Finally, we verify the correctness of the renormalization procedure based on the partition function, and prove that the two point correlation functions of SPDE on lattices converge to the two point correlation functions of the 1D Ising model at the stable fixed point of the RG transformation as $T\rightarrow +\infty$.

math.PR

Asymptotic stability for non-equicontinuous Markov semigroups

We formulate a new criterion of the asymptotic stability for some non-equicontinuous Markov semigroups, the so-called eventually continuous semigroups. In particular, we provide a non-equicontinuous Markov semigroup example with essential randomness, which is asymptotically stable.

math.PR

Ergodicity for eventually continuous Markov--Feller semigroups on Polish spaces

This paper investigates the ergodicity of Markov--Feller semigroups on Polish spaces, focusing on very weak regularity conditions, particularly the Cesàro eventual continuity. First, it is showed that the Cesàro average of such semigroups weakly converges to an ergodic measure when starting from its support. This leads to a characterization of the relationship between Cesàro eventual continuity, Cesàro e-property, and weak-* mean ergodicity. Next, serval criteria are provided for the existence and uniqueness of invariant measures via Cesàro eventual continuity and lower bound conditions, establishing an equivalence relation between weak-* mean ergodicity and a lower bound condition. Additionally, some refined properties of ergodic decomposition are derived. Finally, the results are applied to several non-trivial examples, including iterated function systems, Hopf's turbulence model with random forces, and Lorenz system with noisy perturbations, either with or without Cesàro eventual continuity.

math.PR

Interpreting Deep Learning by Establishing a Rigorous Corresponding Relationship with Renormalization Group

In this paper, we focus on the interpretability of deep neural network. Our work is motivated by the renormalization group (RG) in statistical mechanics. RG plays the role of a bridge connecting microscopical properties and macroscopic properties, the coarse graining procedure of it is quite similar with the calculation between layers in the forward propagation of the neural network algorithm. From this point of view we establish a rigorous corresponding relationship between the deep neural network (DNN) and RG. Concretely, we consider the most general fully connected network structure and real space RG of one dimensional Ising model. We prove that when the parameters of neural network achieve their optimal value, the limit of coupling constant of the output of neural network equals to the fixed point of the coupling constant in RG of one dimensional Ising model. This conclusion shows that the training process of neural network is equivalent to RG and therefore the network extract macroscopic feature from the input data just like RG.

cond-mat.dis-nn

The local Poincare inequality of stochastic dynamic and application to the Ising model

Inspired by the idea of stochastic quantization proposed by Parisi and Wu, we construct the transition probability matrix which plays a central role in the renormalization group through a stochastic differential equation. By establishing the discrete time stochastic dynamics, the renormalization procedure can be characterized from the perspective of probability. Hence, we will focus on the investigation of the infinite dimensional stochastic dynamic. From the stochastic point of view, the discrete time stochastic dynamic can induce a Markov chain. Via calculating the square field operator and the Bakry-Émery curvature for a class of two-points functions, the local Poincaré inequality is established, from which the estimate of correlation functions can also be obtained. Finally, under the condition of ergodicity, by choosing the couple relationship between the system parameter $K$ and the system time $T$ properly when $T\rightarrow +\infty$, the two-points correlation functions for limit system are also estimated.

math.PR

The Variable Volatility Elasticity Model from Commodity Markets

In this paper, we propose and study a novel continuous-time model, based on the well-known constant elasticity of variance (CEV) model, to describe the asset price process. The basic idea is that the volatility elasticity of the CEV model can not be treated as a constant from the perspective of stochastic analysis. To address this issue, we deduce the price process of assets from the perspective of volatility elasticity, propose the constant volatility elasticity (CVE) model, and further derive a more general variable volatility elasticity (VVE) model. Moreover, our model can describe the positive correlation between volatility and asset prices existing in the commodity markets, while CEV model can only describe the negative correlation. Through the empirical research on the financial market, many assets, especially commodities, often show this positive correlation phenomenon in some time periods, which shows that our model has strong practical application value. Finally, we provide the explicit pricing formula of European options based on our model. This formula has an elegant form convenient to calculate, which is similarly to the renowned Black-Scholes formula and of great significance to the research of derivatives market.

q-fin.MF

Null-free False Discovery Rate Control Using Decoy Permutations

The traditional approaches to false discovery rate (FDR) control in multiple hypothesis testing are usually based on the null distribution of a test statistic. However, all types of null distributions, including the theoretical, permutation-based and empirical ones, have some inherent drawbacks. For example, the theoretical null might fail because of improper assumptions on the sample distribution. Here, we propose a null distribution-free approach to FDR control for multiple hypothesis testing. This approach, named target-decoy procedure, simply builds on the ordering of tests by some statistic or score, the null distribution of which is not required to be known. Competitive decoy tests are constructed from permutations of original samples and are used to estimate the false target discoveries. We prove that this approach controls the FDR when the statistics are independent between different tests. Simulation demonstrates that it is more stable and powerful than two existing popular approaches. Evaluation is also made on a real dataset.

stat.ME

Widely distributed clusters of the constraint satisfaction problem model d-k-CSP

Relation between problem hardness and solution space structure is an important research aspect. Model d-k-CSP generates very hard instances when $r=1$ and $r$ is near 1, where $r$ represents normalized constraint density. We find that when $r$ is below and close to 1, the solution space contains many widely distributed well-separated small cluster-regions (a cluster-region is a union of some clusters), which should the reason that the generated instances are hard to solve.

cond-mat.dis-nn

Generate the corresponding Image from Text Description using Modified GAN-CLS Algorithm

Synthesizing images or texts automatically is a useful research area in the artificial intelligence nowadays. Generative adversarial networks (GANs), which are proposed by Goodfellow in 2014, make this task to be done more efficiently by using deep neural networks. We consider generating corresponding images from an input text description using a GAN. In this paper, we analyze the GAN-CLS algorithm, which is a kind of advanced method of GAN proposed by Scott Reed in 2016. First, we find the problem with this algorithm through inference. Then we correct the GAN-CLS algorithm according to the inference by modifying the objective function of the model. Finally, we do the experiments on the Oxford-102 dataset and the CUB dataset. As a result, our modified algorithm can generate images which are more plausible than the GAN-CLS algorithm in some cases. Also, some of the generated images match the input texts better.

cs.LG

A probabilistic proof of the fundamental gap conjecture via the coupling by reflection

Let $Ω\subset\mathbb{R}^n$ be a strictly convex domain with smooth boundary and diameter $D$. The fundamental gap conjecture claims that if $V:\barΩ\to\mathbb{R}$ is convex, then the spectral gap of the Schrödinger operator $-Δ+V$ with Dirichlet boundary condition is greater than $\frac{3π^2}{D^2}$. Using analytic methods, Andrews and Clutterbuck recently proved in [J. Amer. Math. Soc. 24 (2011), no. 3, 899--916] a more general spectral gap comparison theorem which implies this conjecture. In the first part of the current work, we shall give an independent probabilistic proof of their result via the coupling by reflection of the diffusion processes. Moreover, we also present in the second part a simpler probabilistic proof of the original conjecture.

math.PR

Solution space structure of random constraint satisfaction problems with growing domains

In this paper we study the solution space structure of model RB, a standard prototype of Constraint Satisfaction Problem (CSPs) with growing domains. Using rigorous the first and the second moment method, we show that in the solvable phase close to the satisfiability transition, solutions are clustered into exponential number of well-separated clusters, with each cluster contains sub-exponential number of solutions. As a consequence, the system has a clustering (dynamical) transition but no condensation transition. This picture of phase diagram is different from other classic random CSPs with fixed domain size, such as random K-Satisfiability (K-SAT) and graph coloring problems, where condensation transition exists and is distinct from satisfiability transition. Our result verifies the non-rigorous results obtained using cavity method from spin glass theory, and sheds light on the structures of solution spaces of problems with a large number of states.

cond-mat.dis-nn

Impact of heterogenous prior beliefs and disclosed insider trades

In this paper, we present a multi-period trading model by assuming that traders face not only asymmetric information but also heterogenous prior beliefs, under the requirement that the insider publicly disclose his stock trades after the fact. We show that there is an equilibrium in which the irrational insider camouflages his trades with a noise component so that his private information is revealed slowly and linearly whenever he is overconfident or underconfident. We also investigate the relationship between the heterogeneous beliefs and the trade intensity in the presence of trade disclosure, and show that the weights on asymmetric information and heterogeneous prior beliefs are opposite in sign and they change alternatively in the next period. Under the requirement of disclosure, the irrational insider trades more aggressively and leads to smaller market depth. Moreover, the co-existence of "public disclosure requirement" and "heterogeneous prior beliefs" leads to the fluctuant multi-period expected profits and a larger total expected trading volume which is positively related to the degree of heterogeneity. More importantly, even public disclosure may lead to negative profits of the irrational insider's in some periods, inside trading remains profitable from the whole trading period.

q-fin.TR

Inside Trading, Public Disclosure and Imperfect Competition

In this paper, we present a multi-period trading model in the style of Kyle (1985)'s inside trading model, by assuming that there are at least two insiders in the market with long-lived private information, under the requirement that each insider publicly discloses his stock trades after the fact. Based on this model, we study the influences of "public disclosure" and "competition among insiders" on the trading behaviors of insiders. We find that the "competition among insiders" leads to higher effective price and lower insiders' profits, and the "public disclosure" makes each insider play a mixed strategy in every round except the last one. An interesting find is that as the total number of auctions goes to infinity, the market depth and the trading intensity at the first auction are all constants with the requirement of "public disclosure", while the market depth at the first auction goes to zero and the trading intensity of the first period goes to infinity without the requirement of "public disclosure".Moreover, we give the exact speed of the revelation of the private information, and show that all information is revealed immediately and the market depth goes to infinity immediately as trading happens infinitely frequently.

q-fin.TR

Insider Trading in the Market with Rational Expected Price

Kyle (1985) builds a pioneering and influential model, in which an insider with long-lived private information submits an optimal order in each period given the market maker's pricing rule. An inconsistency exists to some extent in the sense that the ``constant pricing rule " actually assumes an adaptive expected price with pricing rule given before insider making the decision, and the ``market efficiency" condition, however, assumes a rational expected price and implies that the pricing rule can be influenced by insider's strategy. We loosen the ``constant pricing rule " assumption by taking into account sufficiently the insider's strategy has on pricing rule. According to the characteristic of the conditional expectation of the informed profits, three different models vary with insider's attitudes regarding to risk are presented. Compared to Kyle (1985), the risk-averse insider in Model 1 can obtain larger guaranteed profits, the risk-neutral insider in Model 2 can obtain a larger ex ante expectation of total profits across all periods and the risk-seeking insider in Model 3 can obtain larger risky profits. Moreover, the limit behaviors of the three models when trading frequency approaches infinity are given, showing that Model 1 acquires a strong-form efficiency, Model 2 acquires the Kyle's (1985) continuous equilibrium, and Model 3 acquires an equilibrium with information released at an increasing speed.

q-fin.TR