Searcharxiv⌕ Search

arXiv subjects

Mehmet Süzen

Publications and source records attributed to Mehmet Süzen.

11 recordsLinked to original sources

H-theorem do-conjecture

A pedagogical formulation of Loschmidt's paradox and H-theorem is presented with basic notation on occupancy on discrete states without invoking velocity collision operators. A conjecture, so called H-theorem do-conjecture, is formulated. Causal inference perspective on the dynamical evolution of classical many-particle system is invoked. This perspectice introduce a probabilistic view on the state of the system conditioning on the thermodyamic ensemble, i.e., function of state-variables representing the ensemble. A numerical simulation of random walkers for deterministic diffusion demonstrate the causal effect of interventional ensemble, showing a dynamical behaviour as a test of the proposed conjecture. Moreover, the chosen game like dynamics provides an accessible practical example, named Ising-Conway Entropy Game, in order to demonstrate increase in entropy over time, as a toy system of statistical physics.

cond-mat.stat-mech↗

Generalised learning of time-series: Ornstein-Uhlenbeck processes

In machine learning, statistics, econometrics and statistical physics, cross-validation (CV) is used asa standard approach in quantifying the generalisation performance of a statistical model. A directapplication of CV in time-series leads to the loss of serial correlations, a requirement of preserving anynon-stationarity and the prediction of the past data using the future data. In this work, we proposea meta-algorithm called reconstructive cross validation (rCV ) that avoids all these issues. At first,k folds are formed with non-overlapping randomly selected subsets of the original time-series. Then,we generate k new partial time-series by removing data points from a given fold: every new partialtime-series have missing points at random from a different entire fold. A suitable imputation or asmoothing technique is used to reconstruct k time-series. We call these reconstructions secondarymodels. Thereafter, we build the primary k time-series models using new time-series coming fromthe secondary models. The performance of the primary models are evaluated simultaneously bycomputing the deviations from the originally removed data points and out-of-sample (OSS) data.Full cross-validation in time-series models can be practiced with rCV along with generating learning curves.

stat.ML↗

Equivalence in Deep Neural Networks via Conjugate Matrix Ensembles

A numerical approach is developed for detecting the equivalence of deep learning architectures. The method is based on generating Mixed Matrix Ensembles (MMEs) out of deep neural network weight matrices and {\it conjugate circular ensemble} matching the neural architecture topology. Following this, the empirical evidence supports the {\it phenomenon} that difference between spectral densities of neural architectures and corresponding {\it conjugate circular ensemble} are vanishing with different decay rates at the long positive tail part of the spectrum i.e., cumulative Circular Spectral Difference (CSD). This finding can be used in establishing equivalences among different neural architectures via analysis of fluctuations in CSD. We investigated this phenomenon for a wide range of deep learning vision architectures and with circular ensembles originating from statistical quantum mechanics. Practical implications of the proposed method for artificial and natural neural architectures discussed such as the possibility of using the approach in Neural Architecture Search (NAS) and classification of biological neural networks.

cs.LG↗

Periodic Spectral Ergodicity: A Complexity Measure for Deep Neural Networks and Neural Architecture Search

Establishing associations between the structure and the generalisation ability of deep neural networks (DNNs) is a challenging task in modern machine learning. Producing solutions to this challenge will bring progress both in the theoretical understanding of DNNs and in building new architectures efficiently. In this work, we address this challenge by developing a new complexity measure based on the concept of {Periodic Spectral Ergodicity} (PSE) originating from quantum statistical mechanics. Based on this measure a technique is devised to quantify the complexity of deep neural networks from the learned weights and traversing the network connectivity in a sequential manner, hence the term cascading PSE (cPSE), as an empirical complexity measure. This measure will capture both topological and internal neural processing complexity simultaneously. Because of this cascading approach, i.e., a symmetric divergence of PSE on the consecutive layers, it is possible to use this measure for Neural Architecture Search (NAS). We demonstrate the usefulness of this measure in practice on two sets of vision models, ResNet and VGG, and sketch the computation of cPSE for more complex network structures.

cs.LG↗

HARK Side of Deep Learning -- From Grad Student Descent to Automated Machine Learning

Recent advancements in machine learning research, i.e., deep learning, introduced methods that excel conventional algorithms as well as humans in several complex tasks, ranging from detection of objects in images and speech recognition to playing difficult strategic games. However, the current methodology of machine learning research and consequently, implementations of the real-world applications of such algorithms, seems to have a recurring HARKing (Hypothesizing After the Results are Known) issue. In this work, we elaborate on the algorithmic, economic and social reasons and consequences of this phenomenon. We present examples from current common practices of conducting machine learning research (e.g. avoidance of reporting negative results) and failure of generalization ability of the proposed algorithms and datasets in actual real-life usage. Furthermore, a potential future trajectory of machine learning research and development from the perspective of accountable, unbiased, ethical and privacy-aware algorithmic decision making is discussed. We would like to emphasize that with this discussion we neither claim to provide an exhaustive argumentation nor blame any specific institution or individual on the raised issues. This is simply a discussion put forth by us, insiders of the machine learning field, reflecting on us.

cs.LG↗

Compressive Transition Path Sampling

Algorithms for rare event complex systems simulations are proposed. Compressed Sensing (CS) has {\it revolutionized} our understanding of limits in signal recovery and has forced us to re-define Shannon-Nyquist sampling theorem for sparse recovery. A formalism to reconstruct trajectories and transition paths via CS is illustrated as proposed algorithms. The implication of under-sampling is quite important. This formalism could increase the tractable time-scales {\it immensely} for simulation of statistical mechanical systems and rare event simulations. While, long time-scales are known to be a major hurdle and a challenge for realistic complex simulations for rare events. The outline of how to implement, test and possible challenges on the proposed approach are discussed in detail.

physics.comp-ph↗

Spectral Ergodicity in Deep Learning Architectures via Surrogate Random Matrices

In this work a novel method to quantify spectral ergodicity for random matrices is presented. The new methodology combines approaches rooted in the metrics of Thirumalai-Mountain (TM) and Kullbach-Leibler (KL) divergence. The method is applied to a general study of deep and recurrent neural networks via the analysis of random matrix ensembles mimicking typical weight matrices of those systems. In particular, we examine circular random matrix ensembles: circular unitary ensemble (CUE), circular orthogonal ensemble (COE), and circular symplectic ensemble (CSE). Eigenvalue spectra and spectral ergodicity are computed for those ensembles as a function of network size. It is observed that as the matrix size increases the level of spectral ergodicity of the ensemble rises, i.e., the eigenvalue spectra obtained for a single realisation at random from the ensemble is closer to the spectra obtained averaging over the whole ensemble. Based on previous results we conjecture that success of deep learning architectures is strongly bound to the concept of spectral ergodicity. The method to compute spectral ergodicity proposed in this work could be used to optimise the size and architecture of deep as well as recurrent neural networks.

stat.ML↗

A testable prediction from entropic gravity

I have shown conceptually that quantum state has a direct relationship to gravitational constant due to entropic force posed by Verlinde's argument and part of the Newton-Schrödinger equation (N-S) in the context of gravity induced collapse of the wavefunction via Diósi-Penrose proposal. This direct relationship can be used to measure gravitational constant using state-of-the-art mater-wave interferometry to test the entropic gravity argument.

physics.gen-ph↗

Evaluating Gaussian processes for sparse irregular spatio-temporal data

A practical approach to evaluate performance of a Gaussian process regression models (GPR) for irregularly sampled sparse time-series is introduced. The approach entails construction of a secondary autoregressive model using the fine scale predictions to forecast a future observation used in GPR. We build different GPR models for Ornstein-Uhlenbeck and Fractional processes for simulated toy data with different sparsity levels to assess the utility of the approach.

stat.ME↗

Effective ergodicity in single-spin-flip dynamics

A quantitative measure of convergence to effective ergodicity, the Thirumalai-Mountain (TM) metric, is applied to Metropolis and Glauber single-spin-flip dynamics. In computing this measure, finite lattice ensemble averages are obtained using the exact solution for a one dimensional Ising model, whereas the time averages are computed with Monte Carlo simulations. The time evolution of the effective ergodic convergence of Ising magnetization is monitored. By this approach, diffusion regimes of the effective ergodic convergence of magnetization are identified for different lattice sizes, nonzero temperature, and nonzero external field values. Results show that caution should be taken when using the TM metric at system parameters that give rise to strong correlations.

cond-mat.stat-mech↗

Adaptive Dynamic Congestion Avoidance with Master Equation

This paper proposes an adaptive variant of Random Early Detection (RED) gateway queue management for packet-switched networks via a discrete state analog of the non-stationary Master Equation i.e. Markov process. The computation of average queue size, which appeared in the original RED algorithm, is altered by introducing a probability $P(l,t)$, which defines the probability of having $l$ number of packets in the queue at the given time $t$, and depends upon the previous state of the queue. This brings the advantage of eliminating a free parameter: queue weight, completely. Computation of transition rates and probabilities are carried out on the fly, and determined by the algorithm automatically. Simulations with unstructured packets illustrate the method, the performance of the adaptive variant of RED algorithm, and the comparison with the standard RED.

cs.NI↗