SearcharxivSearch

arXiv subjects

A. Naumov

Publications and source records attributed to A. Naumov.

9 recordsLinked to original sources

CayleyPy Growth: Efficient growth computations and hundreds of new conjectures on Cayley graphs (Brief version)

This is the third paper of the CayleyPy project applying artificial intelligence to problems in group theory. We announce the first public release of CayleyPy, an open source Python library for computations with Cayley and Schreier graphs. Compared with systems such as GAP and Sage, CayleyPy handles much larger graphs and performs several orders of magnitude faster. Using CayleyPy we obtained about 200 new conjectures on Cayley and Schreier graphs, focused on diameters and growth. For many Cayley graphs of symmetric groups Sn we observe quasi polynomial diameter formulas: a small set of quadratic or linear polynomials indexed by n mod s. We conjecture that this is a general phenomenon, giving efficient diameter computation despite the problem being NP hard. We propose a refinement of the Babai type conjecture on diameters of Sn: n^2/2 + 4n upper bounds in the undirected case, compared to previous O(n^2) bounds. We also provide explicit generator families, related to involutions in a square with whiskers pattern, conjectured to maximize the diameter; search confirms this for all n up to 15. We further conjecture an answer to a question posed by V M Glushkov in 1968 on directed Cayley graphs generated by a cyclic shift and a transposition. For nilpotent groups we conjecture an improvement of J S Ellenberg's results on upper unitriangular matrices over Z/pZ, showing linear dependence of diameter on p. Some conjectures are LLM friendly, naturally stated as sorting problems verifiable by algorithms or Python code. To benchmark path finding we created more than 10 Kaggle datasets. CayleyPy works with arbitrary permutation or matrix groups and includes over 100 predefined generators. Our growth computation code outperforms GAP and Sage up to 1000 times in speed and size.

math.CO

TQCompressor: improving tensor decomposition methods in neural networks via permutations

We introduce TQCompressor, a novel method for neural network model compression with improved tensor decompositions. We explore the challenges posed by the computational and storage demands of pre-trained language models in NLP tasks and propose a permutation-based enhancement to Kronecker decomposition. This enhancement makes it possible to reduce loss in model expressivity which is usually associated with factorization. We demonstrate this method applied to the GPT-2$_{small}$. The result of the compression is TQCompressedGPT-2 model, featuring 81 mln. parameters compared to 124 mln. in the GPT-2$_{small}$. We make TQCompressedGPT-2 publicly available. We further enhance the performance of the TQCompressedGPT-2 through a training strategy involving multi-step knowledge distillation, using only a 3.1% of the OpenWebText. TQCompressedGPT-2 surpasses DistilGPT-2 and KnGPT-2 in comparative evaluations, marking an advancement in the efficient and effective deployment of models in resource-constrained environments.

cs.LG

Tetra-AML: Automatic Machine Learning via Tensor Networks

Neural networks have revolutionized many aspects of society but in the era of huge models with billions of parameters, optimizing and deploying them for commercial applications can require significant computational and financial resources. To address these challenges, we introduce the Tetra-AML toolbox, which automates neural architecture search and hyperparameter optimization via a custom-developed black-box Tensor train Optimization algorithm, TetraOpt. The toolbox also provides model compression through quantization and pruning, augmented by compression using tensor networks. Here, we analyze a unified benchmark for optimizing neural networks in computer vision tasks and show the superior performance of our approach compared to Bayesian optimization on the CIFAR-10 dataset. We also demonstrate the compression of ResNet-18 neural networks, where we use 14.5 times less memory while losing just 3.2% of accuracy. The presented framework is generic, not limited by computer vision problems, supports hardware acceleration (such as with GPUs and TPUs) and can be further extended to quantum hardware and to hybrid quantum machine learning models.

cs.LG

Variance reduction for dependent sequences with applications to Stochastic Gradient MCMC

In this paper we propose a novel and practical variance reduction approach for additive functionals of dependent sequences. Our approach combines the use of control variates with the minimisation of an empirical variance estimate. We analyse finite sample properties of the proposed method and derive finite-time bounds of the excess asymptotic variance to zero. We apply our methodology to Stochastic Gradient MCMC (SGMCMC) methods for Bayesian inference on large data sets and combine it with existing variance reduction methods for SGMCMC. We present empirical results carried out on a number of benchmark examples showing that our variance reduction method achieves significant improvement as compared to state-of-the-art methods at the expense of a moderate increase of computational overhead.

math.ST

Variance reduction for Markov chains with application to MCMC

In this paper we propose a novel variance reduction approach for additive functionals of Markov chains based on minimization of an estimate for the asymptotic variance of these functionals over suitable classes of control variates. A distinctive feature of the proposed approach is its ability to significantly reduce the overall finite sample variance. This feature is theoretically demonstrated by means of a deep non asymptotic analysis of a variance reduced functional as well as by a thorough simulation study. In particular we apply our method to various MCMC Bayesian estimation problems where it favourably compares to the existing variance reduction approaches.

math.ST

Semicircle Law for a Class of Random Matrices with Dependent Entries

In this paper we study ensembles of random symmetric matrices $\X_n = {X_{ij}}_{i,j = 1}^n$ with dependent entries such that $\E X_{ij} = 0$, $\E X_{ij}^2 = σ_{ij}^2$, where $σ_{ij}$ may be different numbers. Assuming that the average of the normalized sums of variances in each row converges to one and Lindeberg condition holds we prove that the empirical spectral distribution of eigenvalues converges to Wigner's semicircle law.

math.PR

Finite-amplitude wave propagation in a stratified fluid of variable depth

Variable-coefficient Korteweg - de Vries equation is applied to describe the interfacial wave transformation in two-layer fluid of variable depth. The soliton dynamics in this fluid is studied. The solitary wave breaks in two transient points. One of them is the point when two-layer fluid transforms to the on-layer flow. The second one is the point where layer thickness are equaled. The soliton amplitude dependence on the thickness of lower layer is found.

physics.ao-ph

Oriented rotational wave-packet dynamics studies via high harmonic generation

We produce oriented rotational wave packets in CO and measure their characteristics via high harmonic generation. The wavepacket is created using an intense, femtosecond laser pulse and its second harmonic. A delayed 800 nm pulse probes the wave packet, generating even-order high harmonics that arise from the broken symmetry induced by the orientation dynamics. The even-order harmonic radiation that we measure appears on a zero background, enabling us to accurately follow the temporal evolution of the wave packet. Our measurements reveal that, for the conditions optimum for harmonic generation, the orientation is produced by preferential ionization which depletes the sample of molecules of one orientation.

physics.atom-ph

Order-dependent structure of High Harmonic Wavefronts

The physics of high harmonics has led to the generation of attosecond pulses and to trains of attosecond pulses. Measurements that confirm the pulse duration are all performed in the far field. All pulse duration measurements tacitly assume that both the beam's wavefront and intensity profile are independent of frequency. However, if one or both are frequency dependent, then the retrieved pulse duration depends on the location where the measurement is made. We measure that each harmonic is very close to a Gaussian, but we also find that both the intensity profile and the beam wavefront depend significantly on the harmonic order. Thus, our findings mean that the pulse duration will depend on where the pulse is observed. Measurement of spectrally resolved wavefronts along with temporal characterization at one single point in the beam would enable complete space-time reconstruction of attosecond pulses. Future attosecond science experiments need not be restricted to spatially averaged observables.

physics.optics