SearcharxivSearch

arXiv subjects

Alexander Shen

Publications and source records attributed to Alexander Shen.

At least 19 recordsLinked to original sources

Keeping Score: Efficiency Improvements in Neural Likelihood Surrogate Training via Score-Augmented Loss Functions

For stochastic process models, parameter inference is often severely bottlenecked by computationally expensive likelihood functions. Simulation-based inference (SBI) bypasses this restriction by constructing amortized surrogate likelihoods, but most SBI methods assume a black-box data generating process. While these surrogates are exact in the limit of infinite training data, practical scenarios force a strict tradeoff between model quality and simulation cost. In this work, we loosen the black-box assumption of SBI to improve this tradeoff for structured stochastic process models. Specifically, for neural network likelihood surrogates trained via probabilistic classification, we propose to augment the standard binary cross-entropy loss with exact score information $\nabla_\theta \log p(x \mid \theta)$ and adaptive weighting based on loss gradients. We evaluate our approach on case studies involving network dynamics and spatial processes, demonstrating that our method improves surrogate quality at a drastically lower computational cost than generating more training data. Notably, in some cases, our approach achieves downstream inference performance equivalent to a 10x increase in training data with less than a 1.1x increase in training time.

stat.ML

Bishop's (up)crossing inequality and lower semicomputable random reals revisited

In this paper we provide an easy proof of Barmpalias--Lewis-Pye result saying that all computable increasing sequences converging to random reals converge with the same speed (up to a $c+o(1)$ factor) by noting that it immediately follows from Bishop's upcrossing inequality. We also provide a simple derivation of this inequality.

math.LO

All Kolmogorov complexity functions are optimal, but are some more optimal?

Kolmogorov (1965) defined the complexity of a string $x$ as the minimal length of a program generating $x$. Obviously this definition depends on the choice of the programming language. Kolmogorov noted that there exist \emph{optimal} programming languages that make the complexity function minimal up to $O(1)$ additive terms, and we should take one of them -- but which one? Is there a chance to agree on some specific programming language in this definition? Or at least should we add some other requirements to optimality? What can we achieve in this way? In this paper we discuss different suggestions of this type that appeared since 1965, specifically a stronger requirement of universality (and show that in many cases this does not change the set of complexity functions).

cs.IT

Local obstructions in sequences revisited

In this article, we consider some simple combinatorial game and a winning strategy in this game. This game is then used to prove several known results about non-repetitive sequences and approximations with denominators from a lacunary sequence. In this way we simplify the proofs, improve the bounds and get for free the computable versions that required a separate treatment.

math.CO

Optimal bounds for dissatisfaction in perpetual voting

In perpetual voting, multiple decisions are made at different moments in time. Taking the history of previous decisions into account allows us to satisfy properties such as proportionality over periods of time. In this paper, we consider the following question: is there a perpetual approval voting method that guarantees that no voter is dissatisfied too many times? We identify a sufficient condition on voter behavior -- which we call 'bounded conflicts' condition -- under which a sublinear growth of dissatisfaction is possible. We provide a tight upper bound on the growth of dissatisfaction under bounded conflicts, using techniques from Kolmogorov complexity. We also observe that the approval voting with binary choices mimics the machine learning setting of prediction with expert advice. This allows us to present a voting method with sublinear guarantees on dissatisfaction under bounded conflicts, based on the standard techniques from prediction with expert advice.

cs.GT

Specialized Foundation Models Struggle to Beat Supervised Baselines

Following its success for vision and text, the "foundation model" (FM) paradigm -- pretraining large models on massive data, then fine-tuning on target tasks -- has rapidly expanded to domains in the sciences, engineering, healthcare, and beyond. Has this achieved what the original FMs accomplished, i.e. the supplanting of traditional supervised learning in their domains? To answer we look at three modalities -- genomics, satellite imaging, and time series -- with multiple recent FMs and compare them to a standard supervised learning workflow: model development, hyperparameter tuning, and training, all using only data from the target task. Across these three specialized domains, we find that it is consistently possible to train simple supervised models -- no more complicated than a lightly modified wide ResNet or UNet -- that match or even outperform the latest foundation models. Our work demonstrates that the benefits of large-scale pretraining have yet to be realized in many specialized areas, reinforces the need to compare new FMs to strong, well-tuned baselines, and introduces two new, easy-to-use, open-source, and automated workflows for doing so.

cs.LG

Kolmogorov complexity as a combinatorial tool

Kolmogorov complexity is often used as a convenient language for counting and/or probabilistic existence proofs. However, there are some applications where Kolmogorov complexity is used in a more subtle way. We provide one (somehow) surprising example where an existence of a winning strategy in a natural combinatorial game is proven (and no direct proof is known).

cs.DM

Comparing angles in Euclid's Elements

The exposition in Euclid's Elements contains an obvious gap (seemingly unnoticed by most commentators): he often compares not just angles, but *groups* of angles, and at the same time he avoids summing angles (and considering angles greater than $\pi$), and does not say what such a comparison of groups could mean. We discuss the problem and suggest a possible interpretation that could make Euclid's exposition consistent.

math.HO

Conditional normality and finite-state dimensions revisited

The notion of a normal bit sequence was introduced by Borel in 1909; it was the first definition of an individual random object. Normality is a weak notion of randomness requiring only that all $2^n$ factors (substrings) of arbitrary length~$n$ appear with the same limit frequency $2^{-n}$. Later many stronger definitions of randomness were introduced, and in this context normality found its place as ``randomness against a finite-memory adversary''. A quantitative measure of finite-state compressibility was also introduced (the finite-state dimension) and normality means that the finite state dimension is maximal (equals~$1$). Recently Nandakumar, Pulari and S (2023) introduced the notion of relative finite-state dimension for a binary sequence with respect to some other binary sequence (treated as an oracle), and the corresponding notion of conditional (relative) normality. (Different notions of conditional randomness were considered before, but not for the finite memory case.) They establish equivalence between the block frequency and the gambling approaches to conditional normality and finite-state dimensions. In this note we revisit their definitions and explain how this equivalence can be obtained easily by generalizing known characterizations of (unconditional) normality and dimension in terms of compressibility (finite-state complexity), superadditive complexity measures and gambling (finite-state gales), thus also answering some questions left open in the above-mentioned paper.

cs.IT

Toward the End-To-End Optimization of the SWGO Array Layout

In this document we consider the problem of finding the optimal layout for the array of water Cherenkov detectors proposed by the SWGO collaboration to study very-high-energy gamma rays in the southern hemisphere. We develop a continuous model of the secondary particles produced by atmospheric showers initiated by high-energy gamma rays and protons, and build an optimization pipeline capable of identifying the most promising configuration of the detector elements. The pipeline employs stochastic gradient descent to maximize a utility function aligned with the scientific goals of the experiment. We demonstrate how the software is capable of finding the global maximum in the high-dimensional parameter space, and discuss its performance and limitations.

astro-ph.IM

Israil Moiseevich Gelfand [In Russian]

I was lucky to meet (and even cooperate at some extent) with Israel M. Gelfand, and tried to write down (mainly in 2003-2013) my recollections about his work style and lessons I learned from him about teaching and writing mathematics (though the views are my own and I.M. is not responsible for them in any way).

math.HO

Ergodic theorem and algorithmic randomness

We prove the constructive version of Birkhoff's ergodic theorem following Vyugin but trying to separate and state explicitly the combinatorial statement on which this proof is based. We pose some questions related to this statement (and the effective ergodic theorem in general).

math.DS

Constructive mathematics and teaching

Constructivists (and intuitionists in general) asked what kind of mental construction is needed to convince ourselves (and others) that some mathematical statement is true. This question has a much more practical (and even cynical) counterpart: a student of a mathematics class wants to know what will the teacher accept as a correct solution of a homework problem. Here the logical structure of the claim is also very important, and we discuss several types of problems and their use in teaching mathematics.

math.HO

The Kraft--Barmpalias--Lewis-Pye lemma revisited

This note provides a simplified exposition of the proof of hierarchical Kraft lemma proven by Barmpalias and Lewis-Pye and its consequences for the oracle use in the Ku\v{c}era--G\'acs theorem (saying that every sequence is Turing reducible to a random one).

cs.IT

Kolmogorov Last Discovery? (Kolmogorov and Algorithmic Statictics)

The last theme of Kolmogorov's mathematics research was algorithmic theory of information, now often called Kolmogorov complexity theory. There are only two main publications of Kolmogorov (1965 and 1968-1969) on this topic. So Kolmogorov's ideas that did not appear as proven (and published) theorems can be reconstructed only partially based on work of his students and collaborators, short abstracts of his talks and the recollections of people who were present at these talks. In this survey we try to reconstruct the development of Kolmogorov's ideas related to algorithmic statistics (resource-bounded complexity, structure function and stochastic objects).

math.LO

Inequalities for entropies and dimensions

We show that linear inequalities for entropies have a natural geometric interpretation in terms of Hausdorff and packing dimensions, using the point-to-set principle and known results about inequalities for complexities, entropies and the sizes of subgroups.

cs.IT

27 Open Problems in Kolmogorov Complexity

The paper proposes open problems in classical Kolmogorov complexity. Each problem is presented with background information and thus the article also surveys some recent studies in the area.

cs.IT

Individual codewords

Algorithmic information theory translates statements about classes of objects into statements about individual objects; it defines individual random sequences, effective Hausdorff dimension of individual points, amount of information in individual strings, etc. We observe that a similar translation is possible for list-decodable codes.

cs.IT