SearcharxivSearch

arXiv subjects

Orit Davidovich

Publications and source records attributed to Orit Davidovich.

7 recordsLinked to original sources

Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers

We formally define algorithmic capture of combinatorial tasks as the ability of a transformer to extrapolate to arbitrary task sizes with controllable error and logarithmic sample adaptation, providing a sharp scaling criterion for distinguishing logic internalization from statistical interpolation. Empirically, across scaling ranges spanning up to 2.5 orders of magnitude, we observe evidence of capture and non-capture. By analyzing infinite-width transformers in both the lazy and rich regimes, we derive upper bounds on the inference-time computational complexity of the combinatorial tasks these networks can capture. We show that, despite their universal expressivity, transformers possess an inductive bias that disfavors higher-complexity algorithmic procedures within the efficient polynomial-time heuristic scheme class, consistent with successful capture on simpler combinatorial tasks such as induction heads, sort, and string matching.

cs.LG

Heuristics for Combinatorial Optimization via Value-based Reinforcement Learning: A Unified Framework and Analysis

Since the 1990s, considerable empirical work has been carried out to train statistical models, such as neural networks (NNs), as learned heuristics for combinatorial optimization (CO) problems. When successful, such an approach eliminates the need for experts to design heuristics per problem type. Due to their structure, many hard CO problems are amenable to treatment through reinforcement learning (RL). Indeed, we find a wealth of literature training NNs using value-based, policy gradient, or actor-critic approaches, with promising results, both in terms of empirical optimality gaps and inference runtimes. Nevertheless, there has been a paucity of theoretical work undergirding the use of RL for CO problems. To this end, we introduce a unified framework to model CO problems through Markov decision processes (MDPs) and solve them using RL techniques. We provide easy-to-test assumptions under which CO problems can be formulated as equivalent undiscounted MDPs that provide optimal solutions to the original CO problems. Moreover, we establish conditions under which value-based RL techniques converge to approximate solutions of the CO problem with a guarantee on the associated optimality gap. Our convergence analysis provides: (1) a sufficient rate of increase in batch size and projected gradient descent steps at each RL iteration; (2) the resulting optimality gap in terms of problem parameters and targeted RL accuracy; and (3) the importance of a choice of state-space embedding. Together, our analysis illuminates the success (and limitations) of the celebrated deep Q-learning algorithm in this problem context.

stat.ML

Mitigating the Curse of Detail: Scaling Arguments for Feature Learning and Sample Complexity

Two pressing topics in the theory of deep learning are the interpretation of feature learning (FL) mechanisms and the determination of implicit bias of networks in the rich regime. Current theories of rich FL often appear in the form of high-dimensional non-linear equations, which require computationally intensive numerical solutions. Given the many details that go into defining a deep learning problem, this analytical complexity is a significant and often unavoidable challenge. Here, we propose a powerful heuristic route for predicting the data and width scales at which various patterns of FL emerge. This form of scale analysis is considerably simpler than such exact theories and reproduces the scaling exponents of various known results. In addition, we make novel predictions on complex toy architectures, such as three-layer non-linear networks and attention heads, thus extending the scope of first-principle theories of deep learning.

cs.LG

Finding Probably Approximate Optimal Solutions by Training to Estimate the Optimal Values of Subproblems

The paper is about developing a solver for maximizing a real-valued function of binary variables. The solver relies on an algorithm that estimates the optimal objective-function value of instances from the underlying distribution of objectives and their respective sub-instances. The training of the estimator is based on an inequality that facilitates the use of the expected total deviation from optimality conditions as a loss function rather than the objective-function itself. Thus, it does not calculate values of policies, nor does it rely on solved instances.

cs.LG

Average pace and horizontal chords

We are motivated by a problem about running: If a race was completed in an average pace of P minutes per mile, is there necessarily some mile of the race that was run in exactly P minutes? The answer is no. We explain why, and describe the history of this celebrated problem, known as the Universal Chord Theorem. We also clarify and streamline the proof of a more powerful result by Heinz Hopf from 1937.

math.HO

On Arithmetic Modular Categories

Modular categories are important algebraic structures in a variety of subjects in mathematics and physics. We provide an explicit, motivated and elementary definition of a modular category over a field of characteristic 0 as an equivalence class of solutions to a set of polynomial equations. We conclude that within each class of solutions, there is one which consists entirely of algebraic numbers. These algebraic solutions make it possible to discuss defining algebraic number fields of modular categories and their Galois twists. One motivation for such a definition is an arithmetic theory of modular categories which plays an important role in their classification. Another is to facilitate implementation of computer-based tools to resolve computational and classification problems intractible by other means. We observe some basic properties of Galois twists of modular categories and make conjectures about their relation to the the intrinsic data of modular categories.

math.QA

Isoscattering on surfaces

We give a number of examples of pairs of non-compact surfaces which are isoscattering, and which are exceptionally simple in one or more senses. We give examples which are of small genus with a small number of ends, and also examles which are congruence surfaces.

math.DG