SearcharxivSearch

arXiv subjects

Daniel Berend

Publications and source records attributed to Daniel Berend.

At least 19 recordsLinked to original sources

Power and Limits of Subset Selection in Statistical Estimation

We study the power and limitations of subset selection in statistical estimation through the framework of \emph{super-teaching}, where a teacher selects a subset of i.i.d. data to optimize a learner's estimator. Unlike prior work focused on specific distributions or fixed subset sizes, we develop a general theory under minimal assumptions. For mean estimation, we prove that super-teaching is possible for any distribution whose density is bounded away from zero in some neighborhood of the mean, allowing subset sizes growing as $k = o(n^{1/3})$ and achieving error on the order of roughly $k!/n^{k}$. This significantly extends existing results on admissible distributions and subset scaling. We also extend the analysis to parameters expressed as smooth functionals of expectations, such as variance and scale parameters in classical parametric families, including settings with heavy tails. Moreover, we show that super-teaching can greatly improve estimation rates for nonlinear estimators like the sample median, achieving rates beyond classical asymptotics. Through examples, including cases where maximum likelihood estimators are inconsistent or fail to be asymptotically normal, we demonstrate that super-teaching can succeed even when standard statistical guarantees break down. Our results establish a unified theory of data selection to enhance statistical efficiency.

math.ST

Fano Geometry and Slow Coupon Collecting

We study the coupon collector's problem in a generalized setting where each draw reveals a fixed number of coupons and the sampling mechanism is required to be \emph{fair}, meaning that every coupon appears with the same frequency among the admissible draws. Grunbaum and Yaakobi conjectured that, among all fair mechanisms with fixed parameters, the fully random model maximizes the expected time to complete coverage. We disprove this conjecture by exhibiting explicit counterexamples arising from finite geometry. In particular, we show that the line set of the Fano plane yields a fair mechanism whose expected coverage time exceeds that of the full model. Further exact and computational results are obtained for projective planes of higher order. In addition, we analyze a simple infinite family of fair mechanisms, the star mechanism, for which the expected coverage time admits a closed form. Depending on the scaling regime, this mechanism can be asymptotically slower or faster than the full model, showing that no universal extremality principle holds for fair mechanisms without additional structural assumptions.

math.CO

Asymptotic Results for Uniform Group Drawing in the Coupon Collector's Problem

The article explores the asymptotic behavior of the expected number of drawings in the Coupon Collector's Problem with group-drawing under the uniform distribution. In this variant, each draw consists of a package of $s$ distinct coupons selected uniformly at random from a set of $n$ coupons. We focus on three regimes of the package size $s$: (i) constant $s$, (ii) $s$ proportional to $n$, and (iii) $s$ "very close" to $n$. For each case, we provide precise asymptotic expressions for the expected collection time. Keywords: Coupon Collector's Problem, Group Drawings, Uniform Distribution, Asymptotic Analysis, Expected Collection Time

math.PR

A cop-robber game on metric graphs

We study a variant of the classical cop-robber game played on compact metric graphs, where each edge is assigned a positive length and identified with a real interval of corresponding length. In this setting, both the cop and the robber move continuously along the edges, subject to upper bounds on their speeds. The cop has no knowledge of the robber's location and must choose a continuous path through the graph that is guaranteed to intersect the robber's trajectory at some point in time. We show that for every compact metric graph, there exists a constant s > 0 such that if the cop's speed exceeds s times the robber's speed, then the cop can guarantee capture.

math.CO

DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing

Efficient cache management is critical for optimizing the system performance, and numerous caching mechanisms have been proposed, each exploring various insertion and eviction strategies. In this paper, we present AdaptiveClimb and its extension, DynamicAdaptiveClimb, two novel cache replacement policies that leverage lightweight, cache adaptation to outperform traditional approaches. Unlike classic Least Recently Used (LRU) and Incremental Rank Progress (CLIMB) policies, AdaptiveClimb dynamically adjusts the promotion distance (jump) of the cached objects based on recent hit and miss patterns, requiring only a single tunable parameter and no per-item statistics. This enables rapid adaptation to changing access distributions while maintaining low overhead. Building on this foundation, DynamicAdaptiveClimb further enhances adaptability by automatically tuning the cache size in response to workload demands. Our comprehensive evaluation across a diverse set of real-world traces, including 1067 traces from 6 different datasets, demonstrates that DynamicAdaptiveClimb consistently achieves substantial speedups and higher hit ratios compared to other state-of-the-art algorithms. In particular, our approach achieves up to a 29% improvement in hit ratio and a substantial reduction in miss penalties compared to the FIFO baseline. Furthermore, it outperforms the next-best contenders, AdaptiveClimb and SIEVE [43], by approximately 10% to 15%, especially in environments characterized by fluctuating working set sizes. These results highlight the effectiveness of our approach in delivering efficient performance, making it well-suited for modern, dynamic caching environments.

cs.OS

On Conjectures concerning the Labeled Coupon Collector Problem

We study a labeled variant of the classical Coupon Collector Problem (CCP), recently introduced by Tan et al., where coupons arrive in groups and only the set of labels is revealed. The goal is to determine the expected number of group drawings required to uniquely identify the labeling of all coupons. We focus on the case where groups consist of pairs ($k=2$), and provide rigorous proofs for two conjectures posed by Tan et al.

math.PR

On a Conjecture on Uniform Group Drawings in the Coupon Collector Problem

We address a conjecture of Schilling concerning the optimality of the uniform distribution in the generalized Coupon Collector's Problem (CCP) where, in each round, a subset (package) of $s$ coupons is drawn from a total of $n$ distinct coupons. While the classical CCP (with single-coupon draws) is well understood, the group-draw variant, where packages of size $s$ are drawn, presents new challenges and has applications in areas such as biological network models. Consider the set of all distributions over the collection of $\binom{n}{s}$ packages of size $s$. Schilling showed that, for $s=n-1$, the uniform distribution yields the minimal expected time for collecting all coupons. She further conjectured that, for $2\le s\le n-2$, the uniform distribution does not yield the minimum. We prove Schilling's conjecture in full by presenting "natural" non-uniform distributions yielding strictly lower expected collection times. Explicit formulas are provided for the expected number of rounds under these and related distributions Keywords: Coupon Collector's Problem, Group Drawings, Uniform Distribution, Expected Collection Time, Schilling's Conjecture, Optimal Distribution.

math.PR

Exceptional Points for Density Modulo 1

It is well known that almost every dilation of a sequence of real numbers, that diverges to $\infty$, is dense modulo~1. This paper studies the exceptional set of points -- those for which the dilation is not dense. Specifically, we consider the Hausdorff and modified box dimensions of the set of exceptional points. In particular, we show that the dimension of this set may be any number between 0 and 1. Similar results are obtained for two ``natural'' subsets of the set of exceptional points. Furthermore, the paper calculates the dimension of several sets of points, defined by certain constraints on their binary expansion.

math.NT

Exact expressions for the maximal probability that all $k$-wise independent bits are 1

Let $M(n, k, p)$ denote the maximum probability of the event $X_1 = X_2 = \cdots = X_n=1$ under a $k$-wise independent distribution whose marginals are Bernoulli random variables with mean $p$. A long-standing question is to calculate $M(n, k, p)$ for all values of $n,k,p$. This question has been partially addressed by several authors, primarily with the goal of answering asymptotic questions. The present paper focuses on obtaining exact expressions for this probability. To this end, we provide closed-form formulas of $M(n,k,p)$ for $p$ near 0 as well as $p$ near 1.

math.PR

Simultaneous visibility in the integer lattice

Two lattice points are visible from one another if there is no lattice point on the open line segment joining them. Let $S$ be a finite subset of $\mathbb{Z}^k$. The asymptotic density of the set of lattice points, visible from all points of $S$, was studied by several authors. Our main result is an improved upper bound on the error term. We also find the Schnirelmann density of the set of visible points from some sets S. Finally, we discuss these questions from the point of view of ergodic theory.

math.NT

Algorithms for Reconstructing DDoS Attack Graphs using Probabilistic Packet Marking

DoS and DDoS attacks are widely used and pose a constant threat. Here we explore Probability Packet Marking (PPM), one of the important methods for reconstructing the attack-graph and detect the attackers. We present two algorithms. Differently from others, their stopping time is not fixed a priori. It rather depends on the actual distance of the attacker from the victim. Our first algorithm returns the graph at the earliest feasible time, and turns out to guarantee high success probability. The second algorithm enables attaining any predetermined success probability at the expense of a longer runtime. We study the performance of the two algorithms theoretically, and compare them to other algorithms by simulation. Finally, we consider the order in which the marks corresponding to the various edges of the attack graph are obtained by the victim. We show that, although edges closer to the victim tend to be discovered earlier in the process than farther edges, the differences are much smaller than previously thought.

cs.CR

The Time for Reconstructing the Attack Graph in DDoS Attacks

Despite their frequency, denial-of-service (DoS\blfootnote{Denial of Service (DoS), Distributed Denial of Service (DDoS), Probabilistic Packet Marking (PPM), coupon collector's problem (CCP)}) and distributed-denial-of-service (DDoS) attacks are difficult to prevent and trace, thus posing a constant threat. One of the main defense techniques is to identify the source of attack by reconstructing the attack graph, and then filter the messages arriving from this source. One of the most common methods for reconstructing the attack graph is Probabilistic Packet Marking (PPM). We focus on edge-sampling, which is the most common method. Here, we study the time, in terms of the number of packets, the victim needs to reconstruct the attack graph when there is a single attacker. This random variable plays an important role in the reconstruction algorithm. Our main result is a determination of the asymptotic distribution and expected value of this time. The process of reconstructing the attack graph is analogous to a version of the well-known coupon collector's problem (with coupons having distinct probabilities). Thus, the results may be used in other applications of this problem.

math.PR

Maximum of Exponential Random Variables, Hurwitz's Zeta Function, and the Partition Function

A natural problem in the context of the coupon collector's problem is the behavior of the maximum of independent geometrically distributed random variables (with distinct parameters). This question has been addressed by Brennan et al. (British J. of Math. & CS. 8 (2015), 330-336). Here we provide explicit asymptotic expressions for the moments of that maximum, as well as of the maximum of exponential random variables with corresponding parameters. We also deal with the probability of each of the variables being the maximal one. The calculations lead to expressions involving Hurwitz's zeta function at certain special points. We find here explicitly the values of the function at these points. Also, the distribution function of the maximum we deal with is closely related to the generating function of the partition function. Thus, our results (and proofs) rely on classical results pertaining to the partition function.

math.PR

On Biased Random Walks, Corrupted Intervals, and Learning Under Adversarial Design

We tackle some fundamental problems in probability theory on corrupted random processes on the integer line. We analyze when a biased random walk is expected to reach its bottommost point and when intervals of integer points can be detected under a natural model of noise. We apply these results to problems in learning thresholds and intervals under a new model for learning under adversarial design.

cs.LG

A Model of Random Industrial SAT

One of the most studied models of SAT is random SAT. In this model, instances are composed from clauses chosen uniformly randomly and independently of each other. This model may be unsatisfactory in that it fails to describe various features of SAT instances, arising in real-world applications. Various modifications have been suggested to define models of industrial SAT. Here, we focus mainly on the aspect of community structure. Namely, here the set of variables consists of a number of disjoint communities, and clauses tend to consist of variables from the same community. Thus, we suggest a model of random industrial SAT, in which the central generalization with respect to random SAT is the additional community structure. There has been a lot of work on the satisfiability threshold of random $k$-SAT, starting with the calculation of the threshold of $2$-SAT, up to the recent result that the threshold exists for sufficiently large $k$. In this paper, we endeavor to study the satisfiability threshold for the proposed model of random industrial SAT. Our main result is that the threshold in this model tends to be smaller than its counterpart for random SAT. Moreover, under some conditions, this threshold even vanishes.

cs.DS

Minimum KL-divergence on complements of $L_1$ balls

Pinsker's widely used inequality upper-bounds the total variation distance $||P-Q||_1$ in terms of the Kullback-Leibler divergence $D(P||Q)$. Although in general a bound in the reverse direction is impossible, in many applications the quantity of interest is actually $D^*(P,\eps)$ --- defined, for an arbitrary fixed $P$, as the infimum of $D(P||Q)$ over all distributions $Q$ that are $\eps$-far away from $P$ in total variation. We show that $D^*(P,\eps)\le C\eps^2 + O(\eps^3)$, where $C=C(P)=1/2$ for "balanced" distributions, thereby providing a kind of reverse Pinsker inequality. An application to large deviations is given, and some of the structural results may be of independent interest. Keywords: Pinsker inequality, Sanov's theorem, large deviations

cs.IT

Consistency of weighted majority votes

We revisit the classical decision-theoretic problem of weighted expert voting from a statistical learning perspective. In particular, we examine the consistency (both asymptotic and finitary) of the optimal Nitzan-Paroush weighted majority and related rules. In the case of known expert competence levels, we give sharp error estimates for the optimal rule. When the competence levels are unknown, they must be empirically estimated. We provide frequentist and Bayesian analyses for this situation. Some of our proof techniques are non-standard and may be of independent interest. The bounds we derive are nearly optimal, and several challenging open problems are posed. Experimental results are provided to illustrate the theory.

math.PR

The state complexity of random DFAs

The state complexity of a Deterministic Finite-state automaton (DFA) is the number of states in its minimal equivalent DFA. We study the state complexity of random $n$-state DFAs over a $k$-symbol alphabet, drawn uniformly from the set $[n]^{[n]\times[k]}\times2^{[n]}$ of all such automata. We show that, with high probability, the latter is $α_k n + O(\sqrt n\log n)$ for a certain explicit constant $α_k$.

math.PR