SearcharxivSearch

arXiv subjects

Daniel Kaiser

Publications and source records attributed to Daniel Kaiser.

14 recordsLinked to original sources

Emergent contagion complexity: Disentangling mechanistic complexity from correlated heterogeneity

Simple and complex contagions differ mechanistically; multiple exposures act synergistically in the latter but independently in the former. Yet correlated mixtures of simple contagions may appear complex when inferring global contagion rules, a phenomenon we call "emergent complexity." We present a measure of contagion complexity and an inferential framework for estimating mixtures of nonparametric contagion rules from time-series data. Our work reframes past studies on complex contagion by offering heterogeneous mixtures of simple contagions as an alternative explanation.

physics.soc-ph

Techreport: Evaluating Tor-based Location Privacy for Ethereum Validators

Privacy and anonymity of validators, especially regarding IP address linkability, are essential to protect the Ethereum network from various attacks. Network-level attacks, such as DoS, can interrupt validators and affect the overall security of the Ethereum network. Correlating the IP addresses of validators with their identities, along with knowledge about their action slots can be exploited by attackers to cause network delays, MEV exploitation, and finality risks. Therefore, ensuring the unlinkability of a validator's IP and identity is crucial for maintaining the network's trust and resilience. In this techreport, we first provide a review of the existing network and consensus layer techniques that have been proposed for maintaining validator privacy in the Ethereum blockchain. Secondly, we evaluate a Tor-based protocol named Tor push that helps unlink validator identities (IDs) from their nodes' IP addresses, thereby making it difficult to determine any end-to-end correlation between validator IDs and IP addresses of validators' beacon nodes. To evaluate the effectiveness of Tor push, we present a working, deployed proof-of-concept (PoC) implementation in the Nimbus Ethereum client. Our PoC deployment pushes attestations, aggregations, and block proposals over Tor to the Goerli testnet. Furthermore, we also analyse the security and latency of Tor push. Our experimental results suggest that Tor can be incorporated into the existing Ethereum network with a tolerable latency overhead of 613.82 ms on average and without compromising the overall network performance while enhancing the location privacy of validators in the Ethereum network.

cs.CR

Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs

As reasoning LLMs increasingly trade tokens for accuracy through deliberation, search, and self-correction, a single accuracy score can no longer tell whether those tokens buy useful reasoning, recovery from hard instances, or unnecessary verbosity. We introduce a trace-optional evaluation protocol that exactly decomposes token efficiency using three observables available even for closed models: completion rate, conditional correctness given completion, and generated length. When instance-level workload metadata is available, we further normalize generated length by declared task-implied work and separate mean verbalization overhead from workload-dependent scaling. When such metadata is absent, we define an auditable solver-derived workload scale and evaluate its stability under leave-self-out, leave-top-k, and held-out-reference-pool perturbations. We evaluate 14 shared open-weight models on CogniLoad, GSM8K, ProofWriter, and ZebraLogic. We further evaluate 11 additional models on CogniLoad, enabling a fine-grained analysis of reasoning-task difficulty factors: task length, intrinsic difficulty, and distractor density. Efficiency and overhead rankings remain stable across all benchmark pairs, more robustly than accuracy rankings, while the decomposition separates logic-limited, context-limited (truncation-driven), and verbosity-limited failure modes that look identical under accuracy-per-token. We release an evaluation artifact and reporting template, which elaborates on why an LLM is inefficient at reasoning.

cs.CL

CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density

Current benchmarks for long-context reasoning in Large Language Models (LLMs) often blur critical factors like intrinsic task complexity, distractor interference, and task length. To enable more precise failure analysis, we introduce CogniLoad, a novel synthetic benchmark grounded in Cognitive Load Theory (CLT). CogniLoad generates natural-language logic puzzles with independently tunable parameters that reflect CLT's core dimensions: intrinsic difficulty ($d$) controls intrinsic load; distractor-to-signal ratio ($\rho$) regulates extraneous load; and task length ($N$) serves as an operational proxy for conditions demanding germane load. Evaluating 22 SotA reasoning LLMs, CogniLoad reveals distinct performance sensitivities, identifying task length as a dominant constraint and uncovering varied tolerances to intrinsic complexity and U-shaped responses to distractor ratios. By offering systematic, factorial control over these cognitive load dimensions, CogniLoad provides a reproducible, scalable, and diagnostically rich tool for dissecting LLM reasoning limitations and guiding future model development.

cs.CL

PREAMBLE and IMRECEIVING for Improved Large Message Handling in libp2p GossipSub

Large message transmissions in libp2p GossipSub lead to longer than expected network-wide message dissemination times and very high bandwidth utilization. This article identifies key issues responsible for this behavior and proposes modifications to the protocol for transmitting large messages. These modifications preserve the GossipSub resilience and fit well into the current algorithm. The proposed changes are rigorously evaluated for performance using the shadow simulator. Results reveal that the suggested changes reduce bandwidth utilization by up to 61% and message dissemination time by up to 35% under different traffic conditions.

cs.NI

Staggering and Fragmentation for Improved Large Message Handling in libp2p GossipSub

The libp2p GossipSub protocol leverages a full-message mesh with a lower node degree and a more densely connected metadata-only (gossip) mesh. This combination allows an efficient dissemination of messages in unstructured peer-to-peer (P2P) networks. However, GossipSub needs to consider message size, which is crucial for the efficient operation of many applications, such as handling large Ethereum blocks. This paper proposes modifications to improve GossipSub's performance when transmitting large messages. We evaluate the proposed improvements using the shadow simulator. Our results show that the proposed improvements significantly enhance GossipSub's performance for large message transmissions in sizeable networks.

cs.NI

Efficient inference of rankings from multi-body comparisons

Many of the existing approaches to assess and predict the performance of players, teams or products in competitive contests rely on the assumption that comparisons occur between pairs of such entities. There are, however, several real contests where more than two entities are part of each comparison, e.g., sports tournaments,multiplayer board and card games, and preference surveys. The Plackett-Luce (PL) model provides a principled approach to infer the ranking of entities involved in such contests characterized by multi-body comparisons. Unfortunately, traditional algorithms used to compute PL rankings suffer from slow convergence limiting the application of the PL model to relatively small-scale systems. We present here an alternative implementation that allows for significant speed-ups and validate its efficiency in both synthetic and real-world sets of data. Further, we perform systematic cross-validation tests concerning the ability of the PL model to predict unobserved comparisons. We find that a PL model trained on a set composed of multi-body comparisons is more predictive than a PL model trained on a set of projected pairwise comparisons derived from the very same training set, emphasizing the need of properly accounting for the true multi-body nature of real-world systems whenever such an information is available.

physics.soc-ph

Reconstruction of multiplex networks via graph embeddings

Multiplex networks are collections of networks with identical nodes but distinct layers of edges. They are genuine representations for a large variety of real systems whose elements interact in multiple fashions or flavors. However, multiplex networks are not always simple to observe in the real world; often, only partial information on the layer structure of the networks is available, whereas the remaining information is in the form of aggregated, single-layer networks. Recent works have proposed solutions to the problem of reconstructing the hidden multiplexity of single-layer networks using tools proper of network science. Here, we develop a machine learning framework that takes advantage of graph embeddings, i.e., representations of networks in geometric space. We validate the framework in systematic experiments aimed at the reconstruction of synthetic and real-world multiplex networks, providing evidence that our proposed framework not only accomplishes its intended task, but often outperforms existing reconstruction techniques.

physics.soc-ph

End-to-end topographic networks as models of cortical map formation and human visual behaviour: moving beyond convolutions

Computational models are an essential tool for understanding the origin and functions of the topographic organisation of the primate visual system. Yet, vision is most commonly modelled by convolutional neural networks that ignore topography by learning identical features across space. Here, we overcome this limitation by developing All-Topographic Neural Networks (All-TNNs). Trained on visual input, several features of primate topography emerge in All-TNNs: smooth orientation maps and cortical magnification in their first layer, and category-selective areas in their final layer. In addition, we introduce a novel dataset of human spatial biases in object recognition, which enables us to directly link models to behaviour. We demonstrate that All-TNNs significantly better align with human behaviour than previous state-of-the-art convolutional models due to their topographic nature. All-TNNs thereby mark an important step forward in understanding the spatial organisation of the visual brain and how it mediates visual behaviour.

q-bio.NC

Community Detection in Hypergraphs via Mutual Information Maximization

The hypergraph community detection problem seeks to identify groups of related nodes in hypergraph data. We propose an information-theoretic hypergraph community detection algorithm which compresses the observed data in terms of community labels and community-edge intersections. This algorithm can also be viewed as maximum-likelihood inference in a degree-corrected microcanonical stochastic blockmodel. We perform the inference/compression step via simulated annealing. Unlike several recent algorithms based on canonical models, our microcanonical algorithm does not require inference of statistical parameters such as node degrees or pairwise group connection rates. Through synthetic experiments, we find that our algorithm succeeds down to recently-conjectured thresholds for sparse random hypergraphs. We also find competitive performance in cluster recovery tasks on several hypergraph data sets.

cs.DM

Multiplex reconstruction with partial information

A multiplex is a collection of network layers, each representing a specific type of edges. This appears to be a genuine representation for many real-world systems. However, due to a variety of potential factors, such as limited budget and equipment, or physical impossibility, multiplex data can be difficult to observe directly. Often, only partial information on the layer structure of the system is available, whereas the remaining information is in the form of a single-layer network. In this work, we face the problem of reconstructing the hidden multiplex structure of an aggregated network from partial information. We propose an algorithm that leverages the layer-wise community structure that can be learned from partial observations to reconstruct the ground-truth topology of the unobserved part of the multiplex. The algorithm is characterized by a computational time that grows linearly with the network size. We perform a systematic study of reconstruction problems for both synthetic and real-world multiplex networks. We show that the ability of the proposed method to solve the reconstruction problem is affected by the heterogeneity of the individual layers and the similarity among the layers. On real-world networks, we observe that the accuracy of the reconstruction saturates quickly as the amount of available information increases. In genetic interaction and scientific collaboration multiplexes for example, we find that 10% of ground-truth information yields 70% accuracy, while 30% information allows for more than 90% accuracy.

physics.soc-ph

On a condition equivalent to the Maximum Distance Separable conjecture

We denote by $\mathcal{P}_q$ the vector space of functions from a finite field $\mathbb{F}_q$ to itself, which can be represented as the space $\mathcal{P}_q := \mathbb{F}_q[x]/(x^q-x)$ of polynomial functions. We denote by $\mathcal{O}_n \subset \mathcal{P}_q$ the set of polynomials that are either the zero polynomial, or have at most $n$ distinct roots in $\mathbb{F}_q$. Given two subspaces $Y,Z$ of $\mathcal{P}_q$, we denote by $\langle Y,Z \rangle$ their span. We prove that the following are equivalent. A) Let $k, q$ integers, with $q$ a prime power and $2 \leq k \leq q$. Suppose that either: 1) $q$ is odd 2) $q$ is even and $k \not\in \{3, q-1\}$. Then there do not exist distinct subspaces $Y$ and $Z$ of $\mathcal{P}_q$ such that: 1') $dim(\langle Y, Z \rangle) = k$ 2') $dim(Y) = dim(Z) = k-1$. 3') $\langle Y, Z \rangle \subset \mathcal{O}_{k-1}$ 4') $Y, Z \subset \mathcal{O}_{k-2}$ 5') $Y\cap Z \subset \mathcal{O}_{k-3}$. B) The MDS conjecture is true for the given $(q,k)$.

cs.IT

Realistic, Extensible DNS and mDNS Models for INET/OMNeT++

The domain name system (DNS) is one of the core services in today's network structures. In local and ad-hoc networks DNS is often enhanced or replaced by mDNS. As of yet, no simulation models for DNS and mDNS have been developed for INET/OMNeT++. We introduce DNS and mDNS simulation models for OMNeT++, which allow researchers to easily prototype and evaluate extensions for these protocols. In addition, we present models for our own experimental extensions, namely Stateless DNS and Privacy-Enhanced mDNS, that are based on the aforementioned models. Using our models we were able to further improve the efficiency of our protocol extensions.

cs.NI