SearcharxivSearch

arXiv subjects

Karsten Müller

Publications and source records attributed to Karsten Müller.

15 recordsLinked to original sources

The Bombieri--van der Poorten Formula for Partial Quotients of Higher Degree Algebraic Irrationals

The fundamental relationship between the partial quotients $b_{n+1}$ of an algebraic irrational $α= \sqrt[m]{k}$ and its corresponding algebraic form $d_n = |p_n^m - k q_n^m|$ was elegantly proposed by Bombieri and van der Poorten. In this paper, we work out the explicit analytical details of the framework for any degree $m \geq 3$. We provide a closed-form derivation of the error term and prove for the cubic case that the remainder $|R_n|$ is strictly bounded by 1 for all convergents with $q_n \geq 2$.

math.NT

Gain Bounds for Diagonal Superelliptic Equations under the Strong ABC Conjecture

We establish a novel framework for bounding the adapted power gain $G_p$ and approximation gain $G_a$ of coprime integer solutions to the generalized diagonal superelliptic equation $By^n = Ax^n + k$ with $x, y \ge 2$. By first deriving a purely structural lower bound for $G_a$, we demonstrate that these equations are inherently predisposed to high ABC-qualities ($q = G_a \cdot G_p$). Combined with the Strong ABC conjecture ($q < q_{max}$), we prove that the power gain is uniformly bounded by $G_p < q_{max}/G_{a,min}$, providing a theoretical foundation for the numerical observation $G_p < 3$ for $n=2$ under the Ultra-Strong conjecture ($q < 1.5$). Specifically, we show that for $k=1$, the structural density forces $q > n/2$, which excludes solutions for $n \ge 4$ under $q < 2$. We validate our theoretical bounds using high-quality ABC triples, specifically analyzing the Reyssat (1987), de Weger (1985), and Nitaj (1993) cases to demonstrate the sharpness of the structural approximation gain.

math.NT

From ABC to Effective Roth and Ridout Constants for Cubic Roots

Enrico Bombieri showed conditionally (1994) that the ABC conjecture implies Roth's theorem, and Van Frankenhuysen (1999) later provided a complete proof. Building on Bombieri's and Van der Poorten's explicit formula for continued-fraction coefficients of algebraic numbers (specialized to cubic roots) we derive an effective bound for a Roth-type constant assuming an effective form of ABC. Roth's original argument establishes existence but does not yield an explicit value; our approach makes the dependence on the ABC parameters explicit and also gives an explicit bound in the corresponding special case of Ridout's theorem. We then introduce the notion of approximation gain as a refinement of the quality of an abc-triple. For c in a large computational range, the approximation gain remains below a strikingly small threshold, motivating the conjecture that the approximation gain is always smaller than 1.5. This suggests a potential strategy for attacking ABC by bounding approximation gain and power gain separately.

math.NT

Optimizing Federated Learning by Entropy-Based Client Selection

Although deep learning has revolutionized domains such as natural language processing and computer vision, its dependence on centralized datasets raises serious privacy concerns. Federated learning addresses this issue by enabling multiple clients to collaboratively train a global deep learning model without compromising their data privacy. However, the performance of such a model degrades under label skew, where the label distribution differs between clients. To overcome this issue, a novel method called FedEntOpt is proposed. In each round, it selects clients to maximize the entropy of the aggregated label distribution, ensuring that the global model is exposed to data from all available classes. Extensive experiments on multiple benchmark datasets show that the proposed method outperforms several state-of-the-art algorithms by up to 6% in classification accuracy under standard settings regardless of the model size, while achieving gains of over 30% in scenarios with low participation rates and client dropout. In addition, FedEntOpt offers the flexibility to be combined with existing algorithms, enhancing their classification accuracy by more than 40%. Importantly, its performance remains unaffected even when differential privacy is applied.

cs.LG

Efficient Federated Learning Tiny Language Models for Mobile Network Feature Prediction

In telecommunications, Autonomous Networks (ANs) automatically adjust configurations based on specific requirements (e.g., bandwidth) and available resources. These networks rely on continuous monitoring and intelligent mechanisms for self-optimization, self-repair, and self-protection, nowadays enhanced by Neural Networks (NNs) to enable predictive modeling and pattern recognition. Here, Federated Learning (FL) allows multiple AN cells - each equipped with NNs - to collaboratively train models while preserving data privacy. However, FL requires frequent transmission of large neural data and thus an efficient, standardized compression strategy for reliable communication. To address this, we investigate NNCodec, a Fraunhofer implementation of the ISO/IEC Neural Network Coding (NNC) standard, within a novel FL framework that integrates tiny language models (TLMs) for various mobile network feature prediction (e.g., ping, SNR or band frequency). Our experimental results on the Berlin V2X dataset demonstrate that NNCodec achieves transparent compression (i.e., negligible performance loss) while reducing communication overhead to below 1%, showing the effectiveness of combining NNC with FL in collaboratively learned autonomous mobile networks.

cs.LG

Evaluation of Torque Ripple and Tooth Forces of a Skewed PMSM by 2D and 3D FE Simulations

In this paper, various skewing configurations for a permanent magnet synchronous machine are evaluated by comparing torque ripple amplitudes and tooth forces. Since high-frequency pure tones emitted by an electrical machine significantly impact a vehicle's noise, vibration, and harshness (NVH) behavior, it is crucial to analyze radial forces. These forces are examined and compared across different skewing configurations and angles using the Maxwell stress tensor in 2D and 3D finite-element (FE) simulations. In addition to conventional investigations in 2D FE simulations, 3D FE simulations are executed. These 3D FE simulations show that axial forces occur at the transition points between the magnetic segments of a linear step skewed rotor.

eess.SY

A Privacy Preserving System for Movie Recommendations Using Federated Learning

Recommender systems have become ubiquitous in the past years. They solve the tyranny of choice problem faced by many users, and are utilized by many online businesses to drive engagement and sales. Besides other criticisms, like creating filter bubbles within social networks, recommender systems are often reproved for collecting considerable amounts of personal data. However, to personalize recommendations, personal information is fundamentally required. A recent distributed learning scheme called federated learning has made it possible to learn from personal user data without its central collection. Consequently, we present a recommender system for movie recommendations, which provides privacy and thus trustworthiness on multiple levels: First and foremost, it is trained using federated learning and thus, by its very nature, privacy-preserving, while still enabling users to benefit from global insights. Furthermore, a novel federated learning scheme, called FedQ, is employed, which not only addresses the problem of non-i.i.d.-ness and small local datasets, but also prevents input data reconstruction attacks by aggregating client updates early. Finally, to reduce the communication overhead, compression is applied, which significantly compresses the exchanged neural network parametrizations to a fraction of their original size. We conjecture that this may also improve data privacy through its lossy quantization stage.

cs.IR

Characteristic Function, Schur Interpolation Problem and Darlington Synthesis

In this paper we would like to show the interrelation between the different mathematical theories concerning the Schur interpolation problem, contractions in Hilbert spaces, pseudocontinuation and Darlington synthesis. The main objects of this article are contractive functions holomorphic in the unit disc (Schur functions). Here they are considered, on the one hand, as characteristic functions of contractions in Hilbert spaces and, on the other hand, as transfer functions of open systems.

math.CV

Characteristic Function, Schur Parameters and Pseudocontinuation of Schur functions

In [19] there is an approach to the investigation of the pseudocontinuability of Schur functions in terms of Schur parameters. In particular, there was obtained a criterion for the pseudocontinuability of Schur functions and the Schur parameters of rational Schur functions were described. This approach is based on the description in terms of the Schur parameters of the relative position of the largest shift and the largest coshift in a completely nonunitary contraction. It should be mentioned that these results received a further development in [8, 21-24]. This paper is aimed to give a survey about essential results on this direction. The main object in the approach is based on considering a Schur function as characteristic function of a contraction (see Section 1.2). This enables us outgoing from Schur parameters to construct a model of the corresponding contraction (see Section 2). In this model, the relative position of the largest shift and the largest coshift in a completely nonunitary contraction is described in Section 3 and then, based on this model, to find characteristics which are responsible for the pseudocontinuability of Schur functions (see Sections 4 and 5). The further parts of this paper (see Sections 6-8) admit applications of the above results to the study of properties of Schur functions and questions related with them.

math.CV

Characteristic function of M. S. Livšic and triangular models of bounded linear operators

This paper is dedicated to the introduction in a circle of ideas and methods, which are connected with the notion of characteristic function of a non-selfadjoint operator. We start with the consideration of closed and open systems (Subsections 2.1.1-2.1.2). In Subsections 2.1.2-2.1.3 we introduce the notion of operator colligation and define the characteristic function of the operator colligation as transfer function of the corresponding open system. In Section 3 we state three basic properties of the c.o.f.. First (Subsection 3.1), we note that the c.o.f. is the full unitary invariant of the operator colligation. Second (see Theorem 3.4), it turns out that the invariant subspaces of the corresponding operator are associated with left divisors of the c.o.f.. Third, the $J$-property of the c.o.f. (see (3.6)-(3.8)) is a basic property which determines the class of c. o. f. (see Section 4). In Chapter 4 we describe the classes of characteristic functions which play an important role in our considerations. In Chapter 5 we state necessary facts on multiplicative integral. Chapter 6 is devoted to the factorization theorem (Theorem 6.7) for matrix-valued characteristic function. In Chapter 7 we construct a triangular Livšic model of bounded linear operator and as application we obtain some known results on dissipative operators.

math.CV

Roth's Theorem implies a Weakened Version of the ABC Conjecture for Special Cases

Enrico Bombieri proved that the ABC Conjecture implies Roth's theorem in 1994. This paper concerns the other direction. In making use of Bombieri's and Van der Poorten's explicit formula for the coefficients of the regular continued fractions of algebraic numbers, we prove that Roth's theorem implies a weakened non-effective version of the ABC Conjecture in certain cases relating to roots.

math.NT

Do algebraic numbers follow Khinchin's Law?

The coefficients of the regular continued fraction for random numbers are distributed by the Gauss-Kuzmin distribution according to Khinchin's law. Their geometric mean converges to Khinchin's constant and their rational approximation speed is Khinchin's speed. It is an open question whether these theorems also apply to algebraic numbers of degree $>2$. Since they apply to almost all numbers it is, however, commonly inferred that it is most likely that non quadratic algebraic numbers also do so. We argue that this inference is not well grounded. There is strong numerical evidence that Khinchin's speed is too fast. For Khinchin's law and Khinchin's constant the numerical evidence is unclear. We apply the Kullback Leibler Divergence (KLD) to show that the Gauss-Kuzmin distribution does not fit well for algebraic numbers of degree $>2$. Our suggestion to truncate the Gauss-Kuzmin distribution for finite parts fits slightly better but its KLD is still much larger than the KLD of a random number. So, if it converges the convergence is non uniform and each algebraic number has its own bound. We conclude that there is no evidence to apply the theorems that hold for random numbers to algebraic numbers.

math.NT

FedAUXfdp: Differentially Private One-Shot Federated Distillation

Federated learning suffers in the case of non-iid local datasets, i.e., when the distributions of the clients' data are heterogeneous. One promising approach to this challenge is the recently proposed method FedAUX, an augmentation of federated distillation with robust results on even highly heterogeneous client data. FedAUX is a partially $(ε, δ)$-differentially private method, insofar as the clients' private data is protected in only part of the training it takes part in. This work contributes a fully differentially private modification, termed FedAUXfdp. We further contribute an upper bound on the $l_2$-sensitivity of regularized multinomial logistic regression. In experiments with deep networks on large-scale image datasets, FedAUXfdp with strong differential privacy guarantees performs significantly better than other equally privatized SOTA baselines on non-iid client data in just a single communication round. Full privatization of the modified method results in a negligible reduction in accuracy at all levels of data heterogeneity.

cs.LG

Adaptive Differential Filters for Fast and Communication-Efficient Federated Learning

Federated learning (FL) scenarios inherently generate a large communication overhead by frequently transmitting neural network updates between clients and server. To minimize the communication cost, introducing sparsity in conjunction with differential updates is a commonly used technique. However, sparse model updates can slow down convergence speed or unintentionally skip certain update aspects, e.g., learned features, if error accumulation is not properly addressed. In this work, we propose a new scaling method operating at the granularity of convolutional filters which 1) compensates for highly sparse updates in FL processes, 2) adapts the local models to new data domains by enhancing some features in the filter space while diminishing others and 3) motivates extra sparsity in updates and thus achieves higher compression ratios, i.e., savings in the overall data transfer. Compared to unscaled updates and previous work, experimental results on different computer vision tasks (Pascal VOC, CIFAR10, Chest X-Ray) and neural networks (ResNets, MobileNets, VGGs) in uni-, bidirectional and partial update FL settings show that the proposed method improves the performance of the central server model while converging faster and reducing the total amount of transmitted data by up to 377 times.

cs.LG

ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs

The remarkable success of deep neural networks (DNNs) in various applications is accompanied by a significant increase in network parameters and arithmetic operations. Such increases in memory and computational demands make deep learning prohibitive for resource-constrained hardware platforms such as mobile devices. Recent efforts aim to reduce these overheads, while preserving model performance as much as possible, and include parameter reduction techniques, parameter quantization, and lossless compression techniques. In this chapter, we develop and describe a novel quantization paradigm for DNNs: Our method leverages concepts of explainable AI (XAI) and concepts of information theory: Instead of assigning weight values based on their distances to the quantization clusters, the assignment function additionally considers weight relevances obtained from Layer-wise Relevance Propagation (LRP) and the information content of the clusters (entropy optimization). The ultimate goal is to preserve the most relevant weights in quantization clusters of highest information content. Experimental results show that this novel Entropy-Constrained and XAI-adjusted Quantization (ECQ$^{\text{x}}$) method generates ultra low-precision (2-5 bit) and simultaneously sparse neural networks while maintaining or even improving model performance. Due to reduced parameter precision and high number of zero-elements, the rendered networks are highly compressible in terms of file size, up to $103\times$ compared to the full-precision unquantized DNN model. Our approach was evaluated on different types of models and datasets (including Google Speech Commands, CIFAR-10 and Pascal VOC) and compared with previous work.

cs.LG