SearcharxivSearch

arXiv subjects

Nabil Mlaiki

Publications and source records attributed to Nabil Mlaiki.

7 recordsLinked to original sources

VORT: Adaptive Power-Law Memory for NLP Transformers

Standard Transformers impose near-exponential decay on the influence of distant tokens, conflicting with the power-law structure of long-range dependencies in natural language. We introduce the \emph{Variable-Order Retention Transformer} (\VORT{}), a memory architecture in which each ingested token is assigned a learnable fractional order \alpha_i\in[\delta,1] that governs a Gr\"unwald--Letnikov power-law retention kernel. Because the fractional weighted sum is non-Markovian, we approximate it through a sum-of-exponentials (SOE) decomposition computed by Gauss--Laguerre quadrature on a Laplace-type integral representation of the kernel weights. Each exponential component admits a one-step Markovian recurrence at O(Sd_v) per step, where S=O(\log(T/\varepsilon)) terms suffice for \varepsilon-uniform accuracy on horizon [1,T]. Retrieval is keyed and associative via a linear-attention accumulator with an exact O(KSd_\phi d_v) -per-step recurrence. Four results are established: (i) an SOE approximation theorem with geometric convergence rate from the analyticity of the integrand after a log-change of variables; (ii) a quantisation bound valid on [\delta,1] with correct analysis near \alpha=0; (iii) a direct L^2 energy argument (Proposition) showing that for \alpha>1/2 any mixture with fixed minimum decay rate \Lambda>0 incurs L^2([1,T]) error at least N_\alpha(T)-C(\Lambda)\to\infty, with the \Lambda-dependence made explicit; and (iv) linear convergence of a gradient plasticity rule under the Polyak--\L{}ojasiewicz condition. Two synthetic experiments confirm the architectural advantage: a Zipf-distributed retrieval benchmark and an entity label-copy task with uniform lag distribution, the latter ruling out prior-matching as an explanation for the power-law kernel's advantage.

cs.LG

An Effective Weight Initialization Method for Deep Learning: Application to Satellite Image Classification

The growing interest in satellite imagery has triggered the need for efficient mechanisms to extract valuable information from these vast data sources, providing deeper insights. Even though deep learning has shown significant progress in satellite image classification. Nevertheless, in the literature, only a few results can be found on weight initialization techniques. These techniques traditionally involve initializing the networks' weights before training on extensive datasets, distinct from fine-tuning the weights of pre-trained networks. In this study, a novel weight initialization method is proposed in the context of satellite image classification. The proposed weight initialization method is mathematically detailed during the forward and backward passes of the convolutional neural network (CNN) model. Extensive experiments are carried out using six real-world datasets. Comparative analyses with existing weight initialization techniques made on various well-known CNN models reveal that the proposed weight initialization technique outperforms the previous competitive techniques in classification accuracy. The complete code of the proposed technique, along with the obtained results, is available at https://github.com/WadiiBoulila/Weight-Initialization

cs.CV

New fixed-circle results related to Fc-contractive and Fc-expanding mappings on metric spaces

The fixed-circle problem is a recent problem about the study of geometric properties of the fixed point set of a self-mapping on metric (resp. generalized metric) spaces. The fixed-disc problem occurs as a natural consequence of this problem. Our aim in this paper, is to investigate new classes of self-mappings which satisfy new specific type of contraction on a metric space. We see that the fixed point set of any member of these classes contains a circle (or a disc) called the fixed circle (resp. fixed disc) of the corresponding self-mapping. For this purpose, we introduce the notions of an $F_{c}$-contractive mapping and an $F_{c}$-expanding mapping. Activation functions with fixed circles (resp. fixed discs) are often seen in the study of neural networks. This shows the effectiveness of our fixed-circle (resp. fixed-disc) results. In this context, our theoretical results contribute to future studies on neural networks.

math.GN

Controlled rectangular metric type spaces and some applications to polynomial equations

In this paper, we introduce a generalization of rectangular $b-$metric spaces, by changing the rectangular inequality as follows \begin{equation*} ρ(x,y)\le θ(x,y,u,v)[ρ(x,u)+ρ(u,v)+ρ(v,y)], \end{equation*}% for all distinct$\ x,y,u,v\in X.$ We prove some fixed-point theorems and we use our results to present a nice application in last section of this paper. Moreover, in the conclusion we present some new open questions.

math.GN

Rectangular metric like type spaces and fixed points

In this paper we introduce the concept of the rectangular metric like spaces, along with its topology and we prove some fixed point theorems under different contraction principles. We introduce the concept of modified metric-like space as well and prove some topological and convergence properties under the symmetric convergence. Some examples are given to illustrate the proven results and enrich the new introduced metric type spaces.

math.GN

Camina triples

In this paper, we study Camina triples. Camina triples are a generalization of Camina pairs. Camina pairs were first introduced in 1978 by A.R. Camina in \cite{camina1}. Camina's work in \cite{camina1} was inspired by the study of Frobenius groups. We show that if $(G,N,M)$ is a Camina triple, then either $G/N$ is a $p$-group, $M$ is nilpotent, or $M$ has a non-trivial nilpotent quotient.

math.GR

A Central series associated with the vanishing off subgroup V(G)

We generalize Lewis's result about a central series associated with the vanishing off subgroup. We write $V_{1}=V(G)$ for the vanishing off subgroup of $G$, and $V_{i}=[V_{i-1},G]$ for the terms in this central series. Lewis proved that there exists a positive integer $n$ such that if $V_{3} < G_{3}$, then $|G:V_{1}|=|G':V_{2}|^{2}=p^{2n}$. Let $D_{3}/V_{3} = C_{G/V_{3}}(G'/V_{3})$. He also showed that if $V_{3} < G_{3}$, then either $|G:D_{3}|=p^{n}$ or $D_{3}=V_{1}$. We show that if $V_{i} <G_{i}$ for $i\ge 4,$ where $G_{i}$ is the $i$-th term in the lower central series of $G$, then $|G_{i-1}:V_{i-1}|=|G:D_{3}|.$

math.GR