SearcharxivSearch

arXiv subjects

Mehmet Aktas

Publications and source records attributed to Mehmet Aktas.

10 recordsLinked to original sources

Reproducing Kernel Hilbert Spaces for Virtual Persistence Diagrams

A persistence diagram is a finite multiset of birth-death pairs representing the lifetimes of topological features across a filtration. Existing functional and kernel representations of persistence diagrams are typically constructed extrinsically through embeddings into auxiliary spaces. For filtrations with finite indexing sets, the associated virtual persistence diagram group obtained by Grothendieck completion of the persistence diagram monoid is a finitely generated lattice. We define a phase map sending each persistence interval to a circular coordinate and a character map aggregating the phases of intervals in a virtual persistence diagram. We introduce heat damping on characters of virtual persistence diagram groups to suppress the unstable frequencies. We derive Lipschitz bounds for the resulting kernels and apply them in a synthetic segmentation experiment.

math.AT

Square Metric Spaces

Product decompositions of metric spaces are built from coordinate maps, but these maps are not part of the resulting metric space. We recover this missing coordinate structure through equivalence relations whose classes are candidate coordinate fibers, and the resulting quotient metrics reconstruct the coordinate factors. This framework characterizes exactly when a metric space admits a finite product or power presentation. We prove an equivalence of categories showing that these equivalence-relation data preserve exactly the ordered coordinate information of power presentations. For spaces with suitable $\ell^\infty$-prime factorizations, we use prime multiplicities to determine the existence and classification of roots. We also study metric spaces satisfying $X\cong X\times_\infty X$, where repeated coordinate splitting gives a family of metric quotients indexed by infinite binary sequences. We prove that these binary tree structures exactly characterize metric spaces satisfying $X\cong X\times_\infty X$. As an application to persistent homology, we show how to recover filtration parameters whose products or powers form a given space of intervals.

math.AT

Higher-order Persistence Diagrams

Many topological data analysis (TDA) pipelines compute large collections of persistence diagrams, yet vectorizations and kernel methods discard the rank-induced implication relations among persistence intervals that are essential for faithful structural comparison and interpretability. We introduce higher-order persistence diagrams, a recursive construction in which containment relations among persistence intervals define higher-order persistence intervals. This construction performs comparison and aggregation directly on persistence diagrams and preserves interval-level structure. We use harmonic analysis to reduce frequency-space evaluations of aggregated diagrams to zeta transforms. This reduction avoids explicit construction of higher-order diagrams and replaces quadratic pair enumeration with nearly linear-time evaluation. Experiments on random network models show substantial speedups over explicit aggregation. Anonymized code is available at https://anonymous.4open.science/r/higher-order-persistence-8201.

cs.CG

Random Walks on Virtual Persistence Diagrams

In the uniformly discrete case of virtual persistence diagram groups $K(X,A)$, we construct a translation-invariant heat semigroup. The kernels are supported on a countable subgroup $H$, and the restriction to $H$ has Fourier exponent $λ_H$ satisfying $λ_H(θ)=\sum_{κ\in H\setminus\{0\}}\bigl(1-\Reθ(κ)\bigr)ν(κ),$ for a symmetric $ν\in\ell^1(H\setminus\{0\})$. This gives a symmetric jump process on $H$. The exponent $λ_H$ determines heat kernels, which define reproducing kernel Hilbert spaces and their associated semimetrics. Convex orders on the mixing measures give monotonicity for the kernels, Hilbert spaces, and semimetrics.

math.PR

Reproducing Kernel Hilbert Spaces on Banach Completions of Virtual Persistence Diagram Groups

Persistent homology maps a simplicial complex filtered by elements in $\mathbb R$ to finite formal sums of elements of $\mathbb R_{\leq}^{2} = \{ (b,d) \in \mathbb R^2 \cup \{ \infty \} \mid b < d \}$ called (finite) persistence diagrams. This map is stable with respect to the $p$--Wasserstein distance for all $p \in \left[1, + \infty \right]$. Bubenik and Elchesen extend the free translation-invariant commutative Lipschitz monoid of finite persistence diagrams $D(X,A) = D(X)/D(A)$ on arbitrary metric pairs $(X,d,A)$ with $A \subset X$ onto the free translation-invariant abelian Lipschitz group of virtual persistence diagrams $K(X,A) = K(X)/K(A)$ as an isometric embedding $D(X,A) \hookrightarrow K(X,A)$ via the Grothendieck group completion. They prove that the $p$-Wasserstein distance is translation invariant on $D(X,A)$ if and only if $p=1$ and define the unique translation-invariant embedding of $W_1[d]$ into $K(X,A)$ as $ρ.$ When $K(X,A)$ is locally compact abelian, translation-invariant kernels can be constructed via positive-definite functions and Bochner's theorem on the Pontryagin dual. We prove that, for the metric topology induced by $ρ$, the group $(K(X,A),ρ)$ is locally compact if and only if it is discrete, equivalently when the pointed metric space $(X/A,d_1,[A])$ is uniformly discrete, and hence this approach fails outside that case. Assuming instead that $(X/A,d_1,[A])$ is separable and not uniformly discrete, we develop a translation-invariant kernel theory for non--locally compact virtual persistence diagram groups. The group $K(X,A)$ embeds isometrically into its canonical Banach-space linearization $B=\widehat V(X,A)\cong\mathcal F(X/A,d_1)$, and each bounded symmetric positive operator $Q\colon B\to B^\ast$ determines a translation-invariant Gaussian kernel $k(x,y)=\exp\!\left(-\tfrac12\,\langle Q(x-y),x-y\rangle_{B,B^\ast}\right).$

math.FA

Controlling Data Access Load in Distributed Systems

Distributed systems store data objects redundantly to balance the data access load over multiple nodes. Load balancing performance depends mainly on 1) the level of storage redundancy and 2) the assignment of data objects to storage nodes. We analyze the performance implications of these design choices by considering four practical storage schemes that we refer to as clustering, cyclic, block and random design. We formulate the problem of load balancing as maintaining the load on any node below a given threshold. Regarding the level of redundancy, we find that the desired load balance can be achieved in a system of $n$ nodes only if the replication factor $d = Ω(\log(n)^{1/3})$, which is a necessary condition for any storage design. For clustering and cyclic designs, $d = Ω(\log(n))$ is necessary and sufficient. For block and random designs, $d = Ω(\log(n))$ is sufficient but unnecessary. Whether $d = Ω(\log(n)^{1/3})$ is sufficient remains open. The assignment of objects to nodes essentially determines which objects share the access capacity on each node. We refer to the number of nodes jointly shared by a set of objects as the \emph{overlap} between those objects. We find that many consistently slight overlaps between the objects (block, random) are better than few but occasionally significant overlaps (clustering, cyclic). However, when the demand is ''skewed beyond a level'' the impact of overlaps becomes the opposite. We derive our results by connecting the load-balancing problem to mathematical constructs that have been used to study other problems. For a class of storage designs containing the clustering and cyclic design, we express load balance in terms of the maximum of moving sums of i.i.d. random variables, which is known as the scan statistic. For random design, we express load balance by using the occupancy metric for random allocation with complexes.

cs.DC

Service Rate Region: A New Aspect of Coded Distributed System Design

Erasure coding has been recognized as a powerful method to mitigate delays due to slow or straggling nodes in distributed systems. This work shows that erasure coding of data objects can flexibly handle skews in the request rates. Coding can help boost the \emph{service rate region}, that is, increase the overall volume of data access requests that the system can handle. This paper aims to postulate the service rate region as an important consideration in the design of erasure-coded distributed systems. We highlight several open problems that can be grouped into two broad threads: 1) characterizing the service rate region of a given code and finding the optimal request allocation, and2) designing the underlying erasure code for a given service rate region. As contributions along the first thread, we characterize the rate regions of maximum-distance-separable, locally repairable, and Simplex codes. We show the effectiveness of hybrid codes that combine replication and erasure coding in terms of code design. We also discover fundamental connections between multi-set batch codes and the problem of maximizing the service rate region.

cs.IT

Network Embedding: on Compression and Learning

Recently, network embedding that encodes structural information of graphs into a vector space has become popular for network analysis. Although recent methods show promising performance for various applications, the huge sizes of graphs may hinder a direct application of existing network embedding method to them. This paper presents NECL, a novel efficient Network Embedding method with two goals. 1) Is there an ideal Compression of a network? 2) Will the compression of a network significantly boost the representation Learning of the network? For the first problem, we propose a neighborhood similarity based graph compression method that compresses the input graph to get a smaller graph without losing any/much information about the global structure of the graph and the local proximity of the vertices in the graph. For the second problem, we use the compressed graph for network embedding instead of the original large graph to bring down the embedding cost. NECL is a general meta-strategy to improve the efficiency of all of the state-of-the-art graph embedding algorithms based on random walks, including DeepWalk and Node2vec, without losing their effectiveness. Extensive experiments on large real-world networks validate the efficiency of NECL method that yields an average improvement of 23 - 57% embedding time, including walking and learning time without decreasing classification accuracy as evaluated on single and multi-label classification tasks on real-world graphs such as DBLP, BlogCatalog, Cora and Wiki.

cs.SI

On the Service Capacity Region of Accessing Erasure Coded Content

Cloud storage systems generally add redundancy in storing content files such that $K$ files are replicated or erasure coded and stored on $N > K$ nodes. In addition to providing reliability against failures, the redundant copies can be used to serve a larger volume of content access requests. A request for one of the files can be either be sent to a systematic node, or one of the repair groups. In this paper, we seek to maximize the service capacity region, that is, the set of request arrival rates for the $K$ files that can be supported by a coded storage system. We explore two aspects of this problem: 1) for a given erasure code, how to optimally split incoming requests between systematic nodes and repair groups, and 2) choosing an underlying erasure code that maximizes the achievable service capacity region. In particular, we consider MDS and Simplex codes. Our analysis demonstrates that erasure coding makes the system more robust to skews in file popularity than simply replicating a file at multiple servers, and that coding and replication together can make the capacity region larger than either alone.

cs.IT

Computing the Braid Monodromy of Completely Reducible $n$-gonal Curves

Braid monodromy is an important tool for computing invariants of curves and surfaces. In this paper, the \emph{rectangular braid diagram (RBD)} method is proposed to compute the braid monodromy of a completely reducible $n$-gonal curve, i.e. the curves in the form $(y-y_1(x))...(y-y_n(x))=0$ where $n\in \mathbb{Z}^{+}$ and $y_i\in \mathbb{C}[x]$. Also, an algorithm is presented to compute the Alexander polynomial of these curve complements using Burau representations of braid groups. Examples for each computation are provided.

math.AT