SearcharxivSearch

arXiv subjects

Sumit Singh

Publications and source records attributed to Sumit Singh.

8 recordsLinked to original sources

LoCCA: Localized Chebyshev Cross Approximation for Kernel Matrix Factorization via Nodal Perturbation Stability

We propose Local Chebyshev Cross-Approximation (LoCCA), a data-driven framework that bridges smooth polynomial interpolation with flexible matrix factorizations. LoCCA dynamically maps ideal grid nodes to their nearest physical neighbors within unstructured point sets. We support this framework with a new perturbation theory, proving that displacing Chebyshev nodes onto physical data points retains mathematical stability and accuracy without explosive error growth. Numerically, LoCCA provides strict, reliable error control even on irregular geometries where standard grid-based methods fail, achieving substantial speedups and near-optimal matrix compression.

math.NA

Continuous Cross Approximation of Matrices Arising Out of Kernel Functions

We propose a residual energy-based framework for constructing low-rank approximations of kernel matrices arising from continuous kernel functions. The method operates in a continuous setting and is based on the adaptive selection of pivot nodes, referred to as \emph{optimal nodes}, which are chosen to minimize the residual energy at each step. This leads to a sequence of rank-$1$ updates of the residual kernel and admits a natural interpretation as a continuous analog of Adaptive Cross Approximation (ACA). From a theoretical perspective, we show that the residual kernels remain in the class of compact operators and that the approximation error is exactly characterized by the residual energy. We provide convergence guarantees showing that the method yields monotonic error reduction under an alignment condition and achieves geometric decay under practically motivated assumptions. Extensive numerical experiments demonstrate that the proposed method achieves approximation errors close to those of the truncated singular value decomposition across a range of kernel functions. The method exhibits strong robustness with respect to sampling and maintains stable performance across different discretizations. Furthermore, the close agreement between the continuous residual energy and the discrete approximation error highlights the consistency of the formulation. These results establish the proposed approach as a theoretically grounded, practically effective continuous counterpart to classical cross-approximation techniques.

math.NA

Rank of Matrices Arising out of Singular Kernel Functions

Kernel functions are frequently encountered in differential equations and machine learning applications. In this work, we study the rank of matrices arising out of the kernel function $K: X \times Y \mapsto \mathbb{R}$, where the sets $X, Y \in \mathbb{R}^d$ are hypercubes that share a boundary. The main contribution of this work is the analysis of the rank of such matrices where the particles (sources/targets) are arbitrarily distributed within these hypercubes. To our knowledge, this is the first work to formally investigate the rank of such matrices for an arbitrary distribution of particles. We model the arbitrary distribution of particles to arise from an underlying random distribution and obtain bounds on the expected rank and variance of the rank of the kernel matrix corresponding to various neighbor interactions. These bounds are useful for understanding the performance and complexity of hierarchical matrix algorithms (especially hierarchical matrices satisfying the weak-admissibility criterion) for an arbitrary distribution of particles. We also present numerical experiments in one-, two-, and three-dimensions, showing the expected rank growth and variance of the rank for different types of interactions. The numerical results, not surprisingly, align with our theoretical predictions.

math.NA

Enhancing Hindi NER in Low Context: A Comparative study of Transformer-based models with vs. without Retrieval Augmentation

One major challenge in natural language processing is named entity recognition (NER), which identifies and categorises named entities in textual input. In order to improve NER, this study investigates a Hindi NER technique that makes use of Hindi-specific pretrained encoders (MuRIL and XLM-R) and Generative Models ( Llama-2-7B-chat-hf (Llama2-7B), Llama-2-70B-chat-hf (Llama2-70B), Llama-3-70B-Instruct (Llama3-70B) and GPT3.5-turbo), and augments the data with retrieved data from external relevant contexts, notably from Wikipedia. We have fine-tuned MuRIL, XLM-R and Llama2-7B with and without RA. However, Llama2-70B, lama3-70B and GPT3.5-turbo are utilised for few-shot NER generation. Our investigation shows that the mentioned language models (LMs) with Retrieval Augmentation (RA) outperform baseline methods that don't incorporate RA in most cases. The macro F1 scores for MuRIL and XLM-R are 0.69 and 0.495, respectively, without RA and increase to 0.70 and 0.71, respectively, in the presence of RA. Fine-tuned Llama2-7B outperforms Llama2-7B by a significant margin. On the other hand the generative models which are not fine-tuned also perform better with augmented data. GPT3.5-turbo adopted RA well; however, Llama2-70B and llama3-70B did not adopt RA with our retrieval context. The findings show that RA significantly improves performance, especially for low-context data. This study adds significant knowledge about how best to use data augmentation methods and pretrained models to enhance NER performance, particularly in languages with limited resources.

cs.CL

Star versions of Lindel\"of spaces

A space $ X $ is said to be set star-Lindel\"{o}f (resp., set strongly star-Lindel\"{o}f) if for each nonempty subset $ A $ of $ X $ and each collection $ \mathcal{U} $ of open sets in $ X $ such that $ \overline{A} \subseteq \bigcup \mathcal{U} $, there is a countable subset $ \mathcal{V}$ of $ \mathcal{U} $ (resp., countable subset $ F $ of $ \overline{A} $) such that $ A \subseteq {\rm St}( \bigcup \mathcal{V}, \mathcal{U})$ (resp., $ A \subseteq {\rm St}( F, \mathcal{U})$). The classes of set star-Lindel\"{o}f spaces and set strongly star-Lindel\"{o}f spaces lie between the class of Lindel\"{o}f spaces and the class of star-Lindel\"{o}f spaces. In this paper, we investigate the relationship among set star-Lindel\"{o}f spaces, set strongly star-Lindel\"{o}f spaces, and other related spaces by providing some suitable examples and study the topological properties of set star-Lindel\"{o}f and set strongly star-Lindel\"{o}f spaces.

math.GM

Star versions of Hurewicz spaces

A space $X$ is said to have the set star Hurewicz property if for each nonempty subset $A$ of $X$ and each sequence $(\mathcal{U}_n: n \in \mathbb{N})$ of sets open in $X$ such that for each $n\in \mathbb N$, $\overline{A} \subset \cup \mathcal{U}_n$, there is a sequence $(\mathcal{V}_n: n \in \mathbb{N})$ such that for each $n \in \mathbb{N}$, $\mathcal{V}_n$ is a finite subset of $\mathcal{U}_n$ and for each $x \in A$, $x \in {\rm St}(\cup \mathcal{V}_n, \mathcal{U}_n)$ for all but finitely many $n$. In this paper, we investigate the relationships among set star Hurewicz, set strongly star Hurewicz and other related covering properties and study the topological properties of these topological spaces.

math.GN

Partial Degree Bounded Edge Packing Problem with Arbitrary Bounds

We study the Partial Degree Bounded Edge Packing (PDBEP) problem introduced in [5] by Zhang. They have shown that this problem is NP-Hard even for uniform degree constraint. They also presented approximation algorithms for the case when all the vertices have degree constraint of 1 and 2 with approximation ratio of 2 and 32=11 respectively. In this work we study general degree constraint case (arbitrary degree constraint for each vertex) and present two combinatorial approximation algorithms with approximation factors 4 and 2. We also study integer program based solution and present an iterative rounding algorithm with approximation factor 3/(1 - ε)^2 for any positive ε. Next we study the same problem with weighted edges. In this case we present an O(log n) approximation algorithm. Zhang has given an exact O(n^2) complexity algorithm for trees in case of uniform degree constraint. We improve their result by giving O(nlog n) complexity exact algorithm for trees with general degree constraint.

cs.DS

A Local Approach for Identifying Clusters in Networks

Graph clustering is a fundamental problem that has been extensively studied both in theory and practice. The problem has been defined in several ways in literature and most of them have been proven to be NP-Hard. Due to their high practical relevancy, several heuristics for graph clustering have been introduced which constitute a central tool for coping with NP-completeness, and are used in applications of clustering ranging from computer vision, to data analysis, to learning. There exist many methodologies for this problem, however most of them are global in nature and are unlikely to scale well for very large networks. In this paper, we propose two scalable local approaches for identifying the clusters in any network. We further extend one of these approaches for discovering the overlapping clusters in these networks. Some experimentation results obtained for the proposed approaches are also presented.

cs.SI