Searcharxiv⌕ Search

arXiv subjects

Neha Gupta

Publications and source records attributed to Neha Gupta.

At least 37 records · Page 2Linked to original sources

Hawkes process with tempered Mittag-Leffler kernel

In this paper, we propose an extension of the Hawkes process by incorporating a kernel based on the tempered Mittag-Leffler distribution. This is the generalization of the work presented in [10]. We derive analytical results for the expectation of the conditional intensity and the expected number of events in the counting process. Additionally, we investigate the limiting behavior of the expectation of the conditional intensity. Finally, we present an empirical comparison of the studied process with its limiting special cases.

math.PR↗

Tempered Fractional Hawkes Process and Its Generalization

Hawkes process (HP) is a point process with a conditionally dependent intensity function. This paper defines the tempered fractional Hawkes process (TFHP) by time-changing the HP with an inverse tempered stable subordinator. We obtained results that generalize the fractional Hawkes process defined in Hainaut (2020) to a tempered version which has \textit{semi-heavy tailed} decay. We derive the mean, the variance, covariance and the governing fractional difference-differential equations of the TFHP. Additionally, we introduce the generalized fractional Hawkes process (GFHP) by time-changing the HP with the inverse Lévy subordinator. This definition encompasses all potential (inverse Lévy) time changes as specific instances. We also explore the distributional characteristics and the governing difference-differential equation of the one-dimensional distribution for the GFHP.

math.PR↗

Language Model Cascades: Token-level uncertainty and beyond

Recent advances in language models (LMs) have led to significant improvements in quality on complex NLP tasks, but at the expense of increased inference costs. Cascading offers a simple strategy to achieve more favorable cost-quality tradeoffs: here, a small model is invoked for most "easy" instances, while a few "hard" instances are deferred to the large model. While the principles underpinning cascading are well-studied for classification tasks - with deferral based on predicted class uncertainty favored theoretically and practically - a similar understanding is lacking for generative LM tasks. In this work, we initiate a systematic study of deferral rules for LM cascades. We begin by examining the natural extension of predicted class uncertainty to generative LM tasks, namely, the predicted sequence uncertainty. We show that this measure suffers from the length bias problem, either over- or under-emphasizing outputs based on their lengths. This is because LMs produce a sequence of uncertainty values, one for each output token; and moreover, the number of output tokens is variable across examples. To mitigate this issue, we propose to exploit the richer token-level uncertainty information implicit in generative LMs. We argue that naive predicted sequence uncertainty corresponds to a simple aggregation of these uncertainties. By contrast, we show that incorporating token-level uncertainty through learned post-hoc deferral rules can significantly outperform such simple aggregation strategies, via experiments on a range of natural language benchmarks with FLAN-T5 models. We further show that incorporating embeddings from the smaller model and intermediate layers of the larger model can give an additional boost in the overall cost-quality tradeoff.

cs.CL↗

Properties of graphs of neural codes

A neural code on $ n $ neurons is a collection of subsets of the set $ [n]=\{1,2,\dots,n\} $. In this paper, we study some properties of graphs of neural codes. In particular, we study codeword containment graph (CCG) given by Chan et al. (SIAM J. on Dis. Math., 37(1):114-145,2017) and general relationship graph (GRG) given by Gross et al. (Adv. in App. Math., 95:65-95, 2018). We provide a sufficient condition for CCG to be connected. We also show that the connectedness and completeness of CCG are preserved under surjective morphisms between neural codes defined by A. Jeffs (SIAM J. on App. Alg. and Geo., 4(1):99-122,2020). Further, we show that if CCG of any neural code $\mathcal{C}$ is complete with $|\mathcal{C}|=m$, then $\mathcal{C} \cong \{\emptyset,1,12,\dots,123\cdots m\}$ as neural codes. We also prove that a code whose CCG is complete is open convex. Later, we show that if a code $\mathcal{C}$ with $|\mathcal{C}|>3$ has its CCG to be connected 2-regular then $|\mathcal{C}| $ is even. The GRG was defined only for degree two neural codes using the canonical forms of its neural ideal. We first define GRG for any neural code. Then, we show the behaviour of GRGs under the various elementary code maps. At last, we compare these two graphs for certain classes of codes and see their properties.

math.CO↗

Neural category

A neural code on $ n $ neurons is a collection of subsets of the set $ [n]=\{1,2,\dots,n\} $. Curto et al. \cite{curto2013neural} associated a ring $\mathcal{R}_{\mathcal{C}}$ (neural ring) to a neural code $\mathcal{C}$. A special class of ring homomorphisms between two neural rings, called neural ring homomorphism, was introduced by Curto and Youngs \cite{curto2020neural}. The main work in this paper comprises constructing two categories. First is the $\mathfrak{C}$ category, a subcategory of SETS consisting of neural codes and code maps. Second is the neural category $\mathfrak{N}$, a subcategory of \textit{Rngs} consisting of neural rings and neural ring homomorphisms. Then, the rest of the paper characterizes the properties of these two categories like initial and final objects, products, coproducts, limits, etc. Also, we show that these two categories are in dual equivalence.

math.CT↗

When Does Confidence-Based Cascade Deferral Suffice?

Cascades are a classical strategy to enable inference cost to vary adaptively across samples, wherein a sequence of classifiers are invoked in turn. A deferral rule determines whether to invoke the next classifier in the sequence, or to terminate prediction. One simple deferral rule employs the confidence of the current classifier, e.g., based on the maximum predicted softmax probability. Despite being oblivious to the structure of the cascade -- e.g., not modelling the errors of downstream models -- such confidence-based deferral often works remarkably well in practice. In this paper, we seek to better understand the conditions under which confidence-based deferral may fail, and when alternate deferral strategies can perform better. We first present a theoretical characterisation of the optimal deferral rule, which precisely characterises settings under which confidence-based deferral may suffer. We then study post-hoc deferral mechanisms, and demonstrate they can significantly improve upon confidence-based deferral in settings where (i) downstream models are specialists that only work well on a subset of inputs, (ii) samples are subject to label noise, and (iii) there is distribution shift between the train and test set.

cs.LG↗

Neural ring homomorphism preserves mandatory sets required for open convexity

It has been studied by Curto et al. (SIAM J. on App. Alg. and Geom., 1(1) : 222 $\unicode{x2013}$ 238, 2017) that a neural code that has an open convex realization does not have any local obstruction relative to the neural code. Further, a neural code $ \mathcal{C} $ has no local obstructions if and only if it contains the set of mandatory codewords, $ \mathcal{C}_{\min}(Δ),$ which depends only on the simplicial complex $Δ=Δ(\mathcal{C})$. Thus if $\mathcal{C} \not \supseteq \mathcal{C}_{\min}(Δ)$, then $\mathcal{C}$ cannot be open convex. However, the problem of constructing $ \mathcal{C}_{\min}(Δ) $ for any given code $ \mathcal{C} $ is undecidable. There is yet another way to capture the local obstructions via the homological mandatory set, $ \mathcal{M}_H(Δ). $ The significance of $ \mathcal{M}_H(Δ) $ for a given code $ \mathcal{C} $ is that $ \mathcal{M}_H(Δ) \subseteq \mathcal{C}_{\min}(Δ)$ and so $ \mathcal{C} $ will have local obstructions if $ \mathcal{C}\not\supseteq\mathcal{M}_H(Δ). $ In this paper we study the affect on the sets $\mathcal{C}_{\min}(Δ) $ and $\mathcal{M}_H(Δ)$ under the action of various surjective elementary code maps. Further, we study the relationship between Stanley-Reisner rings of the simplicial complexes associated with neural codes of the elementary code maps. Moreover, using this relationship, we give an alternative proof to show that $ \mathcal{M}_H(Δ) $ is preserved under the elementary code maps.

math.AT↗

Neural Codes and Neural ring endomorphisms

We investigate combinatorial, topological and algebraic properties of certain classes of neural codes. We look into a conjecture that states if the minimal \textit{open convex} embedding dimension of a neural code is two then its minimal \textit{convex} embedding dimension is also two. We prove the conjecture for two interesting classes of examples and provide a counterexample for the converse of the conjecture. We introduce a new class of neural codes, \textit{Doublet maximal}. We show that a Doublet maximal code is open convex if and only if it is max-intersection complete. We prove that surjective neural ring homomorphisms preserve max-intersection complete property. We introduce another class of neural codes, \textit{Circulant codes}. We give the count of neural ring endomorphisms for several sub-classes of this class.

math.GT↗

Fractional Generalizations of the Compound Poisson Process

This paper introduces the Generalized Fractional Compound Poisson Process (GFCPP), which claims to be a unified fractional version of the compound Poisson process (CPP) that encompasses existing variations as special cases. We derive its distributional properties, generalized fractional differential equations, and martingale properties. Some results related to the governing differential equation about the special cases of jump distributions, including exponential, Mittag-Leffler, Bernstéin, discrete uniform, truncated geometric, and discrete logarithm. Some of processes in the literature such as the fractional Poisson process of order $k$, Pólya-Aeppli process of order $k$, and fractional negative binomial process becomes the special case of the GFCPP. Classification based on arrivals by time-changing the compound Poisson process by the inverse tempered and the inverse of inverse Gaussian subordinators are studied. Finally, we present the simulation of the sample paths of the above-mentioned processes.

math.PR↗

Asymmetric magnetism at the interfaces of MgO/FeCoB bilayers by exchanging the order of MgO and FeCoB

Interfaces in FeCoB/MgO/FeCoB magnetic tunnel junction play a vital role in controlling their magnetic and transport properties for various applications in spintronics and magnetic recording media. In this work, interface structures of a few nm thick FeCoB layers in FeCoB/MgO and MgO/FeCoB bilayers are comprehensively studied using x-ray standing waves (XSW) generated by depositing bilayers between Pt waveguide structures. High interface selectivity of nuclear resonance scattering (NRS) under the XSW technique allowed measuring structure and magnetism at the two interfaces, namely FeCoB-on-MgO and MgO-on-FeCoB, yielding an interesting result that electron density and hyperfine fields are not symmetric at both interfaces. The formation of a high-density FeCoB layer at the MgO/FeCoB (FeCoB-on-MgO) interface with an increased hyperfine field (~34.65 T) is attributed to the increasing volume of FeCo at the interface due to boron diffusion from 57FeCoB to the MgO layer. Furthermore, it caused unusual angular-dependent magnetic properties in MgO/FeCoB bilayer, whereas FeCoB/MgO is magnetically isotropic. In contrast to the literature, where the unusual angular dependent in FeCoB based system is explained in terms of in-plane magnetic anisotropy, present findings attributed the same to the interlayer exchange coupling between bulk and interface layer within the FeCoB layer.

cond-mat.mtrl-sci↗

Berkovich-Uncu type Partition Inequalities Concerning Impermissible Sets and Perfect Power Frequencies

Recently, Rattan and the first author (Ann. Comb. 25 (2021) 697-728) proved a conjectured inequality of Berkovich and Uncu (Ann. Comb. 23 (2019) 263-284) concerning partitions with an impermissible part. In this article, we generalize this inequality upon considering t impermissible parts. We compare these with partitions whose certain parts appear with a frequency which is a perfect t^{th} power. Our inequalities hold after a certain bound, which for given t is a polynomial in s, a major improvement over the previously known bound in the case t=1. To prove these inequalities, our methods involve constructing injective maps between the relevant sets of partitions. The construction of these maps crucially involves concepts from analysis and calculus, such as explicit maps used to prove countability of N^t, and Jensen's inequality for convex functions, and then merge them with techniques from number theory such as Frobenius numbers, congruence classes, binary numbers and quadratic residues. We also show a connection of our results to colored partitions. Finally, we pose an open problem which seems to be related to power residues and the almost universality of diagonal ternary quadratic forms.

math.CO↗

Ensembling over Classifiers: a Bias-Variance Perspective

Ensembles are a straightforward, remarkably effective method for improving the accuracy,calibration, and robustness of models on classification tasks; yet, the reasons that underlie their success remain an active area of research. We build upon the extension to the bias-variance decomposition by Pfau (2013) in order to gain crucial insights into the behavior of ensembles of classifiers. Introducing a dual reparameterization of the bias-variance tradeoff, we first derive generalized laws of total expectation and variance for nonsymmetric losses typical of classification tasks. Comparing conditional and bootstrap bias/variance estimates, we then show that conditional estimates necessarily incur an irreducible error. Next, we show that ensembling in dual space reduces the variance and leaves the bias unchanged, whereas standard ensembling can arbitrarily affect the bias. Empirically, standard ensembling reducesthe bias, leading us to hypothesize that ensembles of classifiers may perform well in part because of this unexpected reduction.We conclude by an empirical analysis of recent deep learning methods that ensemble over hyperparameters, revealing that these techniques indeed favor bias reduction. This suggests that, contrary to classical wisdom, targeting bias reduction may be a promising direction for classifier ensembles.

stat.ML↗

Understanding the bias-variance tradeoff of Bregman divergences

This paper builds upon the work of Pfau (2013), which generalized the bias variance tradeoff to any Bregman divergence loss function. Pfau (2013) showed that for Bregman divergences, the bias and variances are defined with respect to a central label, defined as the mean of the label variable, and a central prediction, of a more complex form. We show that, similarly to the label, the central prediction can be interpreted as the mean of a random variable, where the mean operates in a dual space defined by the loss function itself. Viewing the bias-variance tradeoff through operations taken in dual space, we subsequently derive several results of interest. In particular, (a) the variance terms satisfy a generalized law of total variance; (b) if a source of randomness cannot be controlled, its contribution to the bias and variance has a closed form; (c) there exist natural ensembling operations in the label and prediction spaces which reduce the variance and do not affect the bias.

stat.ML↗

Densities of Inverse Tempered Stable Subordinators and Related Processes With Mellin Transforrm

In this article, the infinite series form of the probability densities of tempered stable and inverse tempered stable subordinators are obtained using Mellin transform. Further, the densities of the products and quotients of stable and inverse stable subordinators are worked out. The asymptotic behaviours of these densities are obtained as $x \rightarrow 0^+$. Similar results for tempered and inverse tempered stable subordinators are discussed. Our results provide alternative methods to find the densities of these subordinators and complement the results available in literature.

math.PR↗

Fractional Poisson Processes of Order k and Beyond

In this article, we introduce fractional Poisson felds of order k in n-dimensional Euclidean space $R_n^+$. We also work on time-fractional Poisson process of order k, space-fractional Poisson process of order k and tempered version of time-space fractional Poisson process of order k in one dimensional Euclidean space $R_1^+$. These processes are defined in terms of fractional compound Poisson processes. Time-fractional Poisson process of order k naturally generalizes the Poisson process and Poisson process of order k to a heavy tailed waiting times counting process. The space-fractional Poisson process of order k, allows on average infinite number of arrivals in any interval. We derive the marginal probabilities, governing difference-differential equations of the introduced processes. We also provide Watanabe martingale characterization for some time-changed Poisson processes.

math.PR↗

Estimating decision tree learnability with polylogarithmic sample complexity

We show that top-down decision tree learning heuristics are amenable to highly efficient learnability estimation: for monotone target functions, the error of the decision tree hypothesis constructed by these heuristics can be estimated with polylogarithmically many labeled examples, exponentially smaller than the number necessary to run these heuristics, and indeed, exponentially smaller than information-theoretic minimum required to learn a good decision tree. This adds to a small but growing list of fundamental learning algorithms that have been shown to be amenable to learnability estimation. En route to this result, we design and analyze sample-efficient minibatch versions of top-down decision tree learning heuristics and show that they achieve the same provable guarantees as the full-batch versions. We further give "active local" versions of these heuristics: given a test point $x^\star$, we show how the label $T(x^\star)$ of the decision tree hypothesis $T$ can be computed with polylogarithmically many labeled examples, exponentially smaller than the number necessary to learn $T$.

cs.LG↗

Universal guarantees for decision tree induction via a higher-order splitting criterion

We propose a simple extension of top-down decision tree learning heuristics such as ID3, C4.5, and CART. Our algorithm achieves provable guarantees for all target functions $f: \{-1,1\}^n \to \{-1,1\}$ with respect to the uniform distribution, circumventing impossibility results showing that existing heuristics fare poorly even for simple target functions. The crux of our extension is a new splitting criterion that takes into account the correlations between $f$ and small subsets of its attributes. The splitting criteria of existing heuristics (e.g. Gini impurity and information gain), in contrast, are based solely on the correlations between $f$ and its individual attributes. Our algorithm satisfies the following guarantee: for all target functions $f : \{-1,1\}^n \to \{-1,1\}$, sizes $s\in \mathbb{N}$, and error parameters $ε$, it constructs a decision tree of size $s^{\tilde{O}((\log s)^2/ε^2)}$ that achieves error $\le O(\mathsf{opt}_s) + ε$, where $\mathsf{opt}_s$ denotes the error of the optimal size $s$ decision tree. A key technical notion that drives our analysis is the noise stability of $f$, a well-studied smoothness measure.

cs.LG↗