Searcharxiv⌕ Search

arXiv subjects

Souvik Ghosh

Publications and source records attributed to Souvik Ghosh.

At least 37 records · Page 2Linked to original sources

LiRank: Industrial Large Scale Ranking Models at LinkedIn

We present LiRank, a large-scale ranking framework at LinkedIn that brings to production state-of-the-art modeling architectures and optimization methods. We unveil several modeling improvements, including Residual DCN, which adds attention and residual connections to the famous DCNv2 architecture. We share insights into combining and tuning SOTA architectures to create a unified model, including Dense Gating, Transformers and Residual DCN. We also propose novel techniques for calibration and describe how we productionalized deep learning based explore/exploit methods. To enable effective, production-grade serving of large ranking models, we detail how to train and compress models using quantization and vocabulary compression. We provide details about the deployment setup for large-scale use cases of Feed ranking, Jobs Recommendations, and Ads click-through rate (CTR) prediction. We summarize our learnings from various A/B tests by elucidating the most effective technical approaches. These ideas have contributed to relative metrics improvements across the board at LinkedIn: +0.5% member sessions in the Feed, +1.76% qualified job applications for Jobs search and recommendations, and +4.3% for Ads CTR. We hope this work can provide practical insights and solutions for practitioners interested in leveraging large-scale deep ranking systems.

cs.LG↗

On best coapproximations in subspaces of diagonal matrices

We characterize the best coapproximation(s) to a given matrix $ T $ out of a given subspace $ \mathbb{Y} $ of the space of diagonal matrices $ \mathcal{D}_n $, by using Birkhoff-James orthogonality techniques and with the help of a newly introduced property, christened the $ * $-Property. We also characterize the coproximinal subspaces and the co-Chebyshev subspaces of $ \mathcal{D}_n $ in terms of the $ * $-Property. We observe that a complete characterization of the best coapproximation problem in $ \ell_{\infty}^n $ follows directly as a particular case of our approach.

math.FA↗

On the best coapproximation problem in $\ell_1^n$

We study the best coapproximation problem in the Banach space $ \ell_1^n, $ by using Birkhoff-James orthogonality techniques. Given a subspace $\mathbb{Y}$ of $\ell_1^n$, we completely identify the elements $x$ in $\ell_1^n,$ for which best coapproximations to $x$ out of $\mathbb{Y}$ exist. The methods developed in this article are computationally effective and it allows us to present an algorithmic approach to the concerned problem. We also identify the coproximinal subspaces and co-Chebyshev subspaces of $\ell_1^n$.

math.FA↗

On isosceles orthogonality and some geometric constants in a normed space

We study the James constant $J(\mathbb{X})$, an important geometric quantity associated with a normed space $ \mathbb{X} $, and explore its connection with isosceles orthogonality $ \perp_I. $ The James constant is defined as $J(\mathbb{X}) := \sup\{\min \{\|x+y\|, \|x-y\|\}: x, y \in \mathbb{X},~ \|x\|=\|y\|=1 \}.$ We prove that if $J(\mathbb{X})$ is attained for unit vectors $x, y \in \mathbb{X},$ then $x\perp_I y.$ We also show that if $\mathbb{X}$ is a two-dimensional polyhedral Banach space then $J(\mathbb{X})$ is always attained at an extreme point $z$ of the unit ball of $\mathbb{X},$ so that $J(\mathbb{X}) = \|z+y\| = \|z-y\|,$ where $ \| y \| = 1 $ and $z\perp_I y.$ This helps us to explicitly compute the James constant of a two-dimensional polyhedral Banach space in an efficient way. We further study some related problems with reference to several other geometric constants in a normed space.

math.FA↗

On some special subspaces of a Banach space, from the perspective of best coapproximation

We study the best coapproximation problem in Banach spaces, by using Birkhoff-James orthogonality techniques. We introduce two special types of subspaces, christened the anti-coproximinal subspaces and the strongly anti-coproximinal subspaces. We obtain a necessary condition for the strongly anti-coproximinal subspaces in a reflexive Banach space whose dual space satisfies the Kadets-Klee Property. On the other hand, we provide a sufficient condition for the strongly anti-coproximinal subspaces in a general Banach space. We also characterize the anti-coproximinal subspaces of a smooth Banach space. Further, we study these special subspaces in a finite-dimensional polyhedral Banach space and find some interesting geometric structures associated with them.

math.FA↗

On T-orthogonality in Banach spaces

Let $\mathbb{X}$ be a Banach space and let $\mathbb{X}^*$ be the dual space of $\mathbb{X}.$ For $x,y \in \mathbb{X},$ $ x$ is said to be $T$-orthogonal to $y$ if $Tx(y) =0,$ where $T$ is a bounded linear operator from $\mathbb{X}$ to $\mathbb{X}^*.$ We study the notion of $T$-orthogonality in a Banach space and investigate its relation with the various geometric properties, like strict convexity, smoothness, reflexivity of the space. We explore the notions of left and right symmetric elements w.r.t. the notion of $T$-orthogonality. We characterize bounded linear operators on $\mathbb{X}$ preserving $T$-orthogonality. Finally we characterize Hilbert spaces among all Banach spaces using $T$-orthogonality. \end{abstract}

math.FA↗

LiGNN: Graph Neural Networks at LinkedIn

In this paper, we present LiGNN, a deployed large-scale Graph Neural Networks (GNNs) Framework. We share our insight on developing and deployment of GNNs at large scale at LinkedIn. We present a set of algorithmic improvements to the quality of GNN representation learning including temporal graph architectures with long term losses, effective cold start solutions via graph densification, ID embeddings and multi-hop neighbor sampling. We explain how we built and sped up by 7x our large-scale training on LinkedIn graphs with adaptive sampling of neighbors, grouping and slicing of training data batches, specialized shared-memory queue and local gradient optimization. We summarize our deployment lessons and learnings gathered from A/B test experiments. The techniques presented in this work have contributed to an approximate relative improvements of 1% of Job application hearing back rate, 2% Ads CTR lift, 0.5% of Feed engaged daily active users, 0.2% session lift and 0.1% weekly active user lift from people recommendation. We believe that this work can provide practical solutions and insights for engineers who are interested in applying Graph neural networks at large scale.

cs.LG↗

Unavoidable emergent biaxiality in chiral molecular-colloidal hybrid liquid crystals

Chiral nematic or cholesteric liquid crystals (LCs) are mesophases with long-ranged orientational order featuring a quasi-layered periodicity imparted by a helical configuration but lacking positional order. Doping molecular cholesteric LCs with thin colloidal rods with a large length-to-width ratio or disks with a large diameter-to-thickness ratio adds another level of complexity to the system because of the interplay between weak surface boundary conditions and bulk-based elastic distortions around the particle-LC interface. By using colloidal disks and rods with different geometric shapes and boundary conditions, we demonstrate that these anisotropic colloidal inclusions exhibit biaxial orientational probability distributions, where they tend to orient with the long rod axes and disk normals perpendicular to the helix axis, thus imparting strong local biaxiality on the hybrid cholesteric LC structure. Unlike the situation in achiral hybrid molecular-colloidal LCs, where biaxial order emerges only at modest to high volume fractions of the anisotropic colloidal particles, the orientational probability distribution of colloidal inclusions immersed in chiral nematic hosts are unavoidably biaxial even at vanishingly low particle volume fractions. In addition, the colloidal inclusions induce local biaxiality in the molecular orientational order of the LC host medium, which enhances the weak biaxiality of the LC in a chiral nematic phase coming from the symmetry breaking caused by the presence of the helical axis. With analytical modeling and computer simulations based on minimizing the Landau de Gennes free energy of the host LC around the colloidals, we explain our experimental findings and conclude that the biaxial order of chiral molecular-colloidal LCs is strongly enhanced as compared to both achiral molecular-colloidal LCs and molecular cholesteric LCs and is rather unavoidable.

cond-mat.soft↗

Adaptive Rate of Convergence of Thompson Sampling for Gaussian Process Optimization

We consider the problem of global optimization of a function over a continuous domain. In our setup, we can evaluate the function sequentially at points of our choice and the evaluations are noisy. We frame it as a continuum-armed bandit problem with a Gaussian Process prior on the function. In this regime, most algorithms have been developed to minimize some form of regret. In this paper, we study the convergence of the sequential point $x^t$ to the global optimizer $x^*$ for the Thompson Sampling approach. Under some assumptions and regularity conditions, we prove concentration bounds for $x^t$ where the probability that $x^t$ is bounded away from $x^*$ decays exponentially fast in $t$. Moreover, the result allows us to derive adaptive convergence rates depending on the function structure.

stat.ML↗

Growth of Common Friends in a Preferential Attachment Model

The number of common friends (or connections) in a graph is a commonly used measure of proximity between two nodes. Such measures are used in link prediction algorithms and recommendation systems in large online social networks. We obtain the rate of growth of the number of common friends in a linear preferential attachment model. We apply our result to develop an estimate for the number of common friends. We also observe a phase transition in the limiting behavior of the number of common friends; depending on the range of the parameters of the model, the growth is either power-law, or, logarithmic, or static with the size of the graph.

math.PR↗

Measuring Long-term Impact of Ads on LinkedIn Feed

Organic updates (from a member's network) and sponsored updates (or ads, from advertisers) together form the newsfeed on LinkedIn. The newsfeed, the default homepage for members, attracts them to engage, brings them value and helps LinkedIn grow. Engagement and Revenue on feed are two critical, yet often conflicting objectives. Hence, it is important to design a good Revenue-Engagement Tradeoff (RENT) mechanism to blend ads in the feed. In this paper, we design experiments to understand how members' behavior evolve over time given different ads experiences. These experiences vary on ads density, while the quality of ads (ensured by relevance models) is held constant. Our experiments have been conducted on randomized member buckets and we use two experimental designs to measure the short term and long term effects of the various treatments. Based on the first three months' data, we observe that the long term impact is at a much smaller scale than the short term impact in our application. Furthermore, we observe different member cohorts (based on user activity level) adapt and react differently over time.

cs.SI↗

Global labor flow network reveals the hierarchical organization and dynamics of geo-industrial clusters in the world economy

Groups of firms often achieve a competitive advantage through the formation of geo-industrial clusters. Although many exemplary clusters, such as Hollywood or Silicon Valley, have been frequently studied, systematic approaches to identify and analyze the hierarchical structure of the geo-industrial clusters at the global scale are rare. In this work, we use LinkedIn's employment histories of more than 500 million users over 25 years to construct a labor flow network of over 4 million firms across the world and apply a recursive network community detection algorithm to reveal the hierarchical structure of geo-industrial clusters. We show that the resulting geo-industrial clusters exhibit a stronger association between the influx of educated-workers and financial performance, compared to existing aggregation units. Furthermore, our additional analysis of the skill sets of educated-workers supplements the relationship between the labor flow of educated-workers and productivity growth. We argue that geo-industrial clusters defined by labor flow provide better insights into the growth and the decline of the economy than other common economic units.

cs.SI↗

Testing for arbitrary interference on experimentation platforms

Experimentation platforms are essential to modern large technology companies, as they are used to carry out many randomized experiments daily. The classic assumption of no interference among users, under which the outcome of one user does not depend on the treatment assigned to other users, is rarely tenable on such platforms. Here, we introduce an experimental design strategy for testing whether this assumption holds. Our approach is in the spirit of the Durbin-Wu-Hausman test for endogeneity in econometrics, where multiple estimators return the same estimate if and only if the null hypothesis holds. The design that we introduce makes no assumptions on the interference model between units, nor on the network among the units, and has a sharp bound on the variance and an implied analytical bound on the type I error rate. We discuss how to apply the proposed design strategy to large experimentation platforms, and we illustrate it in the context of an experiment on the LinkedIn platform.

stat.ME↗

Asymptotic Properties of the Empirical Spatial Extremogram

The extremogram, proposed by Davis and Mikosch (2008), is a useful tool for measuring extremal dependence and checking model adequacy in a time series. We define the extremogram in the spatial domain when the data is observed on a lattice or at locations distributed as a Poisson point process in d-dimensional space. Under mixing and other conditions, we establish a central limit theorem for the empirical spatial extremogram. We show these conditions are applicable for max-moving average processes and Brown-Resnick processes and illustrate the empirical extremogram's performance via simulation. We also demonstrate its practical use with a data set related to rainfall in a region in Florida.

math.ST↗

Detecting tail behavior: mean excess plots with confidence bounds

In many practical situations exploratory plots are helpful in understanding tail behavior of sample data. The Mean Excess plot is often applied in practice to understand the right tail behavior of a data set. It is known that if the underlying distribution of a data sample is in the domain of attraction of a Frechet, Gumbel or Weibull distributions then the ME plot of the data tend to a straight line in an appropriate sense, with positive, zero or negative slopes respectively. In this paper we construct confidence intervals around the ME plots which assist us in ascertaining which particular maximum domain of attraction the data set comes from. We recall weak limit results for the Frechet domain of attraction, already obtained in Das and Ghosh (2013) and derive weak limits for the Gumbel and Weibull domains in order to construct confidence bounds. We test our methods on both simulated and real data sets.

math.ST↗

Weak limits for exploratory plots in the analysis of extremes

Exploratory data analysis is often used to test the goodness-of-fit of sample observations to specific target distributions. A few such graphical tools have been extensively used to detect subexponential or heavy-tailed behavior in observed data. In this paper we discuss asymptotic limit behavior of two such plotting tools: the quantile-quantile plot and the mean excess plot. The weak consistency of these plots to fixed limit sets in an appropriate topology of $\mathbb{R}^2$ has been shown in Das and Resnick (Stoch. Models 24 (2008) 103-132) and Ghosh and Resnick (Stochastic Process. Appl. 120 (2010) 1492-1517). In this paper we find asymptotic distributional limits for these plots when the underlying distributions have regularly varying right-tails. As an application we construct confidence bounds around the plots which enable us to statistically test whether the underlying distribution is heavy-tailed or not.

math.ST↗

A functional large and moderate deviation principle for infinitely divisible processes driven by null-recurrent markov chains

Suppose $ E$ is a space with a null-recurrent Markov kernel $ P$. Furthermore, suppose there are infinite particles with variable weights on $ E$ performing a random walk following $ P$. Let $ X_{t}$ be a weighted functional of the position of particles at time $ t$. Under some conditions on the initial distribution of the particles the process $ (X_{t})$ is stationary over time. Non-Gaussian infinitely divisible (ID) distributions turn out to be natural candidates for the initial distribution and then the process $ (X_{t})$ is ID. We prove a functional large and moderate deviation principle for the partial sums of the process $ (X_{t})$. The recurrence of the Markov Kernel $ P$ induces long memory in the process $ (X_{t})$ and that is reflected in the large deviation principle. It has been observed in certain short memory processes that the large deviation principle is very similar to that of an i.i.d. sequence. Whereas, if the process is long range dependent the large deviations change dramatically. We show that a similar phenomenon is observed for infinitely divisible processes driven by Markov chains. Processes of the form of $ (X_{t})$ gives us a rich class of non-Gaussian long memory models which may be useful in practice.

math.PR↗

When does the mean excess plot look linear?

In risk analysis, the mean excess plot is a commonly used exploratory plotting technique for confirming iid data is consistent with a generalized Pareto assumption for the underlying distribution, since in the presence of such a distribution thresholded data have a mean excess plot that is roughly linear. Does any other class of distributions share this linearity of the plot? Under some extra assumptions, we are able to conclude that only the generalized Pareto family has this property.

math.ST↗