SearcharxivSearch

arXiv subjects

Amit Goyal

Publications and source records attributed to Amit Goyal.

At least 19 recordsLinked to original sources

SAGE: Scalable Automatic Gating Ensemble for Confident Negative Harvesting in Fraud Detection

Music streaming fraud, where bad actors artificially inflate stream counts to manipulate chart rankings and royalty payments, poses a significant threat to streaming services and legitimate content creators. Traditional fraud detection approaches struggle with a critical challenge: many legitimate edge cases, including super-fans and sleep-music sessions, exhibit activity patterns that closely mimic those of coordinated fraud. We present SAGE, a novel counterfactual-aware negative harvesting approach that combines SimHash-based stratified sampling with a modular gating ensemble for confident negative identification from unlabeled data. Our ensemble architecture employs pluggable statistical gates (currently instantiated with Mahalanobis distance and k-NN density) with configurable voting thresholds enabling adaptive precision-recall trade-offs. This addresses the representation bias problem in Positive-Unlabeled learning by ensuring comprehensive coverage of rare behavioral cohorts through floor-constrained sampling. Evaluation demonstrates strong precision and recall on held-out data. The approach generalizes across fraud detection domains, achieving strong performance on both customer-level and artist-level fraud without modification to the core methodology.

cs.LG

Comments on "Fundamental nature of the self-field critical current in superconductors", arxiv:2409.16758

We provide comments on article "Fundamental nature of the self-field critical current in superconductors" by Talantsev and Tallon, arXiv:2409.16758 and https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4978589 related to the predictions of the key equation proposed in article "Universal self-field critical current for thin-film superconductors," by Talantsev and Tallon, Nat. Commun. 6 (2015) 7820. We respond to claims and assertions made in these articles and also show that the so-called "Universal self-field critical current" and the "fundamental limit" for self-field Jc suggested by Talantsev and Tallon, has been proven incorrect experimentally and this invalidates the key claims and findings stated in both these articles.

cond-mat.supr-con

Comments on "Procedures for proper validation of record critical current density claims" by Chiara Tarantini and David C. Larbalestier

We provide comments on article titled "Procedures for proper validation of record critical current density claims" by Chiara Tarantini and David C. Larbalestier, arXiv:2410.22195. We respond to claims and assertions that the expressions for calculation of $J_{c,mag}$ from magnetic moments proposed in Polichetti et al. 2024, arXiv:2410.09197, are incorrect. We show that the analysis presented in Tarantini and Larbalestier is wrong. We also re-emphasize the expression presented in Polichetti et al. for calculation of $J_{c,mag}$ from magnetic moments in which the pre-factor is dimensionless and has NO units. We make the case for establishment of national user facilities for measurement of transport $J_c(H,T,\theta)$ that are easily accessible in an equitable and rapid manner to enable much needed advances in the field.

cond-mat.supr-con

Physics-based reward driven image analysis in microscopy

The rise of electron microscopy has expanded our ability to acquire nanometer and atomically resolved images of complex materials. The resulting vast datasets are typically analyzed by human operators, an intrinsically challenging process due to the multiple possible analysis steps and the corresponding need to build and optimize complex analysis workflows. We present a methodology based on the concept of a Reward Function coupled with Bayesian Optimization, to optimize image analysis workflows dynamically. The Reward Function is engineered to closely align with the experimental objectives and broader context and is quantifiable upon completion of the analysis. Here, cross-section, high-angle annular dark field (HAADF) images of ion-irradiated $(Y, Dy)Ba_2Cu_3O_{7-\delta}$ thin-films were used as a model system. The reward functions were formed based on the expected materials density and atomic spacings and used to drive multi-objective optimization of the classical Laplacian-of-Gaussian (LoG) method. These results can be benchmarked against the DCNN segmentation. This optimized LoG* compares favorably against DCNN in the presence of the additional noise. We further extend the reward function approach towards the identification of partially-disordered regions, creating a physics-driven reward function and action space of high-dimensional clustering. We pose that with correct definition, the reward function approach allows real-time optimization of complex analysis workflows at much higher speeds and lower computational costs than classical DCNN-based inference, ensuring the attainment of results that are both precise and aligned with the human-defined objectives.

cond-mat.mtrl-sci

Optimal tie-breaking rules

We consider two-player contests with the possibility of ties and study the effect of different tie-breaking rules on effort. For ratio-form and difference-form contests that admit pure-strategy Nash equilibrium, we find that the effort of both players is monotone decreasing in the probability that ties are broken in favor of the stronger player. Thus, the effort-maximizing tie-breaking rule commits to breaking ties in favor of the weaker agent. With symmetric agents, we find that the equilibrium is generally symmetric and independent of the tie-breaking rule. We also study the design of random tie-breaking rules that are ex-ante fair and identify sufficient conditions under which breaking ties before the contest actually leads to greater expected effort than the more commonly observed practice of breaking ties after the contest.

econ.TH

Phase modulated domain walls and dark solitons for surface gravity waves

We report theoretical prediction of exact localized solutions for dynamics of surface gravity waves, at the critical point kh=1.363, modelled by higher-order nonlinear Schrodinger equation. The model possess domain walls (kink solitons) and dark solitons modulated through different phase profiles. The parametric domains are delineated for the existence of soliton solutions. The effect of wave parameters have been discussed on the amplitude of surface gravity waves. Our work is motivated by Tsitoura et al. [1], on experimental and analytical observation of phase domain walls for deep water surface gravity waves modelled by nonlinear Schrodinger equation.

nlin.PS

Chirped nonlinear resonant states in femtosecond fiber optics

We show the existence of nonlinear resonant states in a higher-order nonlinear Schr\"odinger model that appertains to the wave propagation in femtosecond fiber optics, under certain parametric regime. These nonlinear resonant states are analytically illustrated in terms of Gaussian beams, Airy beams, and periodic beams that resulted due to the presence of quadratic, linear, and constant type of `smart' potentials, respectively, of the ensuing model. Interestingly, the nonlinear chirp associated with each of these novel resonant states can be efficiently controlled, by varying the self-steepening term and self-frequency shift. Furthermore, we have conducted numerical experiments corroborative of our analytical predictions.

nlin.PS

Generation and controlling of ultrashort self-similar solitons and rogue waves in inhomogeneous optical waveguide

We present exact bright, dark and rogue soliton solutions of generalized higherorder nonlinear Schrodinger equation, describing the ultrashort beam propagation in tapered waveguide amplifier, via a similarity transformation connected with the constant-coefficient Sasa-Satsuma and Hirota equations. Our exact analysis takes recourse to identify the allowed tapering profile in conjunction with appropriate gain function which corresponds to PT-symmetric waveguide. We extend our analysis to study the effect of tapering profiles and higher-order terms on the evolution of self-similar waves and thus enabling one to control the self-similar wave structure and dynamical behavior.

nlin.PS

Chirped Lambert W-kink solitons of the complex cubic-quintic Ginzburg-Landau equation with intrapulse Raman scattering

In this paper, an exact explicit solution for the complex cubic-quintic Ginzburg-Landau equation is obtained, by using Lambert W function or omega function. More pertinently, we term them as Lambert W-kink-type solitons, begotten under the influence of intrapulse Raman scattering. Parameter domains are delineated in which these optical solitons exit in the ensuing model. We report the effect of model coefficients on the amplitude of Lambert W-kink solitons, which enables us to control efficiently the pulse intensity and hence their subsequent evolution. Also, moving fronts or optical shock-type solitons are obtained as a byproduct of this model. We explicate the mechanism to control the intensity of these fronts, by fine tuning the spectral filtering or gain parameter. It is exhibited that the frequency chirp associated with these optical solitons depends on the intensity of the wave and saturates to a constant value as the retarded time approaches its asymptotic value.

nlin.PS

Controlled self-similar matter waves in PT-symmetric waveguide

We study the dynamics of Bose-Einstein condensate coupled to a waveguide with parity-time symmetric potential in the presence of quadratic-cubic nonlinearity modelled by Gross-Pitaevskii equation with external source. We employ the self-similar technique to obtain matter wave solutions, such as bright, kinktype, rational dark and Lorentzian-type self-similar waves for this model. The dynamical behavior of self-similar matter waves can be controlled through variation of trapping potential, external source and nature of nonlinearities present in the system.

nlin.PS

Deeply Supervised Semantic Model for Click-Through Rate Prediction in Sponsored Search

In sponsored search it is critical to match ads that are relevant to a query and to accurately predict their likelihood of being clicked. Commercial search engines typically use machine learning models for both query-ad relevance matching and click-through-rate (CTR) prediction. However, matching models are based on the similarity between a query and an ad, ignoring the fact that a retrieved ad may not attract clicks, while click models rely on click history, being of limited use for new queries and ads. We propose a deeply supervised architecture that jointly learns the semantic embeddings of a query and an ad as well as their corresponding CTR.We also propose a novel cohort negative sampling technique for learning implicit negative signals. We trained the proposed architecture using one billion query-ad pairs from a major commercial web search engine. This architecture improves the best-performing baseline deep neural architectures by 2\% of AUC for CTR prediction and by statistically significant 0.5\% of NDCG for query-ad matching.

cs.IR

Refutations on "Debunking the Myths of Influence Maximization: An In-Depth Benchmarking Study"

In a recent SIGMOD paper titled "Debunking the Myths of Influence Maximization: An In-Depth Benchmarking Study", Arora et al. [1] undertake a performance benchmarking study of several well-known algorithms for influence maximization. In the process, they contradict several published results, and claim to have unearthed and debunked several "myths" that existed around the research of influence maximization. It is the goal of this article to examine their claims objectively and critically, and refute the erroneous ones. Our investigation discovers that first, the overall experimental methodology in Arora et al. [1] is flawed and leads to scientifically incorrect conclusions. Second, the paper [1] is riddled with issues specific to a variety of influence maximization algorithms, including buggy experiments, and draws many misleading conclusions regarding those algorithms. Importantly, they fail to appreciate the trade-off between running time and solution quality, and did not incorporate it correctly in their experimental methodology. In this article, we systematically point out the issues present in [1] and refute 11 of their misclaims.

cs.SI

Convex Factorization Machine for Regression

We propose the convex factorization machine (CFM), which is a convex variant of the widely used Factorization Machines (FMs). Specifically, we employ a linear+quadratic model and regularize the linear term with the $\ell_2$-regularizer and the quadratic term with the trace norm regularizer. Then, we formulate the CFM optimization as a semidefinite programming problem and propose an efficient optimization procedure with Hazan's algorithm. A key advantage of CFM over existing FMs is that it can find a globally optimal solution, while FMs may get a poor locally optimal solution since the objective function of FMs is non-convex. In addition, the proposed algorithm is simple yet effective and can be implemented easily. Finally, CFM is a general factorization method and can also be used for other factorization problems including including multi-view matrix factorization and tensor completion problems. Through synthetic and movielens datasets, we first show that the proposed CFM achieves results competitive to FMs. Furthermore, in a toxicogenomics prediction task, we show that CFM outperforms a state-of-the-art tensor factorization method.

stat.ML

Viral Marketing Meets Social Advertising: Ad Allocation with Minimum Regret

In this paper, we study the problem of allocating ads to users through the viral-marketing lens. Advertisers approach the host with a budget in return for the marketing campaign service provided by the host. We show that allocation that takes into account the propensity of ads for viral propagation can achieve significantly better performance. However, uncontrolled virality could be undesirable for the host as it creates room for exploitation by the advertisers: hoping to tap uncontrolled virality, an advertiser might declare a lower budget for its marketing campaign, aiming at the same large outcome with a smaller cost. This creates a challenging trade-off: on the one hand, the host aims at leveraging virality and the network effect to improve advertising efficacy, while on the other hand the host wants to avoid giving away free service due to uncontrolled virality. We formalize this as the problem of ad allocation with minimum regret, which we show is NP-hard and inapproximable w.r.t. any factor. However, we devise an algorithm that provides approximation guarantees w.r.t. the total budget of all advertisers. We develop a scalable version of our approximation algorithm, which we extensively test on four real-world data sets, confirming that our algorithm delivers high quality solutions, is scalable, and significantly outperforms several natural baselines.

cs.SI

Few-cycle optical solitary waves in cascaded-quadratic-cubic-quintic nonlinear media

We study the propagation of few-cycle optical solitary waves in a nonlinear media under the combined action of quadratic, cubic and quintic nonlinearities in a large phase-mismatched second harmonic (SHG) process. Exact bright and dark soliton solutions to the nonlinear evolution equation for cascaded quadratic media beyond the slowly varying envelope approximations is reported. The analytical solutions obtained are verified through numerical simulations.

physics.optics

Validating Network Value of Influencers by means of Explanations

Recently, there has been significant interest in social influence analysis. One of the central problems in this area is the problem of identifying influencers, such that by convincing these users to perform a certain action (like buying a new product), a large number of other users get influenced to follow the action. The client of such an application is a marketer who would target these influencers for marketing a given new product, say by providing free samples or discounts. It is natural that before committing resources for targeting an influencer the marketer would be interested in validating the influence (or network value) of influencers returned. This requires digging deeper into such analytical questions as: who are their followers, on what actions (or products) they are influential, etc. However, the current approaches to identifying influencers largely work as a black box in this respect. The goal of this paper is to open up the black box, address these questions and provide informative and crisp explanations for validating the network value of influencers. We formulate the problem of providing explanations (called PROXI) as a discrete optimization problem of feature selection. We show that PROXI is not only NP-hard to solve exactly, it is NP-hard to approximate within any reasonable factor. Nevertheless, we show interesting properties of the objective function and develop an intuitive greedy heuristic. We perform detailed experimental analysis on two real world datasets - Twitter and Flixster, and show that our approach is useful in generating concise and insightful explanations of the influence distribution of users and that our greedy algorithm is effective and efficient with respect to several baselines.

cs.SI

Approximation Analysis of Influence Spread in Social Networks

In the context of influence propagation in a social graph, we can identify three orthogonal dimensions - the number of seed nodes activated at the beginning (known as budget), the expected number of activated nodes at the end of the propagation (known as expected spread or coverage), and the time taken for the propagation. We can constrain one or two of these and try to optimize the third. In their seminal paper, Kempe et al. constrained the budget, left time unconstrained, and maximized the coverage: this problem is known as Influence Maximization. In this paper, we study alternative optimization problems which are naturally motivated by resource and time constraints on viral marketing campaigns. In the first problem, termed Minimum Target Set Selection (or MINTSS for short), a coverage threshold n is given and the task is to find the minimum size seed set such that by activating it, at least n nodes are eventually activated in the expected sense. In the second problem, termed MINTIME, a coverage threshold n and a budget threshold k are given, and the task is to find a seed set of size at most k such that by activating it, at least n nodes are activated, in the minimum possible time. Both these problems are NP-hard, which motivates our interest in their approximation. For MINTSS, we develop a simple greedy algorithm and show that it provides a bicriteria approximation. We also establish a generic hardness result suggesting that improving it is likely to be hard. For MINTIME, we show that even bicriteria and tricriteria approximations are hard under several conditions. However, if we allow the budget to be boosted by a logarithmic factor and allow the coverage to fall short, then the problem can be solved exactly in PTIME. Finally, we show the value of the approximation algorithms, by comparing them against various heuristics.

cs.DM

A Data-Based Approach to Social Influence Maximization

Influence maximization is the problem of finding a set of users in a social network, such that by targeting this set, one maximizes the expected spread of influence in the network. Most of the literature on this topic has focused exclusively on the social graph, overlooking historical data, i.e., traces of past action propagations. In this paper, we study influence maximization from a novel data-based perspective. In particular, we introduce a new model, which we call credit distribution, that directly leverages available propagation traces to learn how influence flows in the network and uses this to estimate expected influence spread. Our approach also learns the different levels of influenceability of users, and it is time-aware in the sense that it takes the temporal nature of influence into account. We show that influence maximization under the credit distribution model is NP-hard and that the function that defines expected spread under our model is submodular. Based on these, we develop an approximation algorithm for solving the influence maximization problem that at once enjoys high accuracy compared to the standard approach, while being several orders of magnitude faster and more scalable.

cs.DB