SearcharxivSearch

arXiv subjects

Arpan Kumar

Publications and source records attributed to Arpan Kumar.

3 recordsLinked to original sources

Predictive Subsampling for Scalable Inference in Networks

Current methods for statistical inference in networks often encounter substantial computational bottlenecks when applied to the massive network datasets that are increasingly common across scientific domains. In this paper, we develop \textit{Predictive Subsampling} (\texttt{PredSub}), a scalable framework for estimation and two-sample testing in networks. The central idea is to replace a full-sample estimation procedure by estimation on a random subsample, followed by out-of-sample prediction of the remaining vertices. This construction exploits the fact that the subsample provides an inferential anchor, which enables each remaining vertex to be incorporated in a \textit{predictive} manner through a fast vector operation. Building on this estimator, we develop two scalable procedures for two-sample testing, namely \texttt{PredSubTest} and \texttt{PureSubTest}. We establish finite-sample error bounds as well as estimation and testing consistency of the proposed methods in both Frobenius and two-to-infinity norms. These results formally characterize the trade-offs between statistical accuracy and computational efficiency with respect to the subsample size, the choice of test statistic, and the choice of norm. We demonstrate the empirical performance of the proposed methods through detailed simulation studies and two real-world applications involving DBLP coauthorship networks and the Cannes 2013 social media networks.

stat.ME

Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics

Large language models (LLMs) are increasingly utilized for machine translation, yet their predictions often exhibit uncertainties that hinder interpretability and user trust. Effectively visualizing these uncertainties can enhance the usability of LLM outputs, particularly in contexts where translation accuracy is critical. This paper addresses two primary objectives: (1) providing users with token-level insights into model confidence and (2) developing a web-based visualization tool to quantify and represent translation uncertainties. To achieve these goals, we utilized the T5 model with the WMT19 dataset for translation tasks and evaluated translation quality using established metrics such as BLEU, METEOR, and ROUGE. We introduced three novel uncertainty quantification (UQ) metrics: (1) the geometric mean of token probabilities, (2) the arithmetic mean of token probabilities, and (3) the arithmetic mean of the kurtosis of token distributions. These metrics provide a simple yet effective framework for evaluating translation performance. Our analysis revealed a linear relationship between the traditional evaluation metrics and our UQ metrics, demonstrating the validity of our approach. Additionally, we developed an interactive web-based visualization that uses a color gradient to represent token confidence. This tool offers users a clear and intuitive understanding of translation quality while providing valuable insights into model performance. Overall, we show that our UQ metrics and visualization are both robust and interpretable, offering practical tools for evaluating and accessing machine translation systems.

cs.CL

Demand Analysis with a Thin Price Sample

For about 125 items of food, the Consumer Expenditure Survey (CES) schedule of the Indian National Sample Survey asks the interviewer to obtain both quantity and value of household consumption during the reference period from the respondent. This would appear to put a great burden on the respondent. But it is likely that the price usually paid is almost the same within each first stage unit (fsu). The present work proposes a new sampling scheme to estimate demand elasticities of essential food items. While the conventional sampling method used in practice (e.g. in NSS consumer expenditure survey) involves seeking price information from many households sampled from a fsu, the proposed procedure involves only one household chosen randomly from every fsu for price data collection and thus requires much less interview burden. Using unit records for vegetable items in the NSS's 2011-12 CES, our results show that in spite of requiring much less data, the new scheme captures the household food consumption behavior as precisely as before.

stat.AP