SearcharxivSearch

arXiv subjects

Saumitra Kulkarni

Publications and source records attributed to Saumitra Kulkarni.

4 recordsLinked to original sources

The networks of ingredient combinations as culinary fingerprints of world cuisines

Investigating how different ingredients are combined in popular dishes is crucial to uncover the principles behind food preferences. Here, we use data from public food repositories and network analysis to characterize and compare worldwide cuisines. Ingredients are first grouped into broader types, and each cuisine is then represented as a network in which nodes correspond to ingredient types and weighted links describe how frequently pairs of types co-occur in recipes. Cuisines differ not only in the popularity of ingredient types and range of recipe sizes, but also in the structural organization of ingredient-type combinations. By analyzing these networks, we uncover distinctive patterns of type associations that serve as culinary fingerprints. For example, European cuisines typically distribute ingredients across different types, whereas certain Asian and South American traditions emphasize one dominant type, such as vegetables or spices. The essence of these patterns is well captured by the networks' maximum spanning trees, which offer a simplified yet representative backbone for each cuisine. We demonstrate that both these full and simplified network representations enable machine learning models to identify cuisines from subsets of recipes with very high accuracy. Networks of ingredient combinations also cluster global cuisines into meaningful geo-cultural groups, reflecting shared patterns in culinary traditions. More broadly, our study offers novel insights into the structure of world cuisines, enabling data-driven approaches to their characterization, cross-cultural comparison, and potential adaptation.

physics.soc-ph

Investigation of Indian stock markets using topological data analysis and geometry-inspired network measures

Geometry-inspired measures (such as discrete Ricci curvatures) and topological data analysis (TDA) based methods (such as persistent homology) have become attractive tools for characterizing the higher-order structure of networks representing the financial systems. In this study, our goal is to perform a comparative analysis of both these approaches, especially by assessing the fragility and systemic risk in the Indian stock markets, which is known for its high volatility and risk. To achieve this goal, we analyze the time series of daily log-returns of stocks comprising the National Stock Exchange (NSE) and the Bombay Stock Exchange (BSE). Specifically, our aim is to monitor the changes in standard network measures, edge-centric discrete Ricci curvatures, and persistent homology based topological measures computed from cross-correlation matrices of stocks. In this study, the edge-centric discrete Ricci curvatures have been employed for the first time in the analysis of the Indian stock markets. The Indian stock markets are known to be less diverse in comparison to the US market, and hence provides us an interesting example. Our results point that, among the persistent homology based topological measures, persistent entropy is simple and more robust than $L^1$-norm and $L^2$-norm of persistence landscape. In a broader comparison between network analysis and TDA, we highlight that the network analysis is sensitive to the way of constructing the networks (threshold or minimum spanning tree), as well as the threshold values used to construct the correlation-based threshold networks. On the other hand, the persistent homology is a more robust approach and is able to capture the higher-order interactions and eliminate noisy data in financial systems, since it does not take into account a single value of threshold but rather a range of values.

physics.soc-ph

Detrimental role of fluctuations in the resource dependency networks

Individual components of many real-world complex networks produce and exchange resources among themselves. However, because the resource production in such networks is almost always stochastic, fluctuations in the production are unavoidable. In this paper, we study the effect of fluctuations on the resource dependencies in complex networks. To this end, we consider a modification of a threshold model of resource dependencies in networks that was recently proposed, where each vertex has a fitness that depends on the total amount of resource it has produced, the amount it has procured from its neighbours, and the fitness threshold. We study how the ``network fitness'', defined as the average fitness of vertices in the network, is affected as the fluctuation size is varied. We show that the fluctuations worsen the network fitness even when average production on vertices is kept fixed. This is true independent of whether more than required amount is produced in the network or not. However, this effect saturates for large fluctuations, and hence very large fluctuations cannot worsen the network fitness beyond a limit. We further show that the networks with a homogeneous degree distribution, such as the Erdos-Renyi network, are less affected by fluctuations and also produce lower wastage than the networks with a heterogeneous degree distribution like the Scale-Free network. Our work shows that fluctuations in the resource production should be avoided in resource dependency networks.

physics.soc-ph

Vyākarana: A Colorless Green Benchmark for Syntactic Evaluation in Indic Languages

While there has been significant progress towards developing NLU resources for Indic languages, syntactic evaluation has been relatively less explored. Unlike English, Indic languages have rich morphosyntax, grammatical genders, free linear word-order, and highly inflectional morphology. In this paper, we introduce Vyākarana: a benchmark of Colorless Green sentences in Indic languages for syntactic evaluation of multilingual language models. The benchmark comprises four syntax-related tasks: PoS Tagging, Syntax Tree-depth Prediction, Grammatical Case Marking, and Subject-Verb Agreement. We use the datasets from the evaluation tasks to probe five multilingual language models of varying architectures for syntax in Indic languages. Due to its prevalence, we also include a code-switching setting in our experiments. Our results show that the token-level and sentence-level representations from the Indic language models (IndicBERT and MuRIL) do not capture the syntax in Indic languages as efficiently as the other highly multilingual language models. Further, our layer-wise probing experiments reveal that while mBERT, DistilmBERT, and XLM-R localize the syntax in middle layers, the Indic language models do not show such syntactic localization.

cs.CL