SearcharxivSearch

arXiv subjects

Dipankar Das

Publications and source records attributed to Dipankar Das.

At least 19 recordsLinked to original sources

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints

Large Language Models (LLMs) are increasingly described as possessing strong reasoning capabilities, supported by high performance on mathematical, logical, and planning benchmarks. However, most existing evaluations rely on aggregate accuracy over fixed datasets, obscuring how reasoning behavior evolves as task complexity increases. In this work, we introduce a controlled benchmarking framework to systematically evaluate the robustness of reasoning in Large Reasoning Models (LRMs) under progressively increasing problem complexity. We construct a suite of nine classical reasoning tasks: Boolean Satisfiability, Cryptarithmetic, Graph Coloring, River Crossing, Tower of Hanoi, Water Jug, Checker Jumping, Sudoku, and Rubik's Cube, each parameterized to precisely control complexity while preserving underlying semantics. Using deterministic validators, we evaluate multiple open and proprietary LRMs across low, intermediate, and high complexity regimes, ensuring that only fully valid solutions are accepted. Our results reveal a consistent phase transition like behavior: models achieve high accuracy at low complexity but degrade sharply beyond task specific complexity thresholds. We formalize this phenomenon as reasoning collapse. Across tasks, we observe substantial accuracy declines, often exceeding 50%, accompanied by inconsistent reasoning traces, constraint violations, loss of state tracking, and confidently incorrect outputs. Increased reasoning length does not reliably improve correctness, and gains in one problem family do not generalize to others. These findings highlight the need for evaluation methodologies that move beyond static benchmarks and explicitly measure reasoning robustness under controlled complexity.

cs.CL

Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer

Pretraining Large Language Models (LLMs) from scratch requires massive amount of compute. Aurora super computer is an ExaScale machine with 127,488 Intel PVC (Ponte Vechio) GPU tiles. In this work, we showcase LLM pretraining on Aurora at the scale of 1000s of GPU tiles. Towards this effort, we developed Optimus, an inhouse training library with support for standard large model training techniques. Using Optimus, we first pretrained Mula-1B, a 1 Billion dense model and Mula-7B-A1B, a 7 Billion Mixture of Experts (MoE) model from scratch on 3072 GPU tiles for the full 4 trillion tokens of the OLMoE-mix-0924 dataset. We then demonstrated model scaling by pretraining three large MoE models Mula-20B-A2B, Mula-100B-A7B, and Mula-220B-A10B till 100 Billion tokens on the same dataset. On our largest model Mula-220B-A10B, we pushed the compute scaling from 384 to 12288 GPU tiles and observed scaling efficiency of around 90% at 12288 GPU tiles. We significantly improved the runtime performance of MoE models using custom GPU kernels for expert computation, and a novel EP-Aware sharded optimizer resulting in training speedups up to 1.71x. As part of the Optimus library, we also developed a robust set of reliability and fault tolerant features to improve training stability and continuity at scale.

cs.LG

Soft Symmetry Breaking as a Nonstandard Source of Mass: Phenomenological Insights from the Two-Higgs-Doublet Model

The soft-breaking parameter, $m_{12}^2$, frequently appearing in the 2HDM scalar potential is much more remarkable than being just a nonstandard parameter that helps make the BSM scalars super heavy. In fact, as we show through explicit calculations, it should be treated as the direct but concise embodiment of new non-electroweak spontaneous symmetry breaking effects at very high energy scales, wherein lies its quiddities. Consequently, it is argued that $m_{12}^2$ and the electroweak VEV serve as two distinct sources for the nonstandard scalar masses, which are completely unrelated to each other. Such distinctions allow us to define parameters that conveniently capture the fraction of the nonstandard scalar masses derived from the electroweak VEV. Finally, we demonstrate that constraints can already be placed on such fractions from the current measurements of the diphoton signal strength and from direct searches of new nonstandard scalar resonances in the diphoton channel.

hep-ph

Pairwise Beats All-at-Once: Behavioral Gains from Sequential Choice Presentation

This paper presents the Sequential Rationality Hypothesis, which argues that consumers are better able to make utility-maximizing decisions when products appear in sequential pairwise comparisons rather than in simultaneous multi-option displays. Although this involves higher cognitive costs than the all-at-once format, the current digital market, with its diverse products listed by review ratings, pricing, and paid products, often creates inconsistent choices. The present work shows that preparing the list sequentially supports more rational choice, as the consumer tries to minimize cognitive costs and may otherwise make an irrational decision. If the decision remains the same on both offers, then that is a consistent preference. The platform uses this approach by reducing cognitive costs while still providing the list in an all-at-once format rather than sequentially. To show how sequential exposure reduces cognitive overload and prevents context-dependent errors, we develop a bounded attention model and extend the monotonic attention rule of the random attention model to theorize the sequential rational hypothesis. Using a theoretical design with common consumer goods, we test these hypotheses. This theoretical model helps policymakers in digital market laws, behavioral economics, marketing, and digital platform design consider how choice architectures may improve consumer choices and encourage rational decision-making.

econ.TH

Strategic Bid Shading in Real-Time Bidding Auctions in Ad Exchange Using Minority Game Theory

Traditional auction theory posits that bid value exhibits a positive correlation with the probability of securing the auctioned object in ascending auctions. However, under uncertainty and incomplete information, as is characteristic in real-time advertising markets, truthful bidding may not always represent a dominant strategy or yield a Pure Strategy Nash Equilibrium. Real-Time Bidding (RTB) platforms operationalize impression-level auctions via programmatic interfaces, where advertisers compete in first-price auction settings and often resort to bid shading, i.e., strategically submitting bids below their private valuations to optimize payoff. This paper empirically investigates bid shading behaviors and strategic adaptation using large-scale RTB auction data from the Yahoo Webscope dataset. Integrating Minority Game Theory with clustering algorithms and variance-scaling diagnostics, we analyze equilibrium bidding behavior across temporally segmented impression markets. Our results reveal the emergence of minority-based bidding strategies, wherein agents partition hourly ad slots into submarkets and place bids strategically where they anticipate being in the numerical minority. This strategic heterogeneity facilitates reduced expenditure while enhancing win probability, functioning as an endogenous bid shading mechanism. The analysis highlights the computational and economic implications of minority strategies in shaping bidder dynamics and pricing outcomes in decentralized, high-frequency auction environments.

econ.TH

JU-NLP at Touch\'e: Covert Advertisement in Conversational AI-Generation and Detection Strategies

This paper proposes a comprehensive framework for the generation of covert advertisements within Conversational AI systems, along with robust techniques for their detection. It explores how subtle promotional content can be crafted within AI-generated responses and introduces methods to identify and mitigate such covert advertising strategies. For generation (Sub-Task~1), we propose a novel framework that leverages user context and query intent to produce contextually relevant advertisements. We employ advanced prompting strategies and curate paired training data to fine-tune a large language model (LLM) for enhanced stealthiness. For detection (Sub-Task~2), we explore two effective strategies: a fine-tuned CrossEncoder (\texttt{all-mpnet-base-v2}) for direct classification, and a prompt-based reformulation using a fine-tuned \texttt{DeBERTa-v3-base} model. Both approaches rely solely on the response text, ensuring practicality for real-world deployment. Experimental results show high effectiveness in both tasks, achieving a precision of 1.0 and recall of 0.71 for ad generation, and F1-scores ranging from 0.99 to 1.00 for ad detection. These results underscore the potential of our methods to balance persuasive communication with transparency in conversational AI.

cs.CL

Multi-modal encoder-decoder neural network for forecasting solar wind speed at L1

The solar wind, accelerated within the solar corona, sculpts the heliosphere and continuously interacts with planetary atmospheres. On Earth, high-speed solar-wind streams may lead to severe disruption of satellite operations and power grids. Accurate and reliable forecasting of the ambient solar-wind speed is therefore highly desirable. This work presents an encoder-decoder neural-network framework for simultaneously forecasting the daily averaged solar-wind speed for the subsequent four days. The encoder-decoder framework is trained with the two different modes of solar observations. The history of solar-wind observations from prior solar-rotations and EUV coronal observations up to four days prior to the current time form the input to two different encoders. The decoder is designed to output the daily averaged solar-wind speed from four days prior to the current time to four days into the future. Our model outputs the solar-wind speed with Root-Mean-Square Errors (RMSEs) of 55 km/s, 58 km/s, 58 km/s, and 58 km/s and Pearson correlations of 0.78, 0.66, 0.64 and 0.63 for one to four days in advance respectively. While the model is trained and validated on observations between 2010 - 2018, we demonstrate its robustness via application on unseen test data between 2019 - 2023, yielding RMSEs of 53 km/s and Pearson correlations 0.55 for a four-day advance prediction. Our encoder-decoder model thus produces much improved RMSE values compared to the previous works and paves the way for developing comprehensive multimodal deep learning models for operational solar wind forecasting.

astro-ph.SR

Flavor puzzle in three Higgs-doublet models: Insights from BGL and lessons from flavor data

We study a variant of the 3HDM, referred to as the BGL-3HDM, incorporating a $U(1)_1\times U(1)_2$ symmetry, which can distinguish the primary sources of mass for different fermion generations. In the version considered here, the Yukawa matrices in the down-quark and charged lepton sectors are diagonal, thereby eliminating tree-level FCNCs in these sectors. FCNC interactions mediated by neutral nonstandard Higgses are confined to the up-quark sector only. No new BSM parameters are introduced by the Yukawa sector of the model, making it as economical as the NFC versions of 3HDM with a $U(1)_1\times U(1)_2$ symmetry in terms of the number of free parameters. However, even in the down-quark and in the charged lepton sectors, flavor diagonal but nonuniversal Higgs couplings set this model apart from the NFC versions of the 3HDM.

hep-ph

Drell-Yan constraints on charged scalars: a weak isospin perspective

Charged scalars appear in many motivated extensions beyond the Standard Model. We analyze the constraints on charged scalar pair production via the Drell-Yan process at the Large Hadron Collider and interpret them in terms of weak isospin quantum numbers. Leveraging the experimental limits from existing LHC data and phenomenological recast analyses, we place bounds on the branching ratio of the charged scalar, as a function of its mass, electric charge, and isospin. This approach enables to determine limits on the branching ratios directly from experimental data, without appealing to a specific model. We provide a detailed analysis for singly and doubly charged scalars across various weak isospin scenarios, focusing on decays into leptonic and bosonic final states, and validate this approach in extended Higgs sectors such as the Higgs triplet model and Georgi-Machacek model.

hep-ph

GS-Net: Global Self-Attention Guided CNN for Multi-Stage Glaucoma Classification

Glaucoma is a common eye disease that leads to irreversible blindness unless timely detected. Hence, glaucoma detection at an early stage is of utmost importance for a better treatment plan and ultimately saving the vision. The recent literature has shown the prominence of CNN-based methods to detect glaucoma from retinal fundus images. However, such methods mainly focus on solving binary classification tasks and have not been thoroughly explored for the detection of different glaucoma stages, which is relatively challenging due to minute lesion size variations and high inter-class similarities. This paper proposes a global self-attention based network called GS-Net for efficient multi-stage glaucoma classification. We introduce a global self-attention module (GSAM) consisting of two parallel attention modules, a channel attention module (CAM) and a spatial attention module (SAM), to learn global feature dependencies across channel and spatial dimensions. The GSAM encourages extracting more discriminative and class-specific features from the fundus images. The experimental results on a publicly available dataset demonstrate that our GS-Net outperforms state-of-the-art methods. Also, the GSAM achieves competitive performance against popular attention modules.

cs.CV

Self-similarity of temporal interaction networks arises from hyperbolic geometry with time-varying curvature

The self-similarity of complex systems has been studied intensely across different domains due to its potential applications in system modeling, complexity analysis, etc., as well as for deep theoretical interest. Existing studies rely on scale transformations conceptualized over either a definite geometric structure of the system (very often realized as length-scale transformations) or purely temporal scale transformations. However, many physical and social systems are observed as temporal interactions among agents without any definitive geometry. Yet, one can imagine the existence of an underlying notion of distance as the interactions are mostly localized. Analysing only the time-scale transformations over such systems would uncover only a limited aspect of the complexity. In this work, we propose a novel technique of scale transformation that dissects temporal interaction networks under spatio-temporal scales, namely, flow scales. Upon experimenting with multiple social and biological interaction networks, we find that many of them possess a finite fractal dimension under flow-scale transformation. Finally, we relate the emergence of flow-scale self-similarity to the latent geometry of such networks. We observe strong evidence that justifies the assumption of an underlying, variable-curvature hyperbolic geometry that induces self-similarity of temporal interaction networks. Our work bears implications for modeling temporal interaction networks at different scales and uncovering their latent geometric structures.

physics.soc-ph

Demand Analysis and Customized Product Offering Design on E-Commerce Platform

It can be observed that the purchasing decision of an individual consumer in an electronic marketplace is determined by a set of factors, such as personal characteristics of the consumer, product pricing, minimum price-quantity combination offered, decision-making space, and underlying motivation of the consumer. These factors are combined to form a consumer's choice problem domain, which plays a pivotal role in the product offering. In this study, we attempt to focus on how the products? Offered can be customized by incorporating the quantity and pack size of the products along with the factors above to form a more extensive domain for examining the combined effects of all of these factors on demand. Accordingly, the demand function is defined by a novel method invoking the extended domain of choice problem in the electronic marketplace. Consequently, the predictable uncertainty associated with the consumer's demand function may disappear, increase the likelihood of earning optimum revenue through customized combinations of the components of the extended domain of choice problem, and improve the understanding of the fluctuations in consumer demand. Finally, we propose a generalized price response function with standard properties applicable to E-Commerce.

cs.GT

New physics interpretations for nonstandard values of $h\to Zγ$

Current measurement of the $h\to Zγ$ signal strength invite us to speculate about possible new physics interactions that exclusively affect $μ_{Zγ}$ without altering the other signal strengths. Additional consideration of tree-unitarity enables us to correlate the nonstandard values of $μ_{Zγ}$ with an upper limit on the scale of new physics. We find that even when $μ_{Zγ}$ deviates from the SM value by only $20\%$, the scale of new physics should be well within the reach of the LHC.

hep-ph

Sign of the $hZZ$ coupling and implication for new physics

The magnitudes of the couplings of the scalar resonance at 125 GeV with the SM particles are found to be consistent with those of the SM Higgs boson. However, the signs are not experimentally determined in most of the cases, a prime example being that with the $Z$-boson pair. In other words, $κ_Z^h$, the ratio of the couplings of the actual 125 GeV resonance with $ZZ$ and that of the SM Higgs boson with the same, is consistent with both $+1$ and $-1$, the latter being the `wrong-sign'. We argue that the wrong-sign $hZZ$ coupling will necessitate the intervention of new physics below $\mathcal{O}\left(620\right)$ GeV to safeguard the underlying theory from unitarity violation. The strength of the new nonstandard couplings can be derived from the unitarity sum rules, which are comparable to the SM-Higgs couplings in magnitude. Thus the strong limits from the direct searches at the LHC can help us rule out the existence of such nonstandard particles with unusually large couplings thereby disfavoring the possibility of a wrong-sign $hZZ$ coupling.

hep-ph

New physics implications of VBF searches exemplified through the Georgi-Machacek model

LHC searches for nonstandard scalars in vector boson fusion (VBF) production processes can be particularly efficient in probing scalars belonging to triplet or higher multiplet representations of the Standard Model $SU(2)_L$ gauge group. They can be especially relevant for models where the additional scalars do not have any tree-level couplings to the Standard Model fermions, rendering VBF as their primary production mode at the LHC. In this work, we employ the latest LHC data from VBF resonance searches to constrain the properties of nonstandard scalars, taking the Georgi-Machacek model as a prototypical example. We take into account the theoretical constraints on the potential from unitarity and boundedness-from-below as well as indirect constraints coming from the signal strength measurements of the 125 GeV Higgs boson at the LHC. To facilitate the phenomenological analysis we advocate a convenient reparametrization of the trilinear couplings in the scalar potential. We derive simple correlations among the model parameters corresponding to the decoupling limit of the model. We explicitly demonstrate how a combination of theoretical and phenomenological constraints can push the GM model towards the decoupling limit. Our analysis suggests that the VBF searches can provide key insights into the composition of the electroweak vacuum expectation value.

hep-ph

A Model of Competitive Assortment Planning Algorithm

With a novel search algorithm or assortment planning or assortment optimization algorithm that takes into account a Bayesian approach to information updating and two-stage assortment optimization techniques, the current research provides a novel concept of competitiveness in the digital marketplace. Via the search algorithm, there is competition between the platform, vendors, and private brands of the platform. The current paper suggests a model and discusses how competition and collusion arise in the digital marketplace through assortment planning or assortment optimization algorithm. Furthermore, it suggests a model of an assortment algorithm free from collusion between the platform and the large vendors. The paper's major conclusions are that collusive assortment may raise a product's purchase likelihood but fail to maximize expected revenue. The proposed assortment planning, on the other hand, maintains competitiveness while maximizing expected revenue.

econ.TH

Fingerprinting the Type-Z three Higgs doublet models

There has been great interest in a model with three Higgs doublets in which fermions with a particular charge couple to a single and distinct Higgs field. We study the phenomenological differences between the two common incarnations of this so-called Type-Z 3HDM. We point out that the differences between the two models arise from the scalar potential only. Thus we focus on observables that involve the scalar self-couplings. We find it difficult to uncover features that can uniquely set apart the $Z_3$ variant of the model. However, by studying the dependence of the trilinear Higgs couplings on the nonstandard masses, we have been able to isolate some of the exclusive indicators for the $Z_2\times Z_2$ version of the Type-Z 3HDM. This highlights the importance of precision measurements of the trilinear Higgs couplings.

hep-ph

Complementarity in Demand-side Variables and Educational Participation

Decision to participate in education depends on the circumstances individual inherits and on the returns to education she expects as well. If one person from any socio-economically disadvantaged social group inherits poor circumstances measured in terms of family background, then she is having poor opportunities vis-à-vis her capability set becomes confined. Accordingly, her freedom to choose the best alternative from many is also less, and she fails to expect the potential returns from educational participation. Consequently, a complementary relationship between the circumstances one inherits and the returns to education she expects can be observed. This paper is an attempt to look at this complementarity on the basis of theoretical logic and empirical investigation, which enables us to unearth the origin of inter-group disparity in educational participation, as is existed across the groups defined by taking caste and gender together in Indian society. Furthermore, in the second piece of analysis, we assess the discrimination in the likelihood of educational participation by invoking the method of decomposition of disparity in the likelihood of educational participation applicable in the logistic regression models, which enables us to re-establish the earlier mentioned complementary relationship.

econ.TH