SearcharxivSearch

arXiv subjects

Puja Das

Publications and source records attributed to Puja Das.

6 recordsLinked to original sources

Fortress: A Case Study in Stabilizing Search Recommendations via Temporal Data Augmentation and Feature Pruning

In search and recommendation systems, predictive models often suffer from temporal instability when certain input features introduce volatility in output scores. This instability can degrade model reliability and user experience especially in multi-stage systems where consistent predictions are critical for downstream decision making. We introduce Fortress, a general framework for enhancing model stability and accuracy by identifying and pruning features that contribute to inconsistent prediction scores over time. Fortress leverages historical snapshots temporally partitioned datasets capturing score fluctuations for the same entity across periods and follows a four-step process: (1) collect historical snapshots, (2) identify samples with unstable predictions, (3) isolate and remove instability-inducing features, and (4) retrain models using only stable features. While semantic features from LLMs and BERT-based models improve generalization, they often lack full query or entity coverage. Engagement-based features offer strong predictive power but tend to introduce temporal instability. Fortress mitigates this trade-off by suppressing the volatility of engagement signals while retaining their predictive value leading to more stable and accurate models. We validate Fortress on a query-to-app relevance model in a large-scale app marketplace. Offline experiments demonstrate notable improvements in prediction stability (measured by Coefficient of Variation) and classification performance (measured by PR-AUC).

cs.IR

Nambu-Goldstone boson phenomenology in Domain-Wall Standard Model

We investigate the Domain-Wall Standard Model (DWSM), a five-dimensional framework in which all Standard Model (SM) particles are localized on a domain wall embedded in a non-compact extra spatial dimension. A distinctive feature of this setup is the emergence of a Nambu-Goldstone (NG) boson, arising from the spontaneous breaking of translational invariance in the extra dimension due to the localization of SM chiral fermions. This NG boson couples via Yukawa interactions to SM fermions and their Kaluza-Klein (KK) excitations. We study the phenomenology of this NG boson and derive constraints from astrophysical processes (supernova cooling), Big Bang Nucleosynthesis (BBN), and collider searches for KK-mode fermions at the Large Hadron Collider (LHC). The strongest limits arise from LHC data: we reinterpret existing mass bounds on squarks and sleptons in simplified supersymmetric models (assuming a massless lightest neutralino), as well as limits on exotic hadrons containing long-lived squarks or long-lived charged sleptons in the regime of extremely small Yukawa couplings. From this analysis, we obtain a conservative lower bound of 1 TeV on the masses of KK-mode quarks and charged leptons. Finally, we discuss the prospects for producing KK-mode fermions at future high-energy lepton colliders and outline strategies to distinguish their signatures from those of sfermions.

hep-ph

Finer resolutions and targeted process representations in earth systems models improve hydrologic projections and hydroclimate impacts

Earth system models inform water policy and interventions, but knowledge gaps in hydrologic representations limit the credibility of projections and impacts assessments. The literature does not provide conclusive evidence that incorporating higher resolutions, comprehensive process models, and latest parameterization schemes, will result in improvements. We compare hydroclimate representations and runoff projections across two generations of Coupled Modeling Intercomparison Project (CMIP) models, specifically, CMIP5 and CMIP6, with gridded runoff from Global Runoff Reconstruction (GRUN) and ECMWF Reanalysis V5 (ERA5) as benchmarks. Our results show that systematic embedding of the best available process models and parameterizations, together with finer resolutions, improve runoff projections with uncertainty characterizations in 30 of the largest rivers worldwide in a mechanistically explainable manner. The more skillful CMIP6 models suggest that, following the mid-range SSP370 emissions scenario, 40% of the rivers will exhibit decreased runoff by 2100, impacting 260 million people.

physics.geo-ph

Hybrid physics-AI outperforms numerical weather prediction for extreme precipitation nowcasting

Precipitation nowcasting, critical for flood emergency and river management, has remained challenging for decades, although recent developments in deep generative modeling (DGM) suggest the possibility of improvements. River management centers, such as the Tennessee Valley Authority, have been using Numerical Weather Prediction (NWP) models for nowcasting but have struggled with missed detections even from best-in-class NWP models. While decades of prior research achieved limited improvements beyond advection and localized evolution, recent attempts have shown progress from physics-free machine learning (ML) methods and even greater improvements from physics-embedded ML approaches. Developers of DGM for nowcasting have compared their approaches with optical flow (a variant of advection) and meteorologists' judgment but not with NWP models. Further, they have not conducted independent co-evaluations with water resources and river managers. Here, we show that the state-of-the-art physics-embedded deep generative model, specifically NowcastNet, outperforms the High-Resolution Rapid Refresh (HRRR) model, the latest generation of NWP, along with advection and persistence, especially for heavy precipitation events. For grid-cell extremes over 16 mm/h, NowcastNet demonstrated a median critical success index (CSI) of 0.30, compared with a median CSI of 0.04 for HRRR. However, despite hydrologically relevant improvements in point-by-point forecasts from NowcastNet, caveats include the overestimation of spatially aggregated precipitation over longer lead times. Our co-evaluation with ML developers, hydrologists, and river managers suggests the possibility of improved flood emergency response and hydropower management.

physics.ao-ph

Testing neutrino mass hierarchy under type-II seesaw scenario in $U(1)_X$ from colliders

The origin of tiny neutrino mass is a long standing unsolved puzzle of the Standard Model (SM), which allows us to consider scenarios beyond the Standard Model (BSM) in a variety of ways. One of them being a gauge extension of the SM may be realized as in the form of an anomaly free, general $U(1)_X$ extension of the SM, where an $SU(2)_L$ triplet scalar with a $U(1)_X$ charge is introduced to have Dirac Yukawa couplings with the SM lepton doublets. Once the triplet scalar developes a Vacuum Expectation Value (VEV), light neutrinos acquire their tiny Majorana masses. Hence, the decay modes of the triplet scalar has a direct connection to the neutrino oscillation data for different neutrino mass hierarchies. After the breaking of the $U(1)_X$ gauge symmetry, a neutral $U(1)_X$ gauge boson $(Z^\prime)$ acquires mass, which interacts differently with the left and right handed SM fermions. Satisfying the recent LHC bounds on the triplet scalar and $Z^\prime$ boson productions, we study the pair production of the triplet scalar at LHC, 100 TeV proton proton collider FCC, $e^-e^+$ and $\mu^-\mu^+$ colliders followed by its decay into dominant dilepton modes whose flavor structure depend on the neutrino mass hierarchy. Generating the SM backgrounds, we study the possible signal significance of four lepton final states from the triplet scalar pair production. We also compare our results with the purely SM gauge mediated triplet scalar pair production followed by four lepton final states, which could be significant only in $\mu^- \mu^+$ collider.

hep-ph

Multi-task Sparse Structure Learning

Multi-task learning (MTL) aims to improve generalization performance by learning multiple related tasks simultaneously. While sometimes the underlying task relationship structure is known, often the structure needs to be estimated from data at hand. In this paper, we present a novel family of models for MTL, applicable to regression and classification problems, capable of learning the structure of task relationships. In particular, we consider a joint estimation problem of the task relationship structure and the individual task parameters, which is solved using alternating minimization. The task relationship structure learning component builds on recent advances in structure learning of Gaussian graphical models based on sparse estimators of the precision (inverse covariance) matrix. We illustrate the effectiveness of the proposed model on a variety of synthetic and benchmark datasets for regression and classification. We also consider the problem of combining climate model outputs for better projections of future climate, with focus on temperature in South America, and show that the proposed model outperforms several existing methods for the problem.

cs.LG