SearcharxivSearch

arXiv subjects

Sayantan Choudhury

Publications and source records attributed to Sayantan Choudhury.

At least 19 recordsLinked to original sources

Multi-source conformal prediction: leveraging heterogeneity via localization

Many modern prediction tasks involve data from multiple heterogeneous sources, while the test distribution may differ substantially from any individual source. Although heterogeneity poses challenges, it also offers an opportunity: different sources may provide complementary information, with some regions of the feature space better represented in one source than another. We propose Multi-Source Randomly Localized Conformal Prediction (MS-RLCP), which builds on the local coverage properties of randomly localized conformal prediction (RLCP) (Hore and Barber, 2025) and extends it to multiple sources through data-adaptive source selection. Under the widely adopted assumption of a shared response distribution conditional on the features across sources and the test population, we establish finite-sample coverage bounds using an interpretable notion of envelope distribution that captures their aggregate feature-space representation. Our analysis allows the test feature distribution to be absolutely continuous with respect to the envelope, extending beyond mixtures of source distributions. Under additional regularity conditions, we also establish asymptotic test-conditional coverage. Simulations and real-world experiments demonstrate the effectiveness of MS-RLCP across varying levels of data heterogeneity.

stat.ML

What new physics can we extract from inflation using the ACT DR6 and DESI DR2 Observations?

We present a comprehensive analysis of inflationary models in light of projected sensitivities from forthcoming CMB and gravitational wave experiments, incorporating data from recent ACT DR6, DESI DR2, CMB-S4, LiteBIRD, and SPHEREx. Focusing on precise predictions in the $(n_s, α_s, β_s)$ parameter space, we evaluate a broad class of inflationary scenarios -- including canonical single-field models, non-minimally coupled theories, and string-inspired constructions such as Starobinsky, Higgs, Hilltop, $α$-attractors, and D-brane models. Our results show that next-generation observations will sharply constrain the scale dependence of the scalar power spectrum, elevating $α_s$ and $β_s$ as key discriminants between large-field and small-field dynamics. Strikingly, several widely studied models -- such as quartic Hilltop inflation and specific DBI variants -- are forecast to be excluded at high significance. We further demonstrate that the combined measurement of $β_s$ and the field excursion $Δϕ$ offers a novel diagnostic of kinetic structure and UV sensitivity. These findings underscore the power of upcoming precision cosmology to probe the microphysical origin of inflation and decisively test broad classes of theoretical models.

astro-ph.CO

EFT Perspective On de-Sitter S-Matrix

Non-perturbative limitations on low-energy effective field theories (EFTs) based on the characteristics of high-energy theory are provided by the analyticity of the flat-space version of the S-matrix. Although the analyticity of the flat-space S-matrix is widely established, it is difficult to apply this framework to de Sitter space because the growing backdrop breaks time-translation symmetry and makes it more difficult to define asymptotic states. The flat-space analyticity imprint on the de Sitter S-matrix is examined in this study. On a certain limit, we derive a comprehensive relationship between the flat-space amplitude and the de Sitter S-matrix. In particular, we demonstrate that the relationship is valid for tree-level amplitude exchanging with arbitrary local derivative interactions with a large scalar field. Next, we contend that this specific limit is more consistent with the definition of EFT since, similar to flat space, the Mandelstam variable may be identified as the unique energy scale because the total energy dependence of the de Sitter S-matrix becomes negligible. Finally, we also find an unexpected connection between the idea of generalized energy conservation of an S-matrix of four-dimensional de Sitter and exceptional EFTs in de Sitter space. We restrict the coupling constants in theories of self-interacting scalars dwelling in the exceptional series of de Sitter representations by requiring that such an S-matrix only has support when the total energies of in and out states are equal. We rediscover the Dirac-Born-Infeld (DBI) and Special Galileon theories, in which a single coupling constant uniquely fixes the four-point scalar self-interactions.

hep-th

One-shot Conditional Sampling: MMD meets Nearest Neighbors

How can we generate samples from a conditional distribution that we never fully observe? This question arises across a broad range of applications in both modern machine learning and classical statistics, including image post-processing in computer vision, approximate posterior sampling in simulation-based inference, and conditional distribution modeling in complex data settings. In such settings, compared with unconditional sampling, additional feature information can be leveraged to enable more adaptive and efficient sampling. Building on this, we introduce Conditional Generator using MMD (CGMMD), a novel framework for conditional sampling. Unlike many contemporary approaches, our method frames the training objective as a simple, adversary-free direct minimization problem. A key feature of CGMMD is its ability to produce conditional samples in a single forward pass of the generator, enabling practical one-shot sampling with low test-time complexity. We establish rigorous theoretical bounds on the loss incurred when sampling from the CGMMD sampler, and prove convergence of the estimated distribution to the true conditional distribution. In the process, we also develop a uniform concentration result for nearest-neighbor based functionals, which may be of independent interest. Finally, we show that CGMMD performs competitively on synthetic tasks involving complex conditional densities, as well as on practical applications such as image denoising and image super-resolution.

stat.ML

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters

Gradient clipping is a standard safeguard for training neural networks under noisy, heavy-tailed stochastic gradients; yet, most clipping rules treat all parameters as vectors and ignore the matrix structure of modern architectures. We show empirically that data outliers often amplify only a small number of leading singular values in layer-wise gradient matrices, while the rest of the spectrum remains largely unchanged. Motivated by this phenomenon, we propose spectral clipping, which stabilizes training by clamping singular values that exceed a threshold while preserving the singular directions. This framework generalizes classical gradient norm clipping and can be easily integrated into existing optimizers. We provide a convergence analysis for non-convex optimization with spectrally clipped SGD, yielding the optimal $\mathcal{O}\left(K^{\frac{2 - 2α}{3α- 2}}\right)$ rate for heavy-tailed noise. To minimize hyperparameter tuning, we introduce layer-wise adaptive thresholds based on moving averages or sliding-window quantiles of the top singular values. Finally, we develop efficient implementations that clip only the top $r$ singular values via randomized truncated SVD, avoiding full decompositions for large layers. We demonstrate competitive performance across synthetic heavy-tailed settings and neural network training tasks.

cs.LG

Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition

Most first-order optimizers treat matrix-valued parameters as vectors, ignoring the intrinsic geometry of hidden-layer weights in neural networks. Muon addresses this mismatch by updating along the polar factor of a momentum matrix, but its theoretical understanding has lagged behind practice. In particular, practical implementations incorporate Nesterov momentum, compute the polar factor only approximately, and operate with stochastic gradients that may be heavy-tailed. We close this gap by developing a convergence theory for Muon with Nesterov momentum and inexact polar decomposition in non-convex matrix optimization under heavy-tailed noise. Our analysis builds on a unified framework for inexact polar decomposition that captures practical iterative approximations such as Newton-Schulz and quantifies how their errors propagate through the optimization dynamics. Under this framework, we establish an optimal iteration and sample complexity of $O \left(\varepsilon^{\frac{-(3α-2)}{(α-1)}} \right)$ for finding an $\varepsilon$-stationary point, where $α\in(1,2]$ denotes the heavy-tail index. For the inexact-polar setting with $σ_1=0$, we also provide guarantees that do not require prior knowledge of $α$. We analyze a randomized low-rank polar decomposition that is substantially more efficient than full-space methods while remaining compatible with our theory. Numerical experiments further demonstrate the effectiveness of the proposed inexact and randomized variants.

math.OC

Doubly-Unlinked Regression for Dependent Data

Shuffled regression concerns settings in which covariates and responses are observed without their correct pairing. In dependent-data problems, a second form of missing correspondence can arise when responses are also detached from the latent temporal, spatial, or geometric domain that induces their dependence structure. We study regression under this joint loss of correspondence and, to our knowledge, provide the first systematic treatment of this setting. Specifically, we consider a doubly-unlinked regression model in which both the covariate-response link and the response-domain link are unknown, represented by two latent permutation matrices, while dependence is induced by an unobserved stochastic process. This framework unifies shuffled regression and latent-domain permutation models within a common dependent-data setting. We characterize signal-to-noise regimes governing recovery of the regression parameter and the latent permutations, and show that consistent estimation of the regression coefficient can be achieved under strictly weaker conditions than exact permutation recovery. To address the combinatorial difficulty of inference, we develop REPAIR, a variational Bayes method based on a block-structured permutation model that captures localized scrambling while substantially reducing computational complexity. Simulations and an applied example illustrate the empirical behavior of REPAIR and support the theoretical results.

math.ST

CFT Perspective On de-Sitter Cosmological Correlators

We investigate the principles of quantum field theory using a stiff de Sitter space. We demonstrate that a non-unitary Lagrangian on a Euclidean AdS geometry can produce the perturbative expansion of late-time correlation functions to all orders. This discovery greatly simplifies perturbative computations while also allowing us to prove fundamental features of these correlators, which are part of a Euclidean CFT. This allows us to construct an OPE expansion, limit the operator spectrum, and deduce the analytic structure of the spectral density that captures the conformal partial wave expansion of a late-time four-point function. In general, the standard CFT concept of unitarity does not apply to dimensions and OPE coefficients. Rather, the positivity of the spectral density represents the unitarity of the de Sitter theory. This assertion is non-perturbative and does not depend on the use of Euclidean AdS Lagrangians. In a scalar theory, we compute tree-level and entire one-loop-resummed exchange diagrams to demonstrate and verify these characteristics. In the spectrum density, an exchanged particle shows up as a resonant characteristic that may be helpful in experimental searches.

hep-th

Untangling PBH overproduction in $w$-SIGWs generated by Pulsar Timing Arrays for MST-EFT of single field inflation

Our work highlights the crucial role played by the equation of state (EoS) parameter $w$ within the context of single field inflation with Multiple Sharp Transitions (MSTs) to untangle the current state of the PBH overproduction issue. We examine the situation for a broad interval of EoS parameter that remains most favourable to explain the recent data released by the pulsar timing array (PTA) collaboration. Our analysis yields the interval, $0.2 \leq w \leq 1/3$, to be the most acceptable window from the SIGW interpretation of the PTA signal and where sizeable PBHs abundance, $f_{\rm PBH} \in (10^{-3},1)$, is observed. We also obtain $w=1/3$, radiation-dominated era, to be the best scenario to explain the early stages of the Universe and address the overproduction problem. Within the range of $1 \leq c_{s} \leq 1.17$, we construct a regularized-renormalized-resummed scalar power spectrum whose amplitude obeys the perturbativity criterion while being substantial enough to generate EoS dependent scalar induced gravitational waves ($w$-SIGWs) consistent with NANOGrav-15 data. Working for both $c_{s} = 1\;{\rm and}\;1.17$, we find the $c_{s}=1.17$ case more favourable for generating large mass PBHs, $M_{\rm PBH}\sim {\cal O}(10^{-6}-10^{-3})M_{\odot}$, as potential dark matter candidates with substantial abundance after constraints coming from microlensing experiments.

astro-ph.CO

Quantum Discord in de-Sitter Axiverse

In this work, we compute quantum discord between two causally independent areas in $3+1$ dimensions global de Sitter Axiverse to investigate the signs of quantum entanglement. For this goal, we study a bipartite quantum field theoretic setting driven by an Axiverse that arises from the compactification of Type IIB strings on a Calabi-Yau three fold. We consider a spherical surface that separates the interior and exterior causally unconnected subregions of the spatial slice of the global de Sitter space. The Bunch-Davies state is the most straightforward initial quantum vacuum that may be used for computing purposes. Two observers are introduced, one in an open chart of de Sitter space and the other in a global chart. The observers calculate the quantum discord generated by each detecting a mode. The relationship between an observer in one of the two Rindler charts in flat space and another in a Minkowski chart is comparable to this circumstance. We see that when the curvature of the open chart increases, the state becomes less entangled. Nevertheless, we see that even in the limit when entanglement vanishes, the quantum discord never goes away.

hep-th

Mass gap in non-perturbative quadratic $\mathcal{R}^2$ gravity via Dyson-Schwinger

We apply in a simple model derived from quadratic $\mathcal{R}^2$ gravity the technique of Dyson-Schwinger equations to solve for its corresponding quantum theory. Particularly, we solve the classical equations of motion to get a solution to the hierarchy of Dyson-Schwinger equations in the limit of large Ricci scalar, assumed to be constant and larger than the square of the Starobinsky mass. Moving to the Einstein frame, the model admits Higgs-like solutions with a single particle having a finite mass. We quantize the scalar field showing the appearing of a mass gap through a Higgs-like solution. The presence of the mass gap, that increases with the square root of the Ricci scalar, shows how the effect of the scalar sector at low-energy becomes ineffective, making it relevant only at short distances.

hep-th

Notes On de-Sitter Mellin Barnes Amplitudes

In this paper, we create a Mellin space method for boundary correlation functions in de Sitter (dS) and anti-de Sitter (AdS) spaces. We demonstrate that the analytic continuation between AdS${}_{d+1}$ and dS${}_{d+1}$ is encoded in a set of simple relative phases using the Mellin-Barnes representation of correlators. It helps us to determine the scalar three-point and four-point functions and their corresponding Mellin-Barnes amplitudes in dS${}_{d+1}$ space using the known results from AdS${}_{d+1}$ space. The Mellin-Barnes representation reveals the analytic structure of boundary correlation functions over all $d$ and scaling dimensions. In the present discussion, the {\it split representation} have been used as an instrumental technique in particularly the evaluation of bulk Witten diagrams and is suitable to obtain the {\it Conformal Partial Wave decomposition} of tree-level exchange in the bulk Witten diagrams. The equivalent adjustment to the cosmological three-point and four-point function of generic external scalars may be further extracted from these results, assuming the weak breakdown of the de Sitter isometries. These findings offer a step towards a more methodical comprehension of de Sitter observables utilising Mellin space techniques at the tree level and beyond.

hep-th

Quintessential Inflation in Light of ACT DR6

We perform a precision investigation of smooth quintessential inflation in which a single canonical scalar field unifies the two known phases of cosmic acceleration. Using a CMB-normalized runaway exponential potential, we obtain sharply predictive inflationary observables: a red-tilted spectrum with $n_s = 0.964241$ and an exceptionally suppressed tensor-to-scalar ratio $r = 7.48 \times 10^{-5}$ at $N=60$, lying near the optimal region of current Planck+ACT constraints. Remarkably, all observable scales exit the horizon within an extremely narrow field interval $Δϕ\simeq 0.03\,M_{\rm Pl}$, tightly linking early and late-time dynamics and reducing theoretical ambiguities. While inflationary tensors remain invisible to CMB B-mode surveys, the subsequent stiff epoch-an intrinsic hallmark of quintessential cosmology-imprints a blue-tilted stochastic gravitational-wave background within the discovery reach of future interferometers such as LISA, DECIGO, ALIA, and BBO. Our results demonstrate that this minimal, featureless model not only survives current bounds, but provides concrete, falsifiable predictions across gravitational-wave frequencies spanning over twenty orders of magnitude.

gr-qc

Multiplayer Federated Learning: Reaching Equilibrium with Less Communication

Traditional Federated Learning (FL) approaches assume collaborative clients with aligned objectives working towards a shared global model. However, in many real-world scenarios, clients act as rational players with individual objectives and strategic behaviors, a concept that existing FL frameworks are not equipped to adequately address. To bridge this gap, we introduce Multiplayer Federated Learning (MpFL), a novel framework that models the clients in the FL environment as players in a game-theoretic context, aiming to reach an equilibrium. In this scenario, each player tries to optimize their own utility function, which may not align with the collective goal. Within MpFL, we propose Per-Player Local Stochastic Gradient Descent (PEARL-SGD), an algorithm in which each player/client performs local updates independently and periodically communicates with other players. We theoretically analyze PEARL-SGD and prove that it reaches a neighborhood of equilibrium with less communication in the stochastic setup compared to its non-local counterpart. Finally, we verify our theoretical findings through numerical experiments.

cs.LG

Extragradient Method for $(L_0, L_1)$-Lipschitz Root-finding Problems

Introduced by Korpelevich in 1976, the extragradient method (EG) has become a cornerstone technique for solving min-max optimization, root-finding problems, and variational inequalities (VIs). Despite its longstanding presence and significant attention within the optimization community, most works focusing on understanding its convergence guarantees assume the strong L-Lipschitz condition. In this work, building on the proposed assumptions by Zhang et al. [2024b] for minimization and Vankov et al.[2024] for VIs, we focus on the more relaxed $α$-symmetric $(L_0, L_1)$-Lipschitz condition. This condition generalizes the standard Lipschitz assumption by allowing the Lipschitz constant to scale with the operator norm, providing a more refined characterization of problem structures in modern machine learning. Under the $α$-symmetric $(L_0, L_1)$-Lipschitz condition, we propose a novel step size strategy for EG to solve root-finding problems and establish sublinear convergence rates for monotone operators and linear convergence rates for strongly monotone operators. Additionally, we prove local convergence guarantees for weak Minty operators. We supplement our analysis with experiments validating our theory and demonstrating the effectiveness and robustness of the proposed step sizes for EG.

math.OC

Gravitational Wave Signatures of Periodic Motion near Higher-Derivative Einstein-Æther Black Holes

Higher-derivative modifications of general relativity are generically expected from effective field theory approaches to quantum gravity, and they arise naturally in Lorentz-violating theories such as Einstein-Ether gravity. In this work, we investigate black hole spacetimes within Einstein-Ether theory supplemented by quadratic curvature corrections, including terms proportional to $R^2$, $R_{μν} R^{μν}$, and $R_{μνλρ} R^{μνλρ}$. We derive the corrected static, spherically symmetric metric perturbatively and examine its effects on the geodesic structure and gravitational wave emission. In particular, we analyze periodic timelike orbits in this background and compute the associated tensor-mode gravitational waveforms using the quadrupole approximation. Our results demonstrate that even small higher-derivative corrections can induce distinguishable shifts in the orbital dynamics and imprint characteristic phase modulations and harmonic deformations in the gravitational wave signal. These effects modify the frequency spectrum and amplitude envelope of $h_{+}$ and $h_{\times}$ in a manner sensitive to the coupling constants $α$, $β$, and $γ$, and the Ether parameter $c_{13}$. The resulting signatures provide a potential observational window into ultraviolet deviations from general relativity and Lorentz symmetry in the strong-field regime.

gr-qc

Stochastic origin of primordial fluctuations in the Sky

We provide a study of the effects of the Effective Field Theory (EFT) generalisation of stochastic inflation on the production of primordial black holes (PBHs) in a model-independent single-field context. We demonstrate how the scalar perturbations' Infra-Red (IR) contributions and the emerging Fokker-Planck equation driving the probability distribution characterise the Langevin equations for the ``soft" modes in the quasi-de Sitter background. Both the classical-drift and quantum-diffusion-dominated regimes undergo a specific analysis of the distribution function using the stochastic-$δN$ formalism, which helps us to evade a no-go theorem on the PBH mass. Using the EFT-induced alterations, we evaluate the local non-Gaussian parameters in the drift-dominated limit.

gr-qc