SearcharxivSearch

arXiv subjects

Harsh Mehta

Publications and source records attributed to Harsh Mehta.

At least 19 recordsLinked to original sources

LLMs struggle to simulate human belief updates in controlled environments

LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.

cs.CL

Detecting Axion-like particles using Cosmic Variance Cancellation with CMB and Radio surveys

Axions and axion-like particles (ALPs) arise naturally in many extensions of the Standard Model and are among the well-motivated candidates for dark matter. In the presence of magnetic fields of galaxy clusters, the Cosmic Microwave Background (CMB) photons can convert to ALPs, with the efficiency of the process governed by the cluster electron density and magnetic field profiles, the photon-ALP coupling strength (${g_{a\gamma}}$), as well as the frequency ($\nu$) of the photon at the redshift of the cluster. The CMB blackbody spectrum suggests this resonant conversion takes place at radio wavelengths as well, following the spectral behaviour of the ALP distortion signal. This opens up a new window to search for ALPs using cosmic variance cancellation (CVC), with multi-frequency tracers of the same phenomenon in CMB photon-ALP resonant conversion. The constraints on the ALP signal ratios from different combinations of microwave and radio bands of Simons Observatory (SO) and Square Kilometer Array (SKA), can be significantly improved using CVC as compared to the case of using auto-only spectra from the two experiments. With the large number of galaxy clusters that will be observed by SO and SKA, we will be able to obtain much more information using CVC, especially for the case of low-mass ALPs with stronger signals. Using the auto-only spectra from galaxy clusters up to redshift $z = 1$ for inference of normalized ratio parameter, we obtain a standard deviation of $5.9 \times 10^{-2}$ for ALP mass $m_a = 10^{-14} \, \rm{eV}$, which improves to $1.3 \times 10^{-2}$ using CVC. Not only is this method a universal probe of the ALP distortion signal using its spectral dependence, but will be able to provide a more robust consistency check, helping to identify and mitigate potential spurious signals that might arise in CMB-only analyses, based on its frequency behavior in different bands.

astro-ph.CO

Modified Cosmology or Modified Galaxy Astrophysics is Driving the z>6 JWST Results? CMB Experiments can discover the Origin in the Near Future

The massive and bright galaxies observed by the James Webb Space Telescope (JWST) at high redshifts ($z > 6$) have challenged our understanding of the Universe. This may require revisiting the physics of galaxy formation and evolution, or modifying the $\Lambda$CDM cosmological model to explain these observations, or both. We show that high-resolution CMB experiments such as the Simons Observatory (or CMB-S4) can measure smoking-gun signatures jointly in weak lensing and kinematic Sunyaev-Zeldovich (kSZ) power spectra, which can shed light on both these scenarios. An increase in the matter power spectrum at small scales will enhance the number density of dark matter halos at high redshifts, thereby increasing the galaxy formation rate. This will cause enhanced weak lensing signal from these redshifts and also lead to enhanced patchy-kSZ signal from the epoch of reionization. However, if only galaxy astrophysics is modified, without any modification in the matter power spectrum, then the patchy-kSZ signal gets altered, while the weak lensing signal remains nearly unaltered. We show that we can measure the modified astrophysical and cosmological scenarios at a statistical significance of $10.4\sigma$ (and $29.8\sigma$) from Simons Observatory (and CMB-S4), which will enable a conclusive understanding on what physical process is driving the high-redshift observations of JWST.

astro-ph.CO

Evolution of Entanglement Witness of Dicke State under Noise and Error Mitigation

The experimental verification of multipartite entangled states is essential for advancing quantum information processing. Entanglement witnesses (EWs) provide a widely used and experimentally accessible approach for detecting genuinely multipartite entangled states. In this work, we theoretically derive the entanglement witness for the four-qubit Dicke state and experimentally evaluate it on two distinct IBM 127-qubit Quantum Processing Units (QPUs), namely ibm\_sherbrook and ibm\_brisbane. A negative expectation value of the witness operator serves as a sufficient condition for confirming genuine multipartite entanglement. We report the maximum (negative) values of the witness achieved on these QPUs as $-0.178 \pm 0.009$ and $-0.169 \pm 0.002$, corresponding to two different state preparation protocols. Additionally, we theoretically investigate the effect of various noise channels on the genuine entanglement of a four-qubit Dicke state using the Qiskit Aer simulator. We show the behavior of the EW constructed under the assumption of Markovian and non-Markovian amplitude damping and depolarizing noises, bit-phase flip noise, and readout errors. We also investigate the effect of varying thermal relaxation time on the EW, depicting a bound on the $T_1$ time required for successful generation of a Dicke State on a superconducting QPU.

quant-ph

A White Paper on The Multi-Messenger Science Landscape in India

The multi-messenger science using different observational windows to the Universe such as Gravitational Waves (GWs), Electromagnetic Waves (EMs), Cosmic Rays (CRs), and Neutrinos offer an opportunity to study from the scale of a neutron star to cosmological scales over a large cosmic time. At the smallest scales, we can explore the structure of the neutron star and the different energetics involved in the transition of a pre-merger neutron star to a post-merger neutron star. This will open up a window to study the properties of matter in extreme conditions and a guaranteed discovery space. On the other hand, at the largest cosmological scales, multi-messenger observations allow us to study the long-standing problems in physical cosmology related to the Hubble constant, dark matter, and dark energy by mapping the expansion history of the Universe using GW sources. Moreover, the multi-messenger studies of astrophysical systems such as white dwarfs, neutron stars, and black holes of different masses, all the way up to a high redshift Universe, will bring insightful understanding into the physical processes associated with them that are inaccessible otherwise. This white paper discusses the key cases in the domain of multi-messenger astronomy and the role of observatories in India which can explore uncharted territories and open discovery spaces in different branches of physics ranging from nuclear physics to astrophysics.

astro-ph.HE

The Prospect from the Upcoming CMB Experiment LiteBIRD to Discover Axion-like Particles Using Milky Way

The existence of axion-like particles (ALPs) can be probed from their signatures in the Cosmic Microwave Background (CMB) due to the photon-ALP resonant conversion over the mass range of ALPs that matches with the effective mass of photons in the plasma in the astrophysical systems. Such a conversion can also occur in the Milky Way halo and disk and can cause a unique spatial and spectral distortion. The signal is highly non-Gaussian and cannot be measured precisely by the usual power-spectrum approach. We devise a new technique to search for this signal from the upcoming full-sky CMB experiment LiteBIRD using its multi-frequency band using a template-based spatial profile of the ALP distortion signal. This technique captures the large-scale non-Gaussian aspects of the ALP distortion signal in terms of a spatial template and makes it possible to search for any non-zero ALP signal. We show that the inference of the ALP coupling using the template-based technique from LiteBIRD can provide constraints on the coupling constant approximately $ g_{a\gamma} < 6.5 \times 10^{-12} \, \mathrm{GeV}^{-1}$ for ALP masses below $10^{-14}$ eV at 95\% confidence interval which is an order of magnitude better than the current bounds from CERN Axion Solar Telescope (CAST) at $g_{a\gamma} < 6.6 \times 10^{-11} \, \mathrm{GeV}^{-1}$, This shows the capability of future multi-band CMB experiment LiteBIRD in opening the discovery space towards physics beyond the standard model.

astro-ph.CO

Gemma 3 Technical Report

We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision understanding abilities, a wider coverage of languages and longer context - at least 128K tokens. We also change the architecture of the model to reduce the KV-cache memory that tends to explode with long context. This is achieved by increasing the ratio of local to global attention layers, and keeping the span on local attention short. The Gemma 3 models are trained with distillation and achieve superior performance to Gemma 2 for both pre-trained and instruction finetuned versions. In particular, our novel post-training recipe significantly improves the math, chat, instruction-following and multilingual abilities, making Gemma3-4B-IT competitive with Gemma2-27B-IT and Gemma3-27B-IT comparable to Gemini-1.5-Pro across benchmarks. We release all our models to the community.

cs.CL

Turbulence Induced Non-Gaussian Spectral Distortion in the Microwave Sky from Photon-Axion Conversion in Galaxy Clusters

The conversion of CMB photons to axions (or axion-like particles (ALPs)) can lead to a unique spectral distortion in the temperature and polarization sky which can be explored in upcoming CMB experiments. In this work we have developed a numerical simulation-based technique of photons to ALPs conversion in the galaxy clusters and show for the first time that this physical process can lead to large non-Gaussian signal in the temperature and polarization field, which is impacted by the presence of inhomogeneities and turbulence in the electron density and magnetic field. Our simulation-based technique can simulate the theoretical signal for different scenarios of cluster electron density and magnetic field turbulence and provides testable predictions to discover ALPs from galaxy clusters using spatially non-Gaussian and anisotropic spectral distortion of the microwave sky. We show that the presence of turbulence in the magnetic field and electron density can impact the Gaussian part of the signal captured in terms of the angular power spectra of the signal by more than an order of magnitude. Also, the presence of turbulence in different clusters will lead the temperature and polarization fluctuations around the cluster region to have varying non-Gaussian distribution, with peaks and tails different from the Gaussian statistics of the CMB anisotropy. This new numerical technique has made it possible to calculate also the non-Gaussian signals and can be used in future CMB analysis in synergy with X-ray and radio observations to unveil ALPs coupling with photons in the currently unexplored ranges, for the masses between about $10^{-14}$ eV--$10^{-11}$ eV.

astro-ph.CO

Experimental demonstration of the Bell-type inequalities for four qubit Dicke state using IBM Quantum Processing Units

Violation of the Bell-type inequalities is necessary to confirm the existence of nonlocality in nonclassical (entangled) states. We have designed a customized operator which is made of the sum of the Pauli matrices ($\sigma_x$, $\sigma_y$, and $\sigma_z$). We theoretically and experimentally investigate the violation of Bell-type inequalities using two- and four-qubit Dicke states on IBM Quantum Processing Units (QPUs). We compare two different state preparation methods for the four-qubit Dicke state -- gate-based and statevector-based -- and evaluate their performance on two IBM QPUs, \texttt{ibm\_kyiv} and \texttt{ibm\_sherbrook}. For the two-qubit case, we demonstrate clear violations of the CHSH inequality, with the highest observed Bell parameter reaching $2.821 \pm 0.0019$ using M3 error mitigation, which is within $0.7\sigma$ of the theoretical maximum $2\sqrt{2}$. In the four-qubit case, we employ a Bell-type inequality tailored for Dicke states and achieve a maximum violation of $2.607 \pm 0.029$ without the need for additional mitigation when using the statevector-based method. Our results reveal that advanced error mitigation techniques significantly enhance the observed violations in the gate-based method, while the statevector-based approach inherently yields more robust states with lower noise. This study highlights the critical role of state preparation and mitigation techniques in probing fundamental quantum correlations on near-term quantum hardware.

quant-ph

Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries

We introduce Michelangelo: a minimal, synthetic, and unleaked long-context reasoning evaluation for large language models which is also easy to automatically score. This evaluation is derived via a novel, unifying framework for evaluations over arbitrarily long contexts which measure the model's ability to do more than retrieve a single piece of information from its context. The central idea of the Latent Structure Queries framework (LSQ) is to construct tasks which require a model to ``chisel away'' the irrelevant information in the context, revealing a latent structure in the context. To verify a model's understanding of this latent structure, we query the model for details of the structure. Using LSQ, we produce three diagnostic long-context evaluations across code and natural-language domains intended to provide a stronger signal of long-context language model capabilities. We perform evaluations on several state-of-the-art models and demonstrate both that a) the proposed evaluations are high-signal and b) that there is significant room for improvement in synthesizing long-context information.

cs.CL

SpectrAx: Spectral Search of Axion-Like Particles Using Multi-Band Observations of Galaxy Clusters from SKA, SO, CMB-S4 and eROSITA

The existence of axions or Axion-Like Particles (ALPs) has been predicted by various Beyond Standard Model (BSM) theories, and the proposed photon-ALP interaction is one of the ways to probe them. Such an interaction will lead to photon-ALP resonant conversion in galaxy clusters, resulting in a polarized spectral distortion in the CMB along the cluster line of sight. The estimation of this signal from galaxy clusters requires an estimation of their electron density and magnetic field profiles, as well as their redshifts. We have developed a new Bayesian framework \texttt{SpectrAx} that can use observations from different electromagnetic bands such as radio, CMB, optical, and X-ray to infer the astrophysical properties of a galaxy cluster, such as cluster its redshift, electron density and magnetic field, along with the BSM physics such as ALPs. We use simulated redshifts in our analysis, but that can be obtained by cross-matching with optical surveys having overlapping sky regions with the galaxy clusters. Also, we use radial profiles that are motivated from observations of galaxy clusters at low redshifts. By using the simulated data corresponding to the ALP mass of $10^{-14}$ eV for upcoming CMB surveys such as Simons Observatory (SO) and CMB-S4 in combination with Square Kilometer Array (SKA) and extended ROentgen Survey with an Imaging Telescope Array (eROSITA) we demonstrate the capability in accurately inferring the ALPs coupling strength along with the radial profile of electron density and magnetic field from galaxy clusters. The application of this framework to the data from future surveys by combining SKA+SO+eROSITA and SKA+CMB-S4+eROSITA will make it possible for the first time to explore both astrophysics and BSM physics from low-redshift galaxy clusters using a multi-band approach.

astro-ph.CO

The Road Less Scheduled

Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We propose an approach that avoids the need for this stopping time by eschewing the use of schedules entirely, while exhibiting state-of-the-art performance compared to schedules across a wide family of problems ranging from convex problems to large-scale deep learning problems. Our Schedule-Free approach introduces no additional hyper-parameters over standard optimizers with momentum. Our method is a direct consequence of a new theory we develop that unifies scheduling and iterate averaging. An open source implementation of our method is available at https://github.com/facebookresearch/schedule_free. Schedule-Free AdamW is the core algorithm behind our winning entry to the MLCommons 2024 AlgoPerf Algorithmic Efficiency Challenge Self-Tuning track.

cs.LG

A Diffused Background from Axion-like Particles in the Microwave Sky

The nature of dark matter is an unsolved cosmological problem and axions are one of the weakly interacting cold dark matter candidates. Axions or ALPs (Axion-like particles) are pseudo-scalar bosons predicted by beyond-standard model theories. The weak coupling of ALPs with photons leads to the conversion of CMB photons to ALPs in the presence of a transverse magnetic field. If they have the same mass as the effective mass of a photon in a plasma, the resonant conversion would cause a polarized spectral distortion leading to temperature fluctuations with the distortion spectrum. The probability of resonant conversion depends on the properties of the cluster such as the magnetic field, electron density, and its redshift. We show that this kind of conversion can happen in numerous unresolved galaxy clusters up to high redshifts, which will lead to a diffused polarised anisotropy signal in the microwave sky. The spectrum of the signal and its shape in the angular scale will be different from the lensed CMB polarization signal. This new polarised distortion spectrum will be correlated with the distribution of clusters in the universe and hence, with the large-scale structure. The spectrum can then be probed using its spectral and spatial variation with respect to the CMB and various foregrounds. An SNR of $\sim$ 4.36 and $\sim$ 93.87 are possible in the CMB-S4 145 GHz band and CMB-HD 150 GHz band respectively for a photon-ALPs coupling strength of $\mathrm{g_{a \gamma} = 10^{-12} \, GeV^{-1}}$ using galaxy clusters beyond redshift z $= 1$. The same signal would lead to additional RMS fluctuations of $\sim \mathrm{7.5 \times 10^{-2} \, \mu K}$ at 145 GHz. In the absence of any signal, future CMB experiments such as Simons Observatory (SO), CMB-S4, and CMB-HD can put constraints on coupling strength better than current bounds from particle physics experiment CERN Axion Solar Telescope (CAST).

astro-ph.CO

A power spectrum approach to the search for Axion-like Particles from resolved galaxy clusters using CMB as a backlight

Axions or ALPs are hypothetical particles predicted by BSM theories, which make one of the dark matter candidates. These particles can convert into photons and vice-versa in the presence of magnetic field, with a probability decided by its coupling strength $\mathrm{g_{a\gamma}}$. One of the ways to detect these particles is using the CMB as a backlight. As the CMB photons pass through a galaxy cluster, they can get converted into ALPs in the mass range $10^{-15}$ eV to $10^{-11}$ eV through resonant conversion in the presence of cluster magnetic fields. This leads to a polarized spectral distortion ($\alpha$-distortion) in the CMB as the photon polarization parallel to the magnetic field in the galaxy cluster is involved in the conversion. The fluctuations in the magnetic field and electron density in a galaxy cluster lead to spatially varying $\alpha$-distortion around the cluster, with a power spectrum that is different from the lensed CMB polarization power spectrum for the standard model of cosmology. By measuring the difference in the polarization power spectrum around a galaxy cluster from the all-sky signal, one can find new $\alpha$-distortion in the sky. For galaxy clusters resolvable in multiple EM bands, one can measure the coupling strength $\mathrm{g_{a\gamma}}$ from the ALP power spectrum. Using multi-frequency techniques like ILC to clean the foregrounds, we show that the new power spectrum-based approach of the resolved galaxy clusters from upcoming CMB experiments such as Simons Observatory and CMB-S4 can detect (or put constraints) on the ALP-photon coupling strength of $\mathrm{g_{a\gamma} < 5.2 \times 10^{-12} \, GeV^{-1}}$ and $\mathrm{g_{a\gamma} < 3.6 \times 10^{-12} \, GeV^{-1}}$ at 95\% C.I. respectively for ALPs of masses $10^{-13}$ eV or for smaller $\mathrm{g_{a\gamma}}$ for lighter ALP masses (Abridged).

astro-ph.CO

Optimal Linear Decay Learning Rate Schedules and Further Refinements

Learning rate schedules used in practice bear little resemblance to those recommended by theory. We close much of this theory/practice gap, and as a consequence are able to derive new problem-adaptive learning rate schedules. Our main technical contribution is a refined analysis of learning rate schedules for a wide class of optimization algorithms (including SGD). When considering only worst-case analysis, our theory predicts that the optimal choice is the linear decay schedule where the step-size is set proportional to 1 - t/T, where t is the current iteration and T is the total number of steps. To go beyond this worst-case analysis, we use the observed gradient norms to derive schedules refined for any particular task. These refined schedules exhibit learning rate warm-up and rapid learning rate annealing near the end of training. Ours is the first systematic approach to automatically yield both of these properties. We perform the most comprehensive evaluation of learning rate schedules to date, evaluating across 10 diverse deep learning problems, a series of LLMs, and a suite of logistic regression problems. We validate that overall, the linear-decay schedule outperforms all commonly used default schedules including cosine annealing. Our adaptive schedule refinement method gives further improvements.

cs.LG

Mechanic: A Learning Rate Tuner

We introduce a technique for tuning the learning rate scale factor of any base optimization algorithm and schedule automatically, which we call \textsc{mechanic}. Our method provides a practical realization of recent theoretical reductions for accomplishing a similar goal in online convex optimization. We rigorously evaluate \textsc{mechanic} on a range of large scale deep learning tasks with varying batch sizes, schedules, and base optimization algorithms. These experiments demonstrate that depending on the problem, \textsc{mechanic} either comes very close to, matches or even improves upon manual tuning of learning rates.

cs.LG

Optimal Stochastic Non-smooth Non-convex Optimization through Online-to-Non-convex Conversion

We present new algorithms for optimizing non-smooth, non-convex stochastic objectives based on a novel analysis technique. This improves the current best-known complexity for finding a $(\delta,\epsilon)$-stationary point from $O(\epsilon^{-4}\delta^{-1})$ stochastic gradient queries to $O(\epsilon^{-3}\delta^{-1})$, which we also show to be optimal. Our primary technique is a reduction from non-smooth non-convex optimization to online learning, after which our results follow from standard regret bounds in online learning. For deterministic and second-order smooth objectives, applying more advanced optimistic online learning techniques enables a new complexity of $O(\epsilon^{-1.5}\delta^{-0.5})$. Our techniques also recover all optimal or best-known results for finding $\epsilon$ stationary points of smooth or second-order smooth objectives in both stochastic and deterministic settings.

cs.LG

Simplifying and Understanding State Space Models with Diagonal Linear RNNs

Sequence models based on linear state spaces (SSMs) have recently emerged as a promising choice of architecture for modeling long range dependencies across various modalities. However, they invariably rely on discretization of a continuous state space, which complicates their presentation and understanding. In this work, we dispose of the discretization step, and propose a model based on vanilla Diagonal Linear RNNs ($\mathrm{DLR}$). We empirically show that, despite being conceptually much simpler, $\mathrm{DLR}$ is as performant as previously-proposed SSMs on a variety of tasks and benchmarks including Long Range Arena and raw speech classification. Moreover, we characterize the expressivity of SSMs (including $\mathrm{DLR}$) and attention-based models via a suite of $13$ synthetic sequence-to-sequence tasks involving interactions over tens of thousands of tokens, ranging from simple operations, such as shifting an input sequence, to detecting co-dependent visual features over long spatial ranges in flattened images. We find that while SSMs report near-perfect performance on tasks that can be modeled via $\textit{few}$ convolutional kernels, they struggle on tasks requiring $\textit{many}$ such kernels and especially when the desired sequence manipulation is $\textit{context-dependent}$. Despite these limitations, $\mathrm{DLR}$ reaches high performance on two higher-order reasoning tasks $\mathrm{ListOpsSubTrees}$ and $\mathrm{PathfinderSegmentation}\text{-}\mathrm{256}$ with input lengths $8K$ and $65K$ respectively, and gives encouraging performance on $\mathrm{PathfinderSegmentation}\text{-}\mathrm{512}$ with input length $262K$ for which attention is not a viable choice.

cs.LG