SearcharxivSearch

arXiv subjects

Rahul Shah

Publications and source records attributed to Rahul Shah.

At least 19 recordsLinked to original sources

Overview of 21cm Experiments at high redshift with SKAO

We provide an overview of the eight SKAO Science Book chapters that motivate the Epoch of Reionisation and Cosmic Dawn experiments with SKA-Low. We describe the individual SKA-Low experiments and expected sensitivity - power spectrum, tomography, 21-cm forest, cross-correlations, building on the broad observational plan laid out in the 2015 SKA Science Book. Finally, we outline features of the telescope that will be critical for the success of EoR/CD science, e.g., beam apodization, substations, and multi-beaming.

astro-ph.CO

Probing Interacting Dark Sectors with upcoming Post-Reionization and Galaxy Surveys

We investigate the constraining power of future post-reionization and galaxy surveys on possible interactions between dynamical dark energy and dark matter. The analysis focuses on the interaction strength and the dark energy equation of state parameters, in addition to the six standard cosmological parameters. Using fiducial values obtained from the current observational bounds (Planck 2018 + DESI DR2 + Pantheon+), mock datasets for upcoming 21-cm intensity mapping, galaxy clustering and cosmic shear observations from the SKA-mid, and for the upcoming large-scale survey from the Euclid mission, were generated. Subsequently, Markov chain Monte Carlo analyses combining current cosmological data with these mock datasets were performed to forecast parameter constraints. The results indicate that both SKA-mid and Euclid observations can significantly improve constraints on interacting dark sector parameters. In particular, the interaction strength and dark energy equation of state parameters can be constrained considerably tighter than current combined constraints from Planck 2018, DESI DR2 and Pantheon+. Comparing different probe combinations and survey configurations, it is found that SKA2 provides the tightest projected constraints, particularly on the interaction strength, while Euclid achieves a precision broadly comparable to that of SKA1. The results highlight the potential of these upcoming surveys to probe interactions within the dark sector.

astro-ph.CO

Generative Medical Event Models Improve with Scale

Realizing personalized medicine at scale calls for methods that distill insights from longitudinal patient journeys, which can be viewed as a sequence of medical events. Foundation models pretrained on large-scale medical event data represent a promising direction for scaling real-world evidence generation and generalizing to diverse downstream tasks. Using Epic Cosmos, a dataset with medical events from de-identified longitudinal health records for 16.3 billion encounters over 300 million unique patient records from 310 health systems, we introduce the Curiosity models, a family of decoder-only transformer models pretrained on 118 million patients representing 115 billion discrete medical events (151 billion tokens). We present the largest scaling-law study of medical event data, establishing a methodology for pretraining and revealing power-law scaling relationships for compute, tokens, and model size. Consequently, we pretrained a series of compute-optimal models with up to 1 billion parameters. Conditioned on a patient's real-world history, Curiosity autoregressively predicts the next medical event to simulate patient health timelines. We studied 78 real-world tasks, including diagnosis prediction, disease prognosis, and healthcare operations. Remarkably for a foundation model with generic pretraining and simulation-based inference, Curiosity generally outperformed or matched task-specific supervised models on these tasks, without requiring task-specific fine-tuning or few-shot examples. Curiosity's predictive power consistently improves as the model and pretraining scale. Our results show that Curiosity, a generative medical event foundation model, can effectively capture complex clinical dynamics, providing an extensible and generalizable framework to support clinical decision-making, streamline healthcare operations, and improve patient outcomes.

cs.LG

Interacting Dark Sectors in light of DESI DR2

Possible interaction between dark energy and dark matter has previously shown promise in alleviating the clustering tension, without exacerbating the Hubble tension, when Baryon Acoustic Oscillations (BAO) data from the Sloan Digital Sky Survey (SDSS) DR16 is combined with Cosmic Microwave Background (CMB) and Type-Ia Supernovae (SNIa) data sets. With the recent Dark Energy Spectroscopic Instrument (DESI) BAO DR2, there is now a compelling need to re-evaluate this scenario. We combine DESI DR2 with Planck 2018 and Pantheon+ SNIa data sets to constrain interacting dark matter dark energy models, accounting for interaction effects in both the background and perturbation sectors. Our results exhibit similar trends to those observed with SDSS, albeit with improved precision, reinforcing the consistency between the two BAO data sets. In addition to offering a resolution to the $S_8$ tension, in the phantom-limit, the dark energy equation of state exhibits an early-phantom behaviour, aligning with DESI DR2 findings, before transitioning to $w\sim-1$ at lower redshifts, regardless of the DE parametrization. However, the statistical significance of excluding $w=-1$ is reduced compared to their non-interacting counterparts.

astro-ph.CO

Discrete vs. Continuous Trade-offs for Generative Models

This work explores the theoretical and practical foundations of denoising diffusion probabilistic models (DDPMs) and score-based generative models, which leverage stochastic processes and Brownian motion to model complex data distributions. These models employ forward and reverse diffusion processes defined through stochastic differential equations (SDEs) to iteratively add and remove noise, enabling high-quality data generation. By analyzing the performance bounds of these models, we demonstrate how score estimation errors propagate through the reverse process and bound the total variation distance using discrete Girsanov transformations, Pinsker's inequality, and the data processing inequality (DPI) for an information theoretic lens.

cs.LG

Deep Learning Based Recalibration of SDSS and DESI BAO Alleviates Hubble and Clustering Tensions

Conventional calibration of Baryon Acoustic Oscillations (BAO) data relies on estimation of the sound horizon at drag epoch $r_d$ from early universe observations by assuming a cosmological model. We present a recalibration of two independent BAO datasets, SDSS and DESI, by employing deep learning techniques for model-independent estimation of $r_d$, and explore the impacts on $\Lambda$CDM cosmological parameters. Significant reductions in both Hubble ($H_0$) and clustering ($S_8$) tensions are observed for both the recalibrated datasets. Moderate shifts in some other parameters hint towards further exploration of such data-driven approaches.

astro-ph.CO

Optimizing Nepali PDF Extraction: A Comparative Study of Parser and OCR Technologies

This research compares PDF parsing and Optical Character Recognition (OCR) methods for extracting Nepali content from PDFs. PDF parsing offers fast and accurate extraction but faces challenges with non-Unicode Nepali fonts. OCR, specifically PyTesseract, overcomes these challenges, providing versatility for both digital and scanned PDFs. The study reveals that while PDF parsers are faster, their accuracy fluctuates based on PDF types. In contrast, OCRs, with a focus on PyTesseract, demonstrate consistent accuracy at the expense of slightly longer extraction times. Considering the project's emphasis on Nepali PDFs, PyTesseract emerges as the most suitable library, balancing extraction speed and accuracy.

cs.IR

Reconciling $S_8$: Insights from Interacting Dark Sectors

We do a careful investigation of the prospects of dark energy (DE) interacting with cold dark matter in alleviating the $S_8$ clustering tension. To this end, we consider various well-known parametrizations of the DE equation of state (EoS) and consider perturbations in both the dark sectors, along with an interaction term. Moreover, we perform a separate study for the phantom and non-phantom regimes. Using cosmic microwave background (CMB), baryon acoustic oscillations, and Type Ia supernovae data sets, constraints on the model parameters for each case have been obtained and a generic reduction in the $H_0-\sigma_{8,0}$ correlation has been observed, both for constant and dynamical DE EoS. This reduction, coupled with a significant negative correlation between the interaction term and $\sigma_{8,0}$, contributes to easing the clustering tension by lowering $\sigma_{8,0}$ to somewhere in between the early CMB and late-time clustering measurements for the phantom regime, for almost all the models under consideration. Additionally, this is achieved without exacerbating the Hubble tension. In this regard, the interacting Chevallier-Polarski-Linder and Jassal-Bagla-Padmanabhan models perform the best in relaxing the $S_8$ tension to $<1\sigma$. However, for the non-phantom regime the $\sigma_{8,0}$ tension tends to have worsened, which reassures the merits of phantom DE from latest data. We further investigate the role of redshift space distortion data sets and find an overall reduction in tension, with a $\sigma_{8,0}$ value relatively closer to the CMB value. We finally check whether further extensions of this scenario, such as the inclusion of the sound speed of DE and warm dark matter interacting with DE, can have some effects.

astro-ph.CO

LADDER: Revisiting the Cosmic Distance Ladder with Deep Learning Approaches and Exploring its Applications

We investigate the prospect of reconstructing the ''cosmic distance ladder'' of the Universe using a novel deep learning framework called LADDER - Learning Algorithm for Deep Distance Estimation and Reconstruction. LADDER is trained on the apparent magnitude data from the Pantheon Type Ia supernovae compilation, incorporating the full covariance information among data points, to produce predictions along with corresponding errors. After employing several validation tests with a number of deep learning models, we pick LADDER as the best performing one. We then demonstrate applications of our method in the cosmological context, including serving as a model-independent tool for consistency checks for other datasets like baryon acoustic oscillations, calibration of high-redshift datasets such as gamma ray bursts, and use as a model-independent mock catalog generator for future probes. Our analysis advocates for careful consideration of machine learning techniques applied to cosmological contexts.

astro-ph.CO

EIT: Earnest Insight Toolkit for Evaluating Students' Earnestness in Interactive Lecture Participation Exercises

In today's rapidly evolving educational landscape, traditional modes of passive information delivery are giving way to transformative pedagogical approaches that prioritize active student engagement. Within the context of large-scale hybrid classrooms, the challenge lies in fostering meaningful and active interaction between students and course content. This study delves into the significance of measuring students' earnestness during interactive lecture participation exercises. By analyzing students' responses to interactive lecture poll questions, establishing a clear rubric for evaluating earnestness, and conducting a comprehensive assessment, we introduce EIT (Earnest Insight Toolkit), a tool designed to assess students' engagement within interactive lecture participation exercises - particularly in the context of large-scale hybrid classrooms. Through the utilization of EIT, our objective is to equip educators with valuable means of identifying at-risk students for enhancing intervention and support strategies, as well as measuring students' levels of engagement with course content.

cs.CY

Role of Future SNIa Data from Rubin LSST in Reinvestigating Cosmological Models

We study how future Type-Ia supernovae (SNIa) standard candles detected by the Vera C. Rubin Observatory (LSST) can constrain some cosmological models. We use a realistic three-year SNIa simulated dataset generated by the LSST Dark Energy Science Collaboration (DESC) Time Domain pipeline, which includes a mix of spectroscopic and photometrically identified candidates. We combine this data with Cosmic Microwave Background (CMB) and Baryon Acoustic Oscillation (BAO) measurements to estimate the dark energy model parameters for two models -- the baseline $\Lambda$CDM and Chevallier-Polarski-Linder (CPL) dark energy parametrization. We compare them with the current constraints obtained from joint analysis of the latest real data from the Pantheon SNIa compilation, CMB from Planck 2018 and BAO. Our analysis finds tighter constraints on the model parameters along with a significant reduction of correlation between $H_0$ and $\sigma_{8,0}$. We find that LSST is expected to significantly improve upon the existing SNIa data in the critical analysis of cosmological models.

astro-ph.CO

Reconstructing the Hubble parameter with future Gravitational Wave missions using Machine Learning

We study the prospects of Gaussian processes (GP), a machine learning (ML) algorithm, as a tool to reconstruct the Hubble parameter $H(z)$ with two upcoming gravitational wave missions, namely the evolved Laser Interferometer Space Antenna (eLISA) and the Einstein Telescope (ET). Assuming various background cosmological models, the Hubble parameter has been reconstructed in a non-parametric manner with the help of GP using realistically generated catalogs for each mission. The effects of early-time and late-time priors on the reconstruction of $H(z)$, and hence on the Hubble constant ($H_0$), have also been focused on separately. Our analysis reveals that GP is quite robust in reconstructing the expansion history of the Universe within the observational window of the specific missions under consideration. We further confirm that both eLISA and ET would be able to provide constraints on $H(z)$ and $H_0$ which would be competitive to those inferred from current datasets. In particular, we observe that an eLISA run of $\sim10$-year duration with $\sim80$ detected bright siren events would be able to constrain $H_0$ as good as a $\sim3$-year ET run assuming $\sim 1000$ bright siren event detections. Further improvement in precision is expected for longer eLISA mission durations such as a $\sim15$-year time-frame having $\sim120$ events. Lastly, we discuss the possible role of these future gravitational wave missions in addressing the Hubble tension, for each model, on a case-by-case basis.

astro-ph.CO

A thorough investigation of the prospects of eLISA in addressing the Hubble tension: Fisher Forecast, MCMC and Machine Learning

We carry out an in-depth analysis of the capability of the upcoming space-based gravitational wave mission eLISA in addressing the Hubble tension, with a primary focus on observations at intermediate redshifts ($3<z<8$). We consider six different parametrizations representing different classes of cosmological models, which we constrain using the latest datasets of cosmic microwave background (CMB), baryon acoustic oscillations (BAO), and type Ia supernovae (SNIa) observations, in order to find out the up-to-date tensions with direct measurement data. Subsequently, these constraints are used as fiducials to construct mock catalogs for eLISA. We then employ Fisher analysis to forecast the future performance of each model in the context of eLISA. We further implement traditional Markov Chain Monte Carlo (MCMC) to estimate the parameters from the simulated catalogs. Finally, we utilize Gaussian Processes (GP), a machine learning algorithm, for reconstructing the Hubble parameter directly from simulated data. Based on our analysis, we present a thorough comparison of the three methods as forecasting tools. Our Fisher analysis confirms that eLISA would constrain the Hubble constant ($H_0$) at the sub-percent level. MCMC/GP results predict reduced tensions for models/fiducials which are currently harder to reconcile with direct measurements of $H_0$, whereas no significant change occurs for models/fiducials at lesser tensions with the latter. This feature warrants further investigation in this direction.

astro-ph.CO

Parameterized Pattern Matching -- Succinctly

We consider the $Parameterized$ $Pattern$ $Matching$ problem, where a pattern $P$ matches some location in a text $\mathsf{T}$ iff there is a one-to-one correspondence between the alphabet symbols of the pattern to those of the text. More specifically, assume that the text $\mathsf{T}$ contains $n$ characters from a static alphabet $Σ_s$ and a parameterized alphabet $Σ_p$, where $Σ_s \cap Σ_p = \varnothing$ and $|Σ_s \cup Σ_p|=σ$. A pattern $P$ matches a substring $S$ of $\mathsf{T}$ iff the static characters match exactly, and there exists a one-to-one function that renames the parameterized characters in $S$ to that in $P$. Previous indexing solution [Baker, STOC 1993], known as $Parameterized$ $Suffix$ $Tree$, requires $Θ(n\log n)$ bits of space, and can find all $occ$ occurrences of $P$ in $\mathcal{O}(|P|\log σ+ occ)$ time. In this paper, we present the first succinct index that occupies $n \log σ+ \mathcal{O}(n)$ bits and answers queries in $\mathcal{O}((|P|+ occ\cdot \log n) \logσ\log \log σ)$ time. We also present a compact index that occupies $\mathcal{O}(n\logσ)$ bits and answers queries in $\mathcal{O}(|P|\log σ+ occ\cdot \log n)$ time. Furthermore, the techniques are extended to obtain the first succinct representation of the index of Shibuya for $Structural$ $Matching$ [SWAT, 2000], and of Idury and Schäffer for $Parameterized$ $Dictionary$ $Matching$ [CPM, 1994].

cs.DS

Probabilistic Threshold Indexing for Uncertain Strings

Strings form a fundamental data type in computer systems. String searching has been extensively studied since the inception of computer science. Increasingly many applications have to deal with imprecise strings or strings with fuzzy information in them. String matching becomes a probabilistic event when a string contains uncertainty, i.e. each position of the string can have different probable characters with associated probability of occurrence for each character. Such uncertain strings are prevalent in various applications such as biological sequence data, event monitoring and automatic ECG annotations. We explore the problem of indexing uncertain strings to support efficient string searching. In this paper we consider two basic problems of string searching, namely substring searching and string listing. In substring searching, the task is to find the occurrences of a deterministic string in an uncertain string. We formulate the string listing problem for uncertain strings, where the objective is to output all the strings from a collection of strings, that contain probable occurrence of a deterministic query string. Indexing solution for both these problems are significantly more challenging for uncertain strings than for deterministic strings. Given a construction time probability value $τ$, our indexes can be constructed in linear space and supports queries in near optimal time for arbitrary values of probability threshold parameter greater than $τ$. To the best of our knowledge, this is the first indexing solution for searching in uncertain strings that achieves strong theoretical bound and supports arbitrary values of probability threshold parameter. We also propose an approximate substring search index that can answer substring search queries with an additive error in optimal time. We conduct experiments to evaluate the performance of our indexes.

cs.DB

On Optimal Top-K String Retrieval

Let ${\cal{D}}$ = $\{d_1, d_2, d_3, ..., d_D\}$ be a given set of $D$ (string) documents of total length $n$. The top-$k$ document retrieval problem is to index $\cal{D}$ such that when a pattern $P$ of length $p$, and a parameter $k$ come as a query, the index returns the $k$ most relevant documents to the pattern $P$. Hon et. al. \cite{HSV09} gave the first linear space framework to solve this problem in $O(p + k\log k)$ time. This was improved by Navarro and Nekrich \cite{NN12} to $O(p + k)$. These results are powerful enough to support arbitrary relevance functions like frequency, proximity, PageRank, etc. In many applications like desktop or email search, the data resides on disk and hence disk-bound indexes are needed. Despite of continued progress on this problem in terms of theoretical, practical and compression aspects, any non-trivial bounds in external memory model have so far been elusive. Internal memory (or RAM) solution to this problem decomposes the problem into $O(p)$ subproblems and thus incurs the additive factor of $O(p)$. In external memory, these approaches will lead to $O(p)$ I/Os instead of optimal $O(p/B)$ I/O term where $B$ is the block-size. We re-interpret the problem independent of $p$, as interval stabbing with priority over tree-shaped structure. This leads us to a linear space index in external memory supporting top-$k$ queries (with unsorted outputs) in near optimal $O(p/B + \log_B n + \log^{(h)} n + k/B)$ I/Os for any constant $h${$\log^{(1)}n =\log n$ and $\log^{(h)} n = \log (\log^{(h-1)} n)$}. Then we get $O(n\log^*n)$ space index with optimal $O(p/B+\log_B n + k/B)$ I/Os.

cs.DS

Towards an Optimal Space-and-Query-Time Index for Top-k Document Retrieval

Let $\D = $$ \{d_1,d_2,...d_D\}$ be a given set of $D$ string documents of total length $n$, our task is to index $\D$, such that the $k$ most relevant documents for an online query pattern $P$ of length $p$ can be retrieved efficiently. We propose an index of size $|CSA|+n\log D(2+o(1))$ bits and $O(t_{s}(p)+k\log\log n+poly\log\log n)$ query time for the basic relevance metric \emph{term-frequency}, where $|CSA|$ is the size (in bits) of a compressed full text index of $\D$, with $O(t_s(p))$ time for searching a pattern of length $p$ . We further reduce the space to $|CSA|+n\log D(1+o(1))$ bits, however the query time will be $O(t_s(p)+k(\log σ\log\log n)^{1+ε}+poly\log\log n)$, where $σ$ is the alphabet size and $ε>0$ is any constant.

cs.DS

Fully Dynamic Data Structure for Top-k Queries on Uncertain Data

Top-$k$ queries allow end-users to focus on the most important (top-$k$) answers amongst those which satisfy the query. In traditional databases, a user defined score function assigns a score value to each tuple and a top-$k$ query returns $k$ tuples with the highest score. In uncertain database, top-$k$ answer depends not only on the scores but also on the membership probabilities of tuples. Several top-$k$ definitions covering different aspects of score-probability interplay have been proposed in recent past~\cite{R10,R4,R2,R8}. Most of the existing work in this research field is focused on developing efficient algorithms for answering top-$k$ queries on static uncertain data. Any change (insertion, deletion of a tuple or change in membership probability, score of a tuple) in underlying data forces re-computation of query answers. Such re-computations are not practical considering the dynamic nature of data in many applications. In this paper, we propose a fully dynamic data structure that uses ranking function $PRF^e(α)$ proposed by Li et al.~\cite{R8} under the generally adopted model of $x$-relations~\cite{R11}. $PRF^e$ can effectively approximate various other top-$k$ definitions on uncertain data based on the value of parameter $α$. An $x$-relation consists of a number of $x$-tuples, where $x$-tuple is a set of mutually exclusive tuples (up to a constant number) called alternatives. Each $x$-tuple in a relation randomly instantiates into one tuple from its alternatives. For an uncertain relation with $N$ tuples, our structure can answer top-$k$ queries in $O(k\log N)$ time, handles an update in $O(\log N)$ time and takes $O(N)$ space. Finally, we evaluate practical efficiency of our structure on both synthetic and real data.

cs.DB