Searcharxiv⌕ Search

arXiv subjects

Rahul Ramesh

Publications and source records attributed to Rahul Ramesh.

34 records · Page 2Linked to original sources

Zooming in on the circumgalactic medium: resolving small-scale gas structure with the GIBLE cosmological simulations

We introduce Project GIBLE (Gas Is Better resoLved around galaxiEs), a suite of cosmological zoom-in simulations where gas in the circumgalactic medium (CGM) is preferentially simulated at ultra-high numerical resolution. Our initial sample consists of eight galaxies, all selected as Milky Way-like galaxies at $z=0$ from the TNG50 simulation. Using the same galaxy formation model as IllustrisTNG, and the moving-mesh code AREPO, we re-simulate each of these eight galaxies maintaining a resolution equivalent to TNG50-2 ($m_{\rm{gas}}$ $\sim$ $8 \times 10^5 {\rm M}_{\odot}$). However, we use our super-Lagrangian refinement scheme to more finely resolve gas in the CGM around these galaxies. Our highest resolution runs achieve 512 times better mass resolution ($\sim$ $10^3 {\rm M}_{\odot}$). This corresponds to a median spatial resolution of $\sim$ $75$ pc at $0.15~R_{\rm{200,c}}$, which coarsens with increasing distance to $\sim$ $700$ pc at the virial radius. We make predictions for the covering fractions of several observational tracers of multi-phase CGM gas: HI, MgII, CIV and OVII. We then study the impact of improved resolution on small scale structure. While the abundance of the smallest cold, dense gas clouds continues to increase with improving resolution, the number of massive clouds is well converged. We conclude by quantifying small scale structure with the velocity structure function and the auto-correlation function of the density field, assessing their resolution dependence. The GIBLE cosmological hydrodynamical simulations enable us to improve resolution in a computationally efficient manner, thereby achieving numerical convergence of a subset of key CGM gas properties and observables.

astro-ph.GA↗

Prospective Learning: Principled Extrapolation to the Future

Learning is a process which can update decision rules, based on past experience, such that future performance improves. Traditionally, machine learning is often evaluated under the assumption that the future will be identical to the past in distribution or change adversarially. But these assumptions can be either too optimistic or pessimistic for many problems in the real world. Real world scenarios evolve over multiple spatiotemporal scales with partially predictable dynamics. Here we reformulate the learning problem to one that centers around this idea of dynamic futures that are partially learnable. We conjecture that certain sequences of tasks are not retrospectively learnable (in which the data distribution is fixed), but are prospectively learnable (in which distributions may be dynamic), suggesting that prospective learning is more difficult in kind than retrospective learning. We argue that prospective learning more accurately characterizes many real world problems that (1) currently stymie existing artificial intelligence solutions and/or (2) lack adequate explanations for how natural intelligences solve them. Thus, studying prospective learning will lead to deeper insights and solutions to currently vexing challenges in both natural and artificial intelligences.

cs.LG↗

The Value of Out-of-Distribution Data

We expect the generalization error to improve with more samples from a similar task, and to deteriorate with more samples from an out-of-distribution (OOD) task. In this work, we show a counter-intuitive phenomenon: the generalization error of a task can be a non-monotonic function of the number of OOD samples. As the number of OOD samples increases, the generalization error on the target task improves before deteriorating beyond a threshold. In other words, there is value in training on small amounts of OOD data. We use Fisher's Linear Discriminant on synthetic datasets and deep networks on computer vision benchmarks such as MNIST, CIFAR-10, CINIC-10, PACS and DomainNet to demonstrate and analyze this phenomenon. In the idealistic setting where we know which samples are OOD, we show that these non-monotonic trends can be exploited using an appropriately weighted objective of the target and OOD empirical risk. While its practical utility is limited, this does suggest that if we can detect OOD samples, then there may be ways to benefit from them. When we do not know which samples are OOD, we show how a number of go-to strategies such as data-augmentation, hyper-parameter optimization, and pre-training are not enough to ensure that the target generalization error does not deteriorate with the number of OOD samples in the dataset.

cs.LG↗

Milky Way and Andromeda analogs from the TNG50 simulation

We present the properties of Milky Way- and Andromeda-like (MW/M31-like) galaxies simulated within TNG50, the highest-resolution run of the IllustrisTNG suite of $Λ$CDM magneto-hydrodynamical simulations. We introduce our fiducial selection for MW/M31 analogs, which we propose for direct usage as well as for reference in future analyses. TNG50 contains 198 MW/M31 analogs, i.e. galaxies with stellar disky morphology, with a stellar mass in the range of $M_* = 10^{10.5 - 11.2}$ Msun, and within a MW-like Mpc-scale environment at z=0. These are resolved with baryonic (dark matter) mass resolution of $8.5\times10^4$ Msun ($4.5\times10^5$ Msun) and $\sim150$ pc of average spatial resolution in the star-forming regions: we therefore expand by many factors (2 orders of magnitude) the sample size of cosmologically-simulated analogs with similar ($\times 10$ better) numerical resolution. The majority of TNG50 MW/M31 analogs at $z=0$ exhibit a bar, 60 per cent are star-forming, the sample includes 3 Local Group (LG)-like systems, and a number of galaxies host one or more satellites as massive as e.g. the Magellanic Clouds. Even within such a relatively narrow selection, TNG50 reveals a great diversity in galaxy and halo properties, as well as in past histories. Within the TNG50 sample, it is possible to identify several simulated galaxies whose integral and structural properties are consistent, one or more at a time, with those measured for the Galaxy and Andromeda. With this paper, we document and release a series of broadly applicable data products that build upon the IllustrisTNG public release and aim to facilitate easy access and analysis by public users. These include datacubes across snapshots ($0 \le z \le 7$) for each TNG50 MW/M31-like galaxy, and a series of value-added catalogs that will be continually expanded to provide a convenient and up to date community resource.

astro-ph.GA↗

The Circumgalactic Medium of Milky Way-like Galaxies in the TNG50 Simulation -- II: Cold, Dense Gas Clouds and High-Velocity Cloud Analogs

We use the TNG50 simulation of the IllustrisTNG project to study cold, dense clouds of gas in the circumgalactic media (CGM) of Milky Way-like galaxies. We find that their CGM is typically filled with of order one hundred (thousand) reasonably (marginally) resolved clouds, possible analogs of high-velocity clouds (HVCs). There is a large variation in cloud abundance from galaxy to galaxy, and the physical properties of clouds that we explore -- mass, size, metallicity, pressure, and kinematics -- are also diverse. We quantify the distributions of cloud properties and cloud-background contrasts, providing cosmological inputs for idealized simulations. Clouds characteristically have sub-solar metallicities, diverse shapes, small overdensities ($χ= n_{\rm cold} / n_{\rm hot} \lesssim 10$), are mostly inflowing, and have sub-virial rotation. At TNG50 resolution, resolved clouds have median masses of $\sim 10^6\,\rm{M_\odot}$ and sizes of $\sim 10$ kpc. Larger clouds are well converged numerically, while the abundance of the smallest clouds increases with resolution, as expected. In TNG50 MW-like haloes, clouds are slightly (severely) under-pressurised relative to their surroundings with respect to total (thermal) pressure, implying that magnetic fields may be important. Clouds are not distributed uniformly throughout the CGM, but are clustered around other clouds, often near baryon-rich satellite galaxies. This suggests that at least some clouds originate from satellites, via direct ram-pressure stripping or otherwise. Finally, we compare with observations of intermediate and high velocity clouds from the real Milky Way halo. TNG50 shows a similar cloud velocity distribution as observations, and predicts a significant population of currently difficult-to-detect low velocity clouds.

astro-ph.GA↗

The Origin of Stars in the Inner 500 Parsecs in TNG50 Galaxies

We investigate the origin of stars in the innermost $500\,\mathrm{pc}$ of galaxies spanning stellar masses of $5\times10^{8-12}\,\mathrm{M}_{\odot}$ at $\mathrm{z=0}$ using the cosmological magnetohydrodynamical TNG50 simulation. Three different origins of stars comprise galactic centers: 1) in-situ (born in the center), 2) migrated (born elsewhere in the galaxy and ultimately moved to the center), 3) ex-situ (accreted from other galaxies). In-situ and migrated stars dominate the central stellar mass budget on average with 73% and 23% respectively. The ex-situ fraction rises above 1% for galaxies $\gtrsim10^{11}\,\mathrm{M}_{\odot}$. Yet, only 9% of all galaxies exhibit no ex-situ stars in their centers and the scatter of ex-situ mass is significant ($4-6\,\mathrm{dex}$). Migrated stars predominantly originate closely from the center ($1-2\,\mathrm{kpc}$), but if they travelled together in clumps distances reach $\sim10\,\mathrm{kpc}$. Central and satellite galaxies possess similar amounts and origins of central stars. Star forming galaxies ($\gtrsim10^{10}\,\mathrm{M}_{\odot}$) have on average more ex-situ mass in their centers than quenched ones. We predict readily observable stellar population and dynamical properties: 1) migrated stars are distinctly young ($\sim2\,\mathrm{Gyr}$) and rotationally supported, especially for Milky Way mass galaxies, 2) in-situ stars are most metal-rich and older than migrated stars, 3) ex-situ stars are on random motion dominated orbits and typically the oldest, most metal-poor and $α$-enhanced population. We demonstrate that the interaction history with other galaxies leads to diverse pathways of building up galaxy centers in a $Λ$CDM universe. Our work highlights the necessity for cosmological context in formation scenarios of central galactic components and the potential to use galaxy centers as tracers of overall galaxy assembly.

astro-ph.GA↗

The Circumgalactic Medium of Milky Way-like Galaxies in the TNG50 Simulation -- I: Halo Gas Properties and the Role of SMBH Feedback

We analyze the physical properties of gas in the circumgalactic medium (CGM) of 132 Milky Way (MW)-like central galaxies at $z=0$ from the cosmological magneto-hydrodynamical simulation TNG50, part of the IllustrisTNG project. The properties and abundance of CGM gas across the sample are diverse, and the fractional budgets of different phases (cold, warm, and hot), as well as neutral HI mass and metal mass, vary considerably. Over our stellar mass range of $10^{10.5} < M_\star / \rm{M}_\odot < 10^{10.9}$, radial profiles of gas physical properties from $0.15 < R\rm{ / R_{\rm 200c}} < 1.0$ reveal great CGM structural complexity, with significant variations both at fixed distance around individual galaxies, and across different galaxies. CGM gas is multi-phase: the distributions of density, temperature and entropy are all multimodal, while metallicity and thermal pressure distributions are unimodal; all are broad. We present predictions for magnetic fields in MW-like halos: a median field strength of $|B|\sim\,1μ$G in the inner halo decreases rapidly at larger distance, while magnetic pressure dominates over thermal pressure only within $\sim0.2 \times \rm{R_{200c}}$. Virial temperature gas at $\sim 10^6\,$K coexists with a sub-dominant cool, $< 10^5\,$K component in approximate pressure equilibrium. Finally, the physical properties of the CGM are tightly connected to the galactic star formation rate, in turn dependent on feedback from supermassive black holes (SMBHs). In TNG50, we find that energy from SMBH-driven kinetic winds generates high-velocity outflows ($\gtrsim 500-2000$ km/s), heats gas to super-virial temperatures ($> 10^{6.5-7}$ K), and regulates the net balance of inflows versus outflows in otherwise quasi-static gaseous halos.

astro-ph.GA↗

Model Zoo: A Growing "Brain" That Learns Continually

This paper argues that continual learning methods can benefit by splitting the capacity of the learner across multiple models. We use statistical learning theory and experimental analysis to show how multiple tasks can interact with each other in a non-trivial fashion when a single model is trained on them. The generalization error on a particular task can improve when it is trained with synergistic tasks, but can also deteriorate when trained with competing tasks. This theory motivates our method named Model Zoo which, inspired from the boosting literature, grows an ensemble of small models, each of which is trained during one episode of continual learning. We demonstrate that Model Zoo obtains large gains in accuracy on a variety of continual learning benchmark problems. Code is available at https://github.com/grasp-lyrl/modelzoo_continual.

cs.LG↗

Deep Reference Priors: What is the best way to pretrain a model?

What is the best way to exploit extra data -- be it unlabeled data from the same task, or labeled data from a related task -- to learn a given task? This paper formalizes the question using the theory of reference priors. Reference priors are objective, uninformative Bayesian priors that maximize the mutual information between the task and the weights of the model. Such priors enable the task to maximally affect the Bayesian posterior, e.g., reference priors depend upon the number of samples available for learning the task and for very small sample sizes, the prior puts more probability mass on low-complexity models in the hypothesis space. This paper presents the first demonstration of reference priors for medium-scale deep networks and image-based data. We develop generalizations of reference priors and demonstrate applications to two problems. First, by using unlabeled data to compute the reference prior, we develop new Bayesian semi-supervised learning methods that remain effective even with very few samples per class. Second, by using labeled data from the source task to compute the reference prior, we develop a new pretraining method for transfer learning that allows data from the target task to maximally affect the Bayesian posterior. Empirical validation of these methods is conducted on image classification datasets. Code is available at https://github.com/grasp-lyrl/deep_reference_priors.

stat.ML↗

Gravitational Lensing of Core Collapse Supernova Gravitational Wave Signals

We discuss the prospects of gravitational lensing of gravitational waves (GWs) coming from core-collapse supernovae (CCSN). As the CCSN GW signal can only be detected from within our own Galaxy and the local group by current and upcoming ground-based GW detectors, we focus on microlensing. We introduce a new technique based on analysis of the power spectrum and association of peaks of the power spectrum with the peaks of the amplification factor to identify lensed signals. We validate our method by applying it on the CCSN-like mock signals lensed by a point mass lens. We find that the lensed and unlensed signal can be differentiated using the association of peaks by more than one sigma for lens masses M$_{\rm L} {>} 150{\rm M}_{\odot}$. We also study the correlation integral between the power spectra and corresponding amplification factor. This statistical approach is able to differentiate between unlensed and lensed signals for lenses as small as M$_{\rm L} {\sim} 15{\rm M}_{\odot}$. Further, we demonstrate that this method can be used to estimate the mass of a lens in case the signal is lensed. The power spectrum based analysis is general and can be applied to any broad band signal and is especially useful for incoherent signals.

gr-qc↗

Wave Effects in Double-Plane Lensing

We discuss the wave optical effects in gravitational lens systems with two point mass lenses in two different lens planes. We identify and vary parameters (i.e., lens masses, related distances, and their alignments) related to the lens system to investigate their effects on the amplification factor. We find that due to a large number of parameters, it is not possible to make generalized statements regarding the amplification factor. We conclude by noting that the best approach to study two-plane and multi-plane lensing is to study various possible lens systems case by case in order to explore the possibilities in the parameter space instead of hoping to generalize the results of a few test cases. We present a preliminary analysis of the parameter space for a two-lens system here.

astro-ph.CO↗

Option Encoder: A Framework for Discovering a Policy Basis in Reinforcement Learning

Option discovery and skill acquisition frameworks are integral to the functioning of a Hierarchically organized Reinforcement learning agent. However, such techniques often yield a large number of options or skills, which can potentially be represented succinctly by filtering out any redundant information. Such a reduction can reduce the required computation while also improving the performance on a target task. In order to compress an array of option policies, we attempt to find a policy basis that accurately captures the set of all options. In this work, we propose Option Encoder, an auto-encoder based framework with intelligently constrained weights, that helps discover a collection of basis policies. The policy basis can be used as a proxy for the original set of skills in a suitable hierarchically organized framework. We demonstrate the efficacy of our method on a collection of grid-worlds and on the high-dimensional Fetch-Reach robotic manipulation task by evaluating the obtained policy basis on a set of downstream tasks.

cs.LG↗

Successor Options: An Option Discovery Framework for Reinforcement Learning

The options framework in reinforcement learning models the notion of a skill or a temporally extended sequence of actions. The discovery of a reusable set of skills has typically entailed building options, that navigate to bottleneck states. This work adopts a complementary approach, where we attempt to discover options that navigate to landmark states. These states are prototypical representatives of well-connected regions and can hence access the associated region with relative ease. In this work, we propose Successor Options, which leverages Successor Representations to build a model of the state space. The intra-option policies are learnt using a novel pseudo-reward and the model scales to high-dimensional spaces easily. Additionally, we also propose an Incremental Successor Options model that iterates between constructing Successor Representations and building options, which is useful when robust Successor Representations cannot be built solely from primitive actions. We demonstrate the efficacy of our approach on a collection of grid-worlds, and on the high-dimensional robotic control environment of Fetch.

cs.LG↗

FigureNet: A Deep Learning model for Question-Answering on Scientific Plots

Deep Learning has managed to push boundaries in a wide variety of tasks. One area of interest is to tackle problems in reasoning and understanding, with an aim to emulate human intelligence. In this work, we describe a deep learning model that addresses the reasoning task of question-answering on categorical plots. We introduce a novel architecture FigureNet, that learns to identify various plot elements, quantify the represented values and determine a relative ordering of these statistical values. We test our model on the FigureQA dataset which provides images and accompanying questions for scientific plots like bar graphs and pie charts, augmented with rich annotations. Our approach outperforms the state-of-the-art Relation Networks baseline by approximately $7\%$ on this dataset, with a training time that is over an order of magnitude lesser.

cs.LG↗

AUPCR Maximizing Matchings : Towards a Pragmatic Notion of Optimality for One-Sided Preference Matchings

We consider the problem of computing a matching in a bipartite graph in the presence of one-sided preferences. There are several well studied notions of optimality which include pareto optimality, rank maximality, fairness and popularity. In this paper, we conduct an in-depth experimental study comparing different notions of optimality based on a variety of metrics like cardinality, number of rank-1 edges, popularity, to name a few. Observing certain shortcomings in the standard notions of optimality, we propose an algorithm which maximizes an alternative metric called the Area under Profile Curve ratio (AUPCR). To the best of our knowledge, the AUPCR metric was used earlier but there is no known algorithm to compute an AUPCR maximizing matching. Finally, we illustrate the superiority of the AUPCR-maximizing matching by comparing its performance against other optimal matchings on synthetic instances modeling real-world data.

cs.MA↗

Learning to Factor Policies and Action-Value Functions: Factored Action Space Representations for Deep Reinforcement learning

Deep Reinforcement Learning (DRL) methods have performed well in an increasing numbering of high-dimensional visual decision making domains. Among all such visual decision making problems, those with discrete action spaces often tend to have underlying compositional structure in the said action space. Such action spaces often contain actions such as go left, go up as well as go diagonally up and left (which is a composition of the former two actions). The representations of control policies in such domains have traditionally been modeled without exploiting this inherent compositional structure in the action spaces. We propose a new learning paradigm, Factored Action space Representations (FAR) wherein we decompose a control policy learned using a Deep Reinforcement Learning Algorithm into independent components, analogous to decomposing a vector in terms of some orthogonal basis vectors. This architectural modification of the control policy representation allows the agent to learn about multiple actions simultaneously, while executing only one of them. We demonstrate that FAR yields considerable improvements on top of two DRL algorithms in Atari 2600: FARA3C outperforms A3C (Asynchronous Advantage Actor Critic) in 9 out of 14 tasks and FARAQL outperforms AQL (Asynchronous n-step Q-Learning) in 9 out of 13 tasks.

cs.LG↗