Searcharxiv⌕ Search

arXiv subjects

Yu Feng

Publications and source records attributed to Yu Feng.

At least 253 records · Page 14Linked to original sources

BlueTides simulation: establishing black hole-galaxy relations at high-redshift

The scaling relations between the mass of supermassive black holes ($M_{\bullet}$) and host galaxy properties (stellar mass, $M_{\star}$, and velocity dispersion, $σ$), provide a link between the growth of black holes (BHs) and that of their hosts. Here we investigate if and how the BH-galaxy relations are established in the high-$z$ universe using \textsc{BlueTides}, a high-resolution large volume cosmological hydrodynamic simulation. We find the $M_{\bullet}-M_{\star}$ and $M_{\bullet}-σ$ relations at $z=8$: $\log_{10}(M_{\bullet}) = 8.25 + 1.10 \ \log_{10}(M_{\star}/10^{11}M_{\odot})$ and $\log_{10}(M_{\bullet}) = 8.35 + 5.31 \ \log_{10}(σ/200kms^{-1})$ at $z=8$, both fully consistent with the local measurements. The slope of the $M_{\bullet}-σ$ relation is slightly steeper for high star formation rate and $M_{\star}$ galaxies while it remains unchanged as a function of Eddington accretion rate onto the BH. The intrinsic scatter in $M_{\bullet}-σ$ relation in all cases ($ε\sim 0.4$) is larger at these redshifts than inferred from observations and larger than in $M_{\bullet}-M_{\star}$ relation ($ε\sim 0.14$). We find the gas-to-stellar ratio $f=M_{\rm gas}/M_{\star}$ in the host (which can be very high at these redshifts) to have the most significant impact setting the intrinsic scatter of $M_{\bullet}-σ$. The scatter is significantly reduced when galaxies with high gas fractions ($ε= 0.28$ as $f<10$) are excluded (making the sample more comparable to low-$z$ galaxies); these systems have the largest star formation rates and black hole accretion rates, indicating that these fast-growing systems are still moving toward the relation at these high redshifts. Examining the evolution (from $z=10$ to 8) of high mass black holes in $M_{\bullet}-σ$ plane confirms this trend.

astro-ph.GA↗

The clustering of $z > 7$ galaxies: Predictions from the BLUETIDES simulation

We study the clustering of the highest-z galaxies (from ~ $0.1$ to a few tens Mpc scales) using the BLUETIDES simulation and compare it to current observational constraints from Hubble legacy and Hyper Suprime Cam (HSC) fields (at $z=6-7.2$). With a box length of $400$ $Mpc/h$ on each side and $0.7$ trillion particles, BLUETIDES is the largest high resolution cosmological hydrodynamic simulation to date ideally suited for studies of high-z galaxies. We find that galaxies with magnitude $m_{UV}<27.7$ have a bias ($b_g$) of $8.1\pm 1.2$ at $z=8$, and typical halo masses $M_H \gtrsim 6\times10^{10} M_{\odot}$. Given the redshift evolution between $z=8$ to $z=10$ ($b_g\propto(1+z)^{1.6}$), our inferred values of the bias and halo masses are consistent with measured angular clustering at $z \sim 6.8$ from these brighter samples. The bias of fainter galaxies (in the Hubble legacy field at $H_{160} \lesssim29.5$) is $5.9\pm0.9$ at $z=8$ corresponding to halo masses $M_H \gtrsim 10^{10} M_{\odot}$. We investigate directly the 1-halo term inthe clustering and show that it dominates on scales $r \lesssim 0.1$ Mpc/$h$ ($Θ\lesssim 3"$) with non-linear effect at transition scales between the 1-halo and 2-halo term affecting scales 0.1 $\lesssim r \lesssim $ 20 Mpc/$h$ ($3"\lesssim Θ\lesssim 90"$). Current clustering measurements probe down to the scales in the transition between 1-halo to 2-halo regime where non-linear effects are important. The amplitude of the 1-halo term implies that occupation numbers for satellites in \texttt{BLUETIDES} are somewhat higher than standard HODs adopted in these analyses (which predict amplitudes in the 1-halo regime suppressed by a factor 2-3). That possibly implies a higher number of galaxies detected by JWST (at small scales and even fainter magnitudes) observing these fields.

astro-ph.CO↗

nbodykit: an open-source, massively parallel toolkit for large-scale structure

We present nbodykit, an open-source, massively parallel Python toolkit for analyzing large-scale structure (LSS) data. Using Python bindings of the Message Passing Interface (MPI), we provide parallel implementations of many commonly used algorithms in LSS. nbodykit is both an interactive and scalable piece of scientific software, performing well in a supercomputing environment while still taking advantage of the interactive tools provided by the Python ecosystem. Existing functionality includes estimators of the power spectrum, 2 and 3-point correlation functions, a Friends-of-Friends grouping algorithm, mock catalog creation via the halo occupation distribution technique, and approximate N-body simulations via the FastPM scheme. The package also provides a set of distributed data containers, insulated from the algorithms themselves, that enable nbodykit to provide a unified treatment of both simulation and observational data sets. nbodykit can be easily deployed in a high performance computing environment, overcoming some of the traditional difficulties of using Python on supercomputers. We provide performance benchmarks illustrating the scalability of the software. The modular, component-based approach of nbodykit allows researchers to easily build complex applications using its tools. The package is extensively documented at http://nbodykit.readthedocs.io, which also includes an interactive set of example recipes for new users to explore. As open-source software, we hope nbodykit provides a common framework for the community to use and develop in confronting the analysis challenges of future LSS surveys.

astro-ph.IM↗

Program Synthesis using Conflict-Driven Learning

We propose a new conflict-driven program synthesis technique that is capable of learning from past mistakes. Given a spurious program that violates the desired specification, our synthesis algorithm identifies the root cause of the conflict and learns new lemmas that can prevent similar mistakes in the future. Specifically, we introduce the notion of equivalence modulo conflict and show how this idea can be used to learn useful lemmas that allow the synthesizer to prune large parts of the search space. We have implemented a general-purpose CDCL-style program synthesizer called Neo and evaluate it in two different application domains, namely data wrangling in R and functional programming over lists. Our experiments demonstrate the substantial benefits of conflict-driven learning and show that Neo outperforms two state-of-the-art synthesis tools, Morpheus and Deepcoder, that target these respective domains.

cs.PL↗

Dust Obscured Star Forming Galaxies in the Early Universe

Motivated by recent observational constraints on dust reprocessed emission in star forming galaxies at $z\sim 6$ and above we use the very-large cosmological hydrodynamical simulation \bluetides\ to explore predictions for the amount of dust obscured star formation in the early Universe ($z>8$). \bluetides\ matches current observational constraints on both the UV luminosity function and galaxy stellar mass function and predicts that approximately $90\%$ of the star formation in high-mass ($M_{*}>10^{10}\,{\rm M_{\odot}}$) galaxies at $z=8$ is already obscured by dust. The relationship between dust attenuation and stellar mass predicted by \bluetides\ is consistent with that observed at lower redshift. However, observations of several individual objects at $z>6$ are discrepant with the predictions, though it is possible their uncertainties may have been underestimated. We find that the predicted surface density of $z\ge 8$ sub-mm sources is below that accessible to current {\em Herschel}, SCUBA-2, and ALMA sub-mm surveys. However, as ALMA continues to accrue additional surface area the population of $z>8$ dust-obscured galaxies may become accessible in the near future.

astro-ph.GA↗

FastPM: a new scheme for fast simulations of dark matter and halos

We introduce FastPM, a highly-scalable approximated particle mesh N-body solver, which implements the particle mesh (PM) scheme enforcing correct linear displacement (1LPT) evolution via modified kick and drift factors. Employing a 2-dimensional domain decomposing scheme, FastPM scales extremely well with a very large number of CPUs. In contrast to COmoving-LAgrangian (COLA) approach, we do not require to split the force or track separately the 2LPT solution, reducing the code complexity and memory requirements. We compare FastPM with different number of steps ($N_s$) and force resolution factor ($B$) against 3 benchmarks: halo mass function from Friends of Friends halo finder, halo and dark matter power spectrum, and cross correlation coefficient (or stochasticity), relative to a high resolution TreePM simulation. We show that the modified time stepping scheme reduces the halo stochasticity when compared to COLA with the same number of steps and force resolution. While increasing $N_s$ and $B$ improves the transfer function and cross correlation coefficient, for many applications FastPM achieves sufficient accuracy at low $N_s$ and $B$. For example, $N_s=10$ and $B=2$ simulation provides a substantial saving (a factor of 10) of computing time relative to $N_s=40$, $B=3$ simulation, yet the halo benchmarks are very similar at $z=0$. We find that for abundance matched halos the stochasticity remains low even for $N_s=5$. FastPM compares well against less expensive schemes, being only 7 (4) times more expensive than 2LPT initial condition generator for $N_s=10$ ($N_s=5$). Some of the applications where FastPM can be useful are generating a large number of mocks, producing non-linear statistics where one varies a large number of nuisance or cosmological parameters, or serving as part of an initial conditions solver.

astro-ph.CO↗

Accurate halo-galaxy mocks from automatic bias estimation and particle mesh gravity solvers

Reliable extraction of cosmological information from clustering measurements of galaxy surveys requires estimation of the error covariance matrices of observables. The accuracy of covariance matrices is limited by our ability to generate sufficiently large number of independent mock catalogs that can describe the physics of galaxy clustering across a wide range of scales. Furthermore, galaxy mock catalogs are required to study systematics in galaxy surveys and to test analysis tools. In this investigation, we present a fast and accurate approach for generation of mock catalogs for the upcoming galaxy surveys. Our method relies on low-resolution approximate gravity solvers to simulate the large scale dark matter field, which we then populate with halos according to a flexible nonlinear and stochastic bias model. In particular, we extend the \textsc{patchy} code with an efficient particle mesh algorithm to simulate the dark matter field (the \textsc{FastPM} code), and with a robust MCMC method relying on the \textsc{emcee} code for constraining the parameters of the bias model. Using the halos in the BigMultiDark high-resolution $N$-body simulation as a reference catalog, we demonstrate that our technique can model the bivariate probability distribution function (counts-in-cells), power spectrum, and bispectrum of halos in the reference catalog. Specifically, we show that the new ingredients permit us to reach percentage accuracy in the power spectrum up to $k\sim 0.4\; \,h\,{\rm Mpc}^{-1}$ (within 5\% up to $k\sim 0.6\; \,h\,{\rm Mpc}^{-1}$) with accurate bispectra improving previous results based on Lagrangian perturbation theory.

astro-ph.CO↗

The descendants of the first quasars in the BlueTides simulation

Supermassive blackholes with masses of a billion solar masses or more are known to exist up to $z=7$. However, the present-day environments of the descendants of first quasars is not well understood and it is not known if they live in massive galaxy clusters or more isolated galaxies at $z=0$. We use a dark matter-only realization (BTMassTracer) of the BlueTides cosmological hydrodynamic simulation to study the halo properties of the descendants of the most massive black holes at $z=8$. We find that the descendants of the quasars with most massive black holes are not amongst the most massive halos. They reside in halos of with group-like ($\sim 10^{14}M_{\odot}$) masses, while the most massive halos in the simulations are rich clusters with masses $\sim 10^{15} M_{\odot}$. The distribution of halo masses at low redshift is similar to that of the descendants of least massive black holes, for a similar range of halo masses at $z=8$, which indicates that they are likely to exist in similar environments. By tracing back to the $z = 8$ progenitors of the most massive (cluster sized) halos at $z=0$; we find that their most likely black hole mass is less than $10^7 M_{\odot}$; they are clearly not amongst the most massive black holes. We also provide estimates for the likelihood of finding a high redshift quasar hosting a black hole with masses above $10^{7} M_{\odot}$ for a given halo mass at $z=0$. For halos above $10^{15} M_{\odot}$, there is only $20 \%$ probability that their $z=8$ progenitors hosted a black hole with mass above $10^{7} M_{\odot}$.

astro-ph.GA↗

Structure of spin excitations in heavily electron-doped Li0.8Fe0.2ODFeSe superconductors

Heavily electron-doped iron-selenide (HEDIS) high-transition-temperature (high-$T_{\rm{c}}$) superconductors, which have no hole Fermi pockets, but have a notably high $T_{\rm{c}}$, have challenged the prevailing $s$$_\pm$ pairing scenario originally proposed for iron pnictides containing both electron and hole pockets. The microscopic mechanism underlying the enhanced superconductivity in HEDIS remains unclear. Here, we used neutron scattering to study the spin excitations of the HEDIS material Li$_{0.8}$Fe$_{0.2}$ODFeSe ($T_{\rm{c}}$ = 41 K). Our data revealed nearly ring-shaped magnetic resonant excitations surrounding ($π$, $π$) at $\sim$ 21 meV. As the energy increased, the spin excitations assumed a diamond shape, and they dispersed outward until the energy reached $\sim$ 60 meV and then inward at higher energies. The observed energy-dependent momentum structure and twisted dispersion of spin excitations near ($π$, $π$) are analogous to those of hole-doped cuprates in several aspects, thus implying that such spin excitations are essential for the remarkably high $T_{\rm{c}}$ in these materials.

cond-mat.supr-con↗

Automated Synthesis of Semantic Malware Signatures using Maximum Satisfiability

This paper proposes a technique for automatically learning semantic malware signatures for Android from very few samples of a malware family. The key idea underlying our technique is to look for a maximally suspicious common subgraph (MSCS) that is shared between all known instances of a malware family. An MSCS describes the shared functionality between multiple Android applications in terms of inter-component call relations and their semantic metadata (e.g., data-flow properties). Our approach identifies such maximally suspicious common subgraphs by reducing the problem to maximum satisfiability. Once a semantic signature is learned, our approach uses a combination of static analysis and a new approximate signature matching algorithm to determine whether an Android application matches the semantic signature characterizing a given malware family. We have implemented our approach in a tool called ASTROID and show that it has a number of advantages over state-of-the-art malware detection techniques. First, we compare the semantic malware signatures automatically synthesized by ASTROID with manually-written signatures used in previous work and show that the signatures learned by ASTROID perform better in terms of accuracy as well as precision. Second, we compare ASTROID against two state-of-the-art malware detection tools and demonstrate its advantages in terms of interpretability and accuracy. Finally, we demonstrate that ASTROID's approximate signature matching algorithm is resistant to behavioral obfuscation and that it can be used to detect zero-day malware. In particular, we were able to find 22 instances of zero-day malware in Google Play that are not reported as malware by existing tools.

cs.CR↗

A fast algorithm for identifying Friends-of-Friends halos

We describe a simple and fast algorithm for identifying friends-of-friends features and prove its correctness. The algorithm avoids unnecessary expensive neighbor queries, uses minimal memory overhead, and rejects slowdown in high over-density regions. We define our algorithm formally based on pair enumeration, a problem that has been heavily studied in fast 2-point correlation codes and our reference implementation employs a dual KD-tree correlation function code. We construct features in a hierarchical tree structure, and use a splay operation to reduce the average cost of identifying the root of a feature from $O[\log L]$ to $O[1]$ ($L$ is the size of a feature) without additional memory costs. This reduces the overall time complexity of merging trees from $O[L\log L]$ to $O[L]$, reducing the number of operations per splay by orders of magnitude. We next introduce a pruning operation that skips merge operations between two fully self-connected KD-tree nodes. This improves the robustness of the algorithm, reducing the number of merge operations in high density peaks from $O[δ^2]$ to $O[δ]$. We show that for cosmological data set the algorithm eliminates more than half of merge operations for typically used linking lengths $b \sim 0.2$ (relative to mean separation). Furthermore, our algorithm is extremely simple and easy to implement on top of an existing pair enumeration code, reusing the optimization effort that has been invested in fast correlation function codes.

astro-ph.IM↗

The properties of the first galaxies in the BLUETIDES simulation

We employ the very large cosmological hydrodynamical simulation BLUETIDES to investigate the predicted properties of the galaxy population during the epoch of reionisation ($z>8$). BLUETIDES has a resolution and volume ($(400/h\approx 577)^{3}\,{\rm cMpc^3}$) providing a population of galaxies which is well matched to depth and area of current observational surveys targeting the high-redshift Universe. At $z=8$ BLUETIDES includes almost 160,000 galaxies with stellar masses $>10^{8}\,{\rm M_{\odot}}$. The population of galaxies predicted by BLUETIDES closely matches observational constraints on both the galaxy stellar mass function and far-UV ($150\,{\rm nm}$) luminosity function. Galaxies in BLUETIDES are characterised by rapidly increasing star formation histories. Specific star formation rates decrease with redshift though remain largely insensitive to stellar mass. As a result of the enhanced surface density of metals more massive galaxies are predicted to have higher dust attenuation resulting in a significant steepening of the observed far-UV luminosity function at high luminosities. The contribution of active SMBHs to the UV luminosities of galaxies with stellar masses $10^{9-10}\,{\rm M_{\odot}}$ is around $3\%$ on average. Approximately $25\%$ of galaxies with $M_{*}\approx 10^{10}\,{\rm M_{\odot}}$ are predicted to have active SMBH which contribute $>10\%$ of the total UV luminosity.

astro-ph.GA↗

Is the colour-octet mechanism consistent with the double $J/ψ$ production measurement at B-factories?

Double $J/ψ$ production in $e^+e^-$ collisions involving colour-octet channels are evaluated up to order $α^2α_s^3$. Having implemented the variation of the parameters ($m_c$, $μ_r$ and long-distance matrix elements), we found that the cross sections for producing double $J/ψ$ at B-factories range from $-0.016$fb to $0.245$fb, which are even much smaller than that via the colour-siglet mechanism. Accordingly, this result is consistent with the measurement by the Belle and BABAR Collaborations.

hep-ph↗

The DESI Experiment Part I: Science,Targeting, and Survey Design

DESI (Dark Energy Spectroscopic Instrument) is a Stage IV ground-based dark energy experiment that will study baryon acoustic oscillations (BAO) and the growth of structure through redshift-space distortions with a wide-area galaxy and quasar redshift survey. To trace the underlying dark matter distribution, spectroscopic targets will be selected in four classes from imaging data. We will measure luminous red galaxies up to $z=1.0$. To probe the Universe out to even higher redshift, DESI will target bright [O II] emission line galaxies up to $z=1.7$. Quasars will be targeted both as direct tracers of the underlying dark matter distribution and, at higher redshifts ($ 2.1 < z < 3.5$), for the Ly-$α$ forest absorption features in their spectra, which will be used to trace the distribution of neutral hydrogen. When moonlight prevents efficient observations of the faint targets of the baseline survey, DESI will conduct a magnitude-limited Bright Galaxy Survey comprising approximately 10 million galaxies with a median $z\approx 0.2$. In total, more than 30 million galaxy and quasar redshifts will be obtained to measure the BAO feature and determine the matter power spectrum, including redshift space distortions.

astro-ph.IM↗

The DESI Experiment Part II: Instrument Design

DESI (Dark Energy Spectropic Instrument) is a Stage IV ground-based dark energy experiment that will study baryon acoustic oscillations and the growth of structure through redshift-space distortions with a wide-area galaxy and quasar redshift survey. The DESI instrument is a robotically-actuated, fiber-fed spectrograph capable of taking up to 5,000 simultaneous spectra over a wavelength range from 360 nm to 980 nm. The fibers feed ten three-arm spectrographs with resolution $R= λ/Δλ$ between 2000 and 5500, depending on wavelength. The DESI instrument will be used to conduct a five-year survey designed to cover 14,000 deg$^2$. This powerful instrument will be installed at prime focus on the 4-m Mayall telescope in Kitt Peak, Arizona, along with a new optical corrector, which will provide a three-degree diameter field of view. The DESI collaboration will also deliver a spectroscopic pipeline and data management system to reduce and archive all data for eventual public use.

astro-ph.IM↗

Component-based Synthesis of Table Consolidation and Transformation Tasks from Examples

This paper presents an example-driven synthesis technique for automating a large class of data preparation tasks that arise in data science. Given a set of input tables and an out- put table, our approach synthesizes a table transformation program that performs the desired task. Our approach is not restricted to a fixed set of DSL constructs and can synthesize programs from an arbitrary set of components, including higher-order combinators. At a high-level, our approach performs type-directed enumerative search over partial pro- grams but incorporates two key innovations that allow it to scale: First, our technique can utilize any first-order specification of the components and uses SMT-based deduction to reject partial programs. Second, our algorithm uses partial evaluation to increase the power of deduction and drive enumerative search. We have evaluated our synthesis algorithm on dozens of data preparation tasks obtained from on-line forums, and we show that our approach can automatically solve a large class of problems encountered by R users.

cs.PL↗

Type-Directed Code Reuse using Integer Linear Programming

In many common scenarios, programmers need to implement functionality that is already provided by some third party library. This paper presents a tool called Hunter that facilitates code reuse by finding relevant methods in large code bases and automatically synthesizing any necessary wrapper code. The key technical idea underlying our approach is to use types to both improve search results and guide synthesis. Specifically, our method computes similarity metrics between types and uses this information to solve an integer linear programming (ILP) problem in which the objective is to minimize the cost of synthesis. We have implemented Hunter as an Eclipse plug-in and evaluate it by (a) comparing it against S6, a state-of-the-art code reuse tool, and (b) performing a user study. Our evaluation shows that Hunter compares favorably with S6 and significantly increases programmer productivity.

cs.SE↗

Impact of Baryonic Physics on Intrinsic Alignments

We explore the effects of specific assumptions in the subgrid models of star formation and stellar and AGN feedback on intrinsic alignments of galaxies in cosmological simulations of "MassiveBlack-II" family. Using smaller volume simulations, we explored the parameter space of the subgrid star formation and feedback model and found remarkable robustness of the observable statistical measures to the details of subgrid physics. The one observational probe most sensitive to modeling details is the distribution of misalignment angles. We hypothesize that the amount of angular momentum carried away by the galactic wind is the primary physical quantity that controls the orientation of the stellar distribution. Our results are also consistent with a similar study by the EAGLE simulation team.

astro-ph.CO↗