Searcharxiv⌕ Search

arXiv subjects

Yu Feng

Publications and source records attributed to Yu Feng.

At least 181 records · Page 10Linked to original sources

Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent

Distributed Deep Learning (DDL) is essential for large-scale Deep Learning (DL) training. Synchronous Stochastic Gradient Descent (SSGD) 1 is the de facto DDL optimization method. Using a sufficiently large batch size is critical to achieving DDL runtime speedup. In a large batch setting, the learning rate must be increased to compensate for the reduced number of parameter updates. However, a large learning rate may harm convergence in SSGD and training could easily diverge. Recently, Decentralized Parallel SGD (DPSGD) has been proposed to improve distributed training speed. In this paper, we find that DPSGD not only has a system-wise run-time benefit but also a significant convergence benefit over SSGD in the large batch setting. Based on a detailed analysis of the DPSGD learning dynamics, we find that DPSGD introduces additional landscape-dependent noise that automatically adjusts the effective learning rate to improve convergence. In addition, we theoretically show that this noise smoothes the loss landscape, hence allowing a larger learning rate. We conduct extensive studies over 18 state-of-the-art DL models/tasks and demonstrate that DPSGD often converges in cases where SSGD diverges for large learning rates in the large batch setting. Our findings are consistent across two different application domains: Computer Vision (CIFAR10 and ImageNet-1K) and Automatic Speech Recognition (SWB300 and SWB2000), and two different types of neural network models: Convolutional Neural Networks and Long Short-Term Memory Recurrent Neural Networks.

cs.LG↗

Not all peaks are created equal: the early growth of Supermassive Black Holes

In this work, we use the constrained Gaussian realization technique to study the early growth of supermassive black holes (SMBHs) in cosmological hydrodynamic simulations, exploring its relationship with features of the initial density peaks on large scales, ~1 Mpc/h. Our constrained simulations of volume (20 Mpc/h)^3 successfully reconstruct the large-scale structure as well as the black hole growth for the hosts of the rare 10^9 Msun SMBHs found in the BlueTides simulation at z~7. We run a set of simulations with constrained initial conditions by imposing a 5 σ_0(R_G) peak on scale of R_G = 1 Mpc/h varying different peak features, such as the shape and compactness as well as the tidal field surrounding the peak. We find that initial density peaks with high compactness and low tidal field induce the most rapid BH growth at early epochs. This is because compact density peaks with a more spherical large scale matter distribution lead to the formation of high density gas clumps in the centers of halos, and thus boost early BH accretion. Moreover, such initially compact density peaks in low tidal field regions also lead to a more compact BH host galaxy morphology. This can explain the tight correlation between BH growth and host galaxy compactness seen in observations.

astro-ph.GA↗

Global existence of a non-local semilinear parabolic equation with advection and applications to shear flow

In this paper, we consider the following non-local semi-linear parabolic equation with advection: for $1 \le p<1+\frac{2}{N}$, \begin{equation*} \begin{cases} u_t+v \cdot \nabla u-Δu=|u|^p-\int_{\mathbb T^N} |u|^p \quad & \textrm{on} \quad \mathbb T^N, \\ \\ u \ \textrm{periodic} \quad & \textrm{on} \quad \partial \mathbb T^N \end{cases} \end{equation*} with initial data $u_0$ defined on $\mathbb T^N$. Here $v$ is an incompressible flow, and $\mathbb T^N=[0, 1]^N$ is the $N$-torus with $N$ being the dimension. We first prove the local existence of mild solutions to the above equation for arbitrary data in $L^2$. We then study the global existence of the solutions under the following two scenarios: (1). when $v$ is a mixing flow; (2). when $v$ is a shear flow. More precisely, we show that under these assumptions, there exists a global solution to the above equation in the sense of $L^2$.

math.AP↗

AI-assisted super-resolution cosmological simulations II: Halo substructures, velocities and higher order statistics

In this work, we expand and test the capabilities of our recently developed super-resolution (SR) model to generate high-resolution (HR) realizations of the full phase-space matter distribution, including both displacement and velocity, from computationally cheap low-resolution (LR) cosmological N-body simulations. The SR model enhances the simulation resolution by generating 512 times more tracer particles, extending into the deeply non-linear regime where complex structure formation processes take place. We validate the SR model by deploying the model in 10 test simulations of box size 100 Mpc/h, and examine the matter power spectra, bispectra and 2D power spectra in redshift space. We find the generated SR field matches the true HR result at percent level down to scales of k ~ 10 h/Mpc. We also identify and inspect dark matter halos and their substructures. Our SR model generate visually authentic small-scale structures, that cannot be resolved by the LR input, and are in good statistical agreement with the real HR results. The SR model performs satisfactorily on the halo occupation distribution, halo correlations in both real and redshift space, and the pairwise velocity distribution, matching the HR results with comparable scatter, thus demonstrating its potential in making mock halo catalogs. The SR technique can be a powerful and promising tool for modelling small-scale galaxy formation physics in large cosmological volumes.

astro-ph.CO↗

Fast and Accurate: Video Enhancement using Sparse Depth

This paper presents a general framework to build fast and accurate algorithms for video enhancement tasks such as super-resolution, deblurring, and denoising. Essential to our framework is the realization that the accuracy, rather than the density, of pixel flows is what is required for high-quality video enhancement. Most of prior works take the opposite approach: they estimate dense (per-pixel)-but generally less robust-flows, mostly using computationally costly algorithms. Instead, we propose a lightweight flow estimation algorithm; it fuses the sparse point cloud data and (even sparser and less reliable) IMU data available in modern autonomous agents to estimate the flow information. Building on top of the flow estimation, we demonstrate a general framework that integrates the flows in a plug-and-play fashion with different task-specific layers. Algorithms built in our framework achieve 1.78x - 187.41x speedup while providing a 0.42 dB - 6.70 dB quality improvement over competing methods.

cs.CV↗

Field-Tuned Quantum Effects in a Triangular-Lattice Ising Magnet

We report thermodynamic and neutron scattering measurements of the triangular-lattice quantum Ising magnet TmMgGaO 4 in longitudinal magnetic fields. Our experiments reveal a quasi-plateau state induced by quantum fluctuations. This state exhibits an unconventional non-monotonic field and temperature dependence of the magnetic order and excitation gap. In the high field regime where the quantum fluctuations are largely suppressed, we observed a disordered state with coherent magnon-like excitations despite the suppression of the spin excitation intensity. Through detailed semi-classical calculations, we are able to understand these behaviors quantitatively from the subtle competition between quantum fluctuations and frustrated Ising interactions.

cond-mat.str-el↗

Suppression of epitaxial thin film growth by mixing

We consider following fourth-order parabolic equation with gradient nonlinearity on the two-dimensional torus with and without advection of an incompressible vector field in the case $2<p<3$: \begin{equation*} \partial_t u + (-Δ)^2 u = -\nabla\cdot(|\nabla u|^{p-2}\nabla u). \end{equation*} The study of this form of equations arises from mathematical models that simulate the epitaxial growth of the thin film. We prove the local existence of mild solutions for any initial data lies in $L^2$ in both cases. Our main result is: in the advective case, if the imposed advection is sufficiently mixing, then the global existence of solution can be proved, and the solution will converge exponentially to a homogeneous mixed state. While in the absence of advection, there exist initial data in $H^2\cap W^{1,\infty}$ such that the solution will blow up in finite time.

math.AP↗

The Quijote simulations

The Quijote simulations are a set of 44,100 full N-body simulations spanning more than 7,000 cosmological models in the $\{Ω_{\rm m}, Ω_{\rm b}, h, n_s, σ_8, M_ν, w \}$ hyperplane. At a single redshift the simulations contain more than 8.5 trillions of particles over a combined volume of 44,100 $(h^{-1}{\rm Gpc})^3$; each simulation follow the evolution of $256^3$, $512^3$ or $1024^3$ particles in a box of $1~h^{-1}{\rm Gpc}$ length. Billions of dark matter halos and cosmic voids have been identified in the simulations, whose runs required more than 35 million core hours. The Quijote simulations have been designed for two main purposes: 1) to quantify the information content on cosmological observables, and 2) to provide enough data to train machine learning algorithms. In this paper we describe the simulations and show a few of their applications. We also release the Petabyte of data generated, comprising hundreds of thousands of simulation snapshots at multiple redshifts, halo and void catalogs, together with millions of summary statistics such as power spectra, bispectra, correlation functions, marked power spectra, and estimated probability density functions.

astro-ph.CO↗

Polarized neutron scattering studies of magnetic excitations in iron-selenide superconductor (Li$_{0.8}$Fe$_{0.2}$)ODFeSe ($T_c$ = 41 K)

We report polarized neutron scattering measurements of the low energy spin fluctuations of the iron-selenide superconductor Li$_{0.8}$Fe$_{0.2}$ODFeSe below and above its superconducting transition temperature $T_c=41$ K. Our experiments confirmed that the resonance mode near 21 meV is magnetic. Moreover, the spin excitations are essentially isotropic in spin space at 5$\leq E\leq$ 29 meV in the superconducting and normal states. Our results suggest that the resonance mode in iron-based superconductors becomes isotropic when the influence of spin-orbit coupling and magnetic/nematic order is minimized, similar to those observed in cuprate superconductors.

cond-mat.supr-con↗

AI-assisted super-resolution cosmological simulations

Cosmological simulations of galaxy formation are limited by finite computational resources. We draw from the ongoing rapid advances in Artificial Intelligence (specifically Deep Learning) to address this problem. Neural networks have been developed to learn from high-resolution (HR) image data, and then make accurate super-resolution (SR) versions of different low-resolution (LR) images. We apply such techniques to LR cosmological N-body simulations, generating SR versions. Specifically, we are able to enhance the simulation resolution by generating 512 times more particles and predicting their displacements from the initial positions. Therefore our results can be viewed as new simulation realizations themselves rather than projections, e.g., to their density fields. Furthermore, the generation process is stochastic, enabling us to sample the small-scale modes conditioning on the large-scale environment. Our model learns from only 16 pairs of small-volume LR-HR simulations, and is then able to generate SR simulations that successfully reproduce the HR matter power spectrum to percent level up to $16\,h^{-1}\mathrm{Mpc}$, and the HR halo mass function to within $10 \%$ down to $10^{11} \, M_\odot$. We successfully deploy the model in a box 1000 times larger than the training simulation box, showing that high-resolution mock surveys can be generated rapidly. We conclude that AI assistance has the potential to revolutionize modeling of small-scale galaxy formation physics in large cosmological volumes.

astro-ph.CO↗

Dynamical Friction Modeling of Massive Black Holes in Cosmological Simulations and Effects on Merger Rate Predictions

In this work we establish and test methods for implementing dynamical friction for massive black hole pairs that form in large volume cosmological hydrodynamical simulations which include galaxy formation and black hole growth. We verify our models and parameters both for individual black hole dynamics and for the black hole population in cosmological volumes. Using our model of dynamical friction (DF) from collisionless particles, black holes can effectively sink close to the galaxy center, provided that the black hole's dynamical mass is at least twice that of the lowest mass resolution particles in the simulation. Gas drag also plays a role in assisting the black holes' orbital decay, but it is typically less effective than that from collisionless particles, especially after the first billion years of the black hole's evolution. DF from gas becomes less than $1\%$ of DF from collisionless particles for BH masses $> 10^{7}$ M$_{\odot}$. Using our best DF model, we calculate the merger rate down to $z=1.1$ using an $L_{\rm box}=35$ Mpc$/h$ simulation box. We predict $\sim 2$ mergers per year for $z>1.1$ peaking at $z\sim 2$. These merger rates are within the range obtained in previous work using similar-resolution hydro-dynamical simulations. We show that the rate is enhanced by factor of $\sim 2$ when DF is taken into account in the simulations compared to the no-DF run. This is due to $>40\%$ more black holes reaching the center of their host halo when DF is added.

astro-ph.GA↗

Falx: Synthesis-Powered Visualization Authoring

Modern visualization tools aim to allow data analysts to easily create exploratory visualizations. When the input data layout conforms to the visualization design, users can easily specify visualizations by mapping data columns to visual channels of the design. However, when there is a mismatch between data layout and the design, users need to spend significant effort on data transformation. We propose Falx, a synthesis-powered visualization tool that allows users to specify visualizations in a similarly simple way but without needing to worry about data layout. In Falx, users specify visualizations using examples of how concrete values in the input are mapped to visual channels, and Falx automatically infers the visualization specification and transforms the data to match the design. In a study with 33 data analysts on four visualization tasks involving data transformation, we found that users can effectively adopt Falx to create visualizations they otherwise cannot implement.

cs.HC↗

Phases of learning dynamics in artificial neural networks: with or without mislabeled data

Despite tremendous success of deep neural network in machine learning, the underlying reason for its superior learning capability remains unclear. Here, we present a framework based on statistical physics to study dynamics of stochastic gradient descent (SGD) that drives learning in neural networks. By using the minibatch gradient ensemble, we construct order parameters to characterize dynamics of weight updates in SGD. Without mislabeled data, we find that the SGD learning dynamics transitions from a fast learning phase to a slow exploration phase, which is associated with large changes in order parameters that characterize the alignment of SGD gradients and their mean amplitude. In the case with randomly mislabeled samples, SGD learning dynamics falls into four distinct phases. The system first finds solutions for the correctly labeled samples in phase I, it then wanders around these solutions in phase II until it finds a direction to learn the mislabeled samples during phase III, after which it finds solutions that satisfy all training samples during phase IV. Correspondingly, the test error decreases during phase I and remains low during phase II; however, it increases during phase III and reaches a high plateau during phase IV. The transitions between different phases can be understood by changes of order parameters that characterize the alignment of mean gradients for the correctly and incorrectly labeled samples and their (relative) strength during learning. We find that individual sample losses for the two datasets are most separated during phase II, which leads to a cleaning process to eliminate mislabeled samples for improving generalization.

cs.LG↗

A fast particle-mesh simulation of non-linear cosmological structure formation with massive neutrinos

Quasi-N-body simulations, such as FastPM, provide a fast way to simulate cosmological structure formation, but have yet to adequately include the effects of massive neutrinos. We present a method to include neutrino particles in FastPM, enabling computation of the CDM and total matter power spectra to percent-level accuracy in the non-linear regime. The CDM-neutrino cross-power can also be computed at a sufficient accuracy to constrain cosmological observables. To avoid the shot noise that typically plagues neutrino particle simulations, we employ a quasi-random algorithm to sample the relevant Fermi-Dirac distribution when setting the initial neutrino thermal velocities. We additionally develop an effective distribution function to describe a set of non-degenerate neutrinos as a single particle to speed up non-degenerate simulations. The simulation is accurate for the full range of physical interest, $M_ν\lesssim 0.6$eV, and applicable to redshifts $z\lesssim2$. Such accuracy can be achieved by initializing particles with the two-fluid approximation transfer functions (using the REPS package). Convergence can be reached in $\sim 25$ steps, with a starting redshift of $z=99$. Probing progressively smaller scales only requires an increase in the number of CDM particles being simulated, while the number of neutrino particles can remain fixed at a value less than or similar to the number of CDM particles. In turn, the percentage increase in runtime-per-step due to neutrino particles is between $\sim 5-20\%$ for runs with $1024^3$ CDM particles, and decreases as the number of CDM particles is increased. The code has been made publicly available, providing an invaluable resource to produce fast predictions for cosmological surveys and studying reconstruction.

astro-ph.CO↗

MADLens, a python package for fast and differentiable non-Gaussian lensing simulations

We present MADLens a python package for producing non-Gaussian lensing convergence maps at arbitrary source redshifts with unprecedented precision. MADLens is designed to achieve high accuracy while keeping computational costs as low as possible. A MADLens simulation with only $256^3$ particles produces convergence maps whose power agree with theoretical lensing power spectra up to $L{=}10000$ within the accuracy limits of HaloFit. This is made possible by a combination of a highly parallelizable particle-mesh algorithm, a sub-evolution scheme in the lensing projection, and a machine-learning inspired sharpening step. Further, MADLens is fully differentiable with respect to the initial conditions of the underlying particle-mesh simulations and a number of cosmological parameters. These properties allow MADLens to be used as a forward model in Bayesian inference algorithms that require optimization or derivative-aided sampling. Another use case for MADLens is the production of large, high resolution simulation sets as they are required for training novel deep-learning-based lensing analysis tools. We make the MADLens package publicly available under a Creative Commons License (https://github.com/VMBoehm/MADLens).

astro-ph.CO↗

Interval-driven discrete-time general nonlinear robust control: stabilization with closed-loop robust DOA enlargement

This paper presents new results that allow one to address the discrete-time general nonlinear robust control problem. The uncertain system is described by a general nonlinear function set characterized by the nominal model and the corresponding modeling error bound. Traditional synthesis methods design parameters of a structured robust controller. The key aim of this paper is to find an unstructured robust controller set in the state-control space, which enlarges the estimate of the closed-loop robust domain of attraction (RDOA). Based on the interval analysis arithmetic, a numerical method to estimate the unstructured robust controller set is proposed and the rigorous convergence analysis is given. The existing RDOA results are constrained by the level-set of the Lyapunov function, whereas the results in this paper remove this limitation. Furthermore, a solvable optimization problem is formulated so the estimate of RDOA is enlarged by selecting a Lyapunov function from a Lyapunov function set of sum-of-squares polynomials. The method is then validated by a specific case simulation study and results show more extensive RDOA than the previous methods.

math.OC↗

Stacking Redshifted 21cm Images of HII Regions Around High Redshift Galaxies as a Probe of Early Reionization

A number of current and future experiments aim to detect the reionization of neutral hydrogen by the first stars and galaxies in the Universe via the redshifted 21cm line. Using the \textsc{BlueTides} simulation, we investigate the measurement of an \textit{average} ionised region towards the beginning of reionization by stacking redshifted 21cm images around optically identified bright galaxies using mock observations. We find that with an SKA 1000 hour observation, assuming perfect foreground subtraction, a $5σ$ detection of a stacked HII region can be made with 30 images around some of the brightest galaxies in \textsc{bluetides} (brighter than $M_{UV} < -22.75$) at $z=9$ (corresponding to a neutral fraction of 90.1 \% in our model). We present simulated relationships between the UV magnitude of galaxies, the sizes of the ionised regions they reside in, and the shape of the stacked profiles. These mock observations can also distinguish between scenarios where the IGM is in net emission or absorption of 21cm photons. Once 21cm foreground contamination is included, we find that even with up to 200 images around these rare, bright galaxies, only a tentative $> 1σ$ detection will be possible. However, partial foreground subtraction substantially improves signal-to-noise. For example, we predict that reducing the area of Fourier space dominated by foregrounds by 50 (80) percent will allow $> 3σ$ ($> 5σ$) detections of ionised regions at $z=9$.

astro-ph.CO↗

On the possibility of Baryon Acoustic Oscillation measurements at redshift $z>7.6$ with the Roman Space Telescope

The Nancy Grace Roman Space Telescope (RST), with its field of view and high sensitivity will make surveys of cosmological large-scale structure possible at high redshifts. We investigate the possibility of detecting Baryon Acoustic Oscillations (BAO) at redshifts $z>7.6$ for use as a standard ruler. We use data from the hydrodynamic simulation \textsc{BlueTides} in conjunction with the gigaparsec-scale Outer Rim simulation and a model for patchy reionization to create mock RST High Latitude Survey grism data for Lyman-alpha emission line selected galaxies at redshifts $z=7.4$ to $z=10$, covering 2280 square degrees. We measure the monopoles of galaxies in the mock catalogues and fit the BAO features. We find that for a line flux of $L = 7\times 10^{-17} \ {\rm erg/s/cm}^{2}$, the $5 σ$ detection limit for the current design, the BAO feature is partially detectable (measured in three out of four survey quadrants analysed independently). The resulting root mean square error on the angular diameter distance to $z=7.7$ is 7.9$\%$. If we improve the detection sensitivity by a factor of two (i.e. $L = 3.5\times 10^{-17} \ {\rm erg/s/cm}^{2}$), the distance error reduces to $1.4\%$. We caution that many more factors are yet to be modelled, including dust obscuration, the damping wing due to the intergalactic medium, and low redshift interlopers. If these issues do not strongly affect the results, or different observational techniques (such as use of multiple lines) can mitigate them, RST or similar instruments may be able to constrain the angular diameter distance to the high redshift Universe.

astro-ph.CO↗