Searcharxiv⌕ Search

arXiv subjects

Yu Feng

Publications and source records attributed to Yu Feng.

At least 163 records · Page 9Linked to original sources

Crescent: Taming Memory Irregularities for Accelerating Deep Point Cloud Analytics

3D perception in point clouds is transforming the perception ability of future intelligent machines. Point cloud algorithms, however, are plagued by irregular memory accesses, leading to massive inefficiencies in the memory sub-system, which bottlenecks the overall efficiency. This paper proposes Crescent, an algorithm-hardware co-design system that tames the irregularities in deep point cloud analytics while achieving high accuracy. To that end, we introduce two approximation techniques, approximate neighbor search and selectively bank conflict elision, that "regularize" the DRAM and SRAM memory accesses. Doing so, however, necessarily introduces accuracy loss, which we mitigate by a new network training procedure that integrates approximation into the network training process. In essence, our training procedure trains models that are conditioned upon a specific approximate setting and, thus, retain a high accuracy. Experiments show that Crescent doubles the performance and halves the energy consumption compared to an optimized baseline accelerator with < 1% accuracy loss. The code of our paper is available at: https://github.com/horizon-research/crescent.

cs.AR↗

FIBA: Frequency-Injection based Backdoor Attack in Medical Image Analysis

In recent years, the security of AI systems has drawn increasing research attention, especially in the medical imaging realm. To develop a secure medical image analysis (MIA) system, it is a must to study possible backdoor attacks (BAs), which can embed hidden malicious behaviors into the system. However, designing a unified BA method that can be applied to various MIA systems is challenging due to the diversity of imaging modalities (e.g., X-Ray, CT, and MRI) and analysis tasks (e.g., classification, detection, and segmentation). Most existing BA methods are designed to attack natural image classification models, which apply spatial triggers to training images and inevitably corrupt the semantics of poisoned pixels, leading to the failures of attacking dense prediction models. To address this issue, we propose a novel Frequency-Injection based Backdoor Attack method (FIBA) that is capable of delivering attacks in various MIA tasks. Specifically, FIBA leverages a trigger function in the frequency domain that can inject the low-frequency information of a trigger image into the poisoned image by linearly combining the spectral amplitude of both images. Since it preserves the semantics of the poisoned image pixels, FIBA can perform attacks on both classification and dense prediction models. Experiments on three benchmarks in MIA (i.e., ISIC-2019 for skin lesion classification, KiTS-19 for kidney tumor segmentation, and EAD-2019 for endoscopic artifact detection), validate the effectiveness of FIBA and its superiority over state-of-the-art methods in attacking MIA models as well as bypassing backdoor defense. Source code will be available at https://github.com/HazardFY/FIBA.

cs.CV↗

Fast Parallel Hypertree Decompositions in Logarithmic Recursion Depth

Modern trends in data collection are bringing current mainstream techniques for database query processing to their limits. Consequently, various novel approaches for efficient query processing are being actively studied. One such approach is based on hypertree decompositions (HDs), which have been shown to carry great potential to process complex queries more efficiently and with stronger theoretical guarantees. However, using HDs for query execution relies on the difficult task of computing decompositions of the query structure, which guides the efficient execution of the query. From theoretical results we know that the performance of purely sequential methods is inherently limited, yet the problem is susceptible to parallelisation. In this paper we propose the first algorithm for computing hypertree decompositions that is well-suited for parallelisation. The proposed algorithm log-k-decomp requires only a logarithmic number of recursion levels and additionally allows for highly parallelised pruning of the search space by restriction to balanced separators. We provide detailed experimental evaluation over the HyperBench benchmark and demonstrate that our approach is highly effective especially for complex queries.

cs.DB↗

The ASTRID Simulation: Galaxy Formation and Reionization

We introduce the ASTRID simulation, a large-scale cosmological hydrodynamic simulation in a $250$ Mpc/h box with $2\times 5500^3$ particles. ASTRID contains a large number of high redshift galaxies, which can be compared to future survey data, and resolves galaxies in halos more massive than $2\times 10^9 M_\odot$. ASTRID has been run from $z=99$ to $z=3$. As a particular focus is modelling the high redshift Universe, it contains models for inhomogeneous hydrogen and helium reionization, baryon relative velocities and massive neutrinos, as well as supernova and AGN feedback. The black hole model includes mergers driven by dynamical friction rather than repositioning. We briefly summarise the implemented models, and the technical choices we took when developing the simulation code. We validate the model, showing good agreement with observed UV luminosity functions, galaxy stellar mass functions and specific star formation rates. We show that the redshift at which a given galaxy underwent hydrogen reionization has a large effect on the halo gas fraction. Finally, at $z=6$, halos with $M \sim 2\times 10^9 M_\odot$ which have been reionized have a star formation rate $1.5$ times greater than those which have not yet been reionized.

astro-ph.GA↗

The ASTRID simulation: the evolution of Supermassive Black Holes

We present the evolution of black holes (BHs) and their relationship with their host galaxies in Astrid, a large-volume cosmological hydrodynamical simulation with box size 250 $h^{-1} \rm Mpc$ containing $2\times5500^3$ particles evolved to z=3. Astrid statistically models BH gas accretion and AGN feedback to their environments, applies a power-law distribution for BH seed mass $M_{\rm sd}$, uses a dynamical friction model for BH dynamics and executes a physical treatment of BH mergers. The BH population is broadly consistent with empirical constraints on the BH mass function, the bright end of the luminosity functions, and the time evolution of BH mass and accretion rate density. The BH mass and accretion exhibit a tight correlation with host stellar mass and star formation rate. We trace BHs seeded before z>10 down to z=3, finding that BHs carry virtually no imprint of the initial $M_{\rm sd}$ except those with the smallest $M_{\rm sd}$, where less than 50\% of them have doubled in mass. Gas accretion is the dominant channel for BH growth compared to BH mergers. With dynamical friction, Astrid predicts a significant delay for BH mergers after the first encounter of a BH pair, with a typical elapse time of about 200 Myrs. There are in total $4.5 \times 10^5$ BH mergers in Astrid at z>3, $\sim 10^3$ of which have X-ray detectable EM counterparts: a bright kpc scale dual AGN with $L_X>10^{43}$ erg/s. BHs with $M_{\rm BH} \sim 10^{7-8} M_{\odot}$ experience the most frequent mergers. Galaxies that host BH mergers are unbiased tracers of the overall $M_{\rm BH} - M_{*}$ relation. Massive ($>10^{11} M_{\odot}$) galaxies have a high occupation number (>10) of BHs, and hence host the majority of BH mergers.

astro-ph.GA↗

Automated Transpilation of Imperative to Functional Code using Neural-Guided Program Synthesis (Extended Version)

While many mainstream languages such as Java, Python, and C# increasingly incorporate functional APIs to simplify programming and improve parallelization/performance, there are no effective techniques that can be used to automatically translate existing imperative code to functional variants using these APIs. Motivated by this problem, this paper presents a transpilation approach based on inductive program synthesis for modernizing existing code. Our method is based on the observation that the overwhelming majority of source/target programs in this setting satisfy an assumption that we call trace-compatibility: not only do the programs share syntactically identical low-level expressions, but these expressions also take the same values in corresponding execution traces. Our method leverages this observation to design a new neural-guided synthesis algorithm that (1) uses a novel neural architecture called cognate grammar network (CGN) and (2) leverages a form of concolic execution to prune partial programs based on intermediate values that arise during a computation. We have implemented our approach in a tool called NGST2 and use it to translate imperative Java and Python code to functional variants that use the Stream and functools APIs respectively. Our experiments show that NGST2 significantly outperforms several baselines and that our proposed neural architecture and pruning techniques are vital for achieving good results.

cs.PL↗

Predicting the Future Performance of the Planned Seismic Network in Mainland China

The new broadband seismic network in China will increase the number of stations from approximately 950 to 2000. The higher-resolution monitoring of the frequent smaller earthquakes expected inside Mainland China can be quantified via the completeness magnitude (Mc) metric. Using the Bayesian Magnitude of Completeness (BMC) method, we generate the spatial distribution of Mc predicted for the new network, based on the prior model calibrated on the current earthquake catalog (2012 to 2021) and network configuration. If 99% of Mainland China is at present covered down to Mc = 2.7, this threshold will soon fall to Mc = 2.0. This means approximately 3 times more earthquakes available per year. Based on the observation that seismic precursors are most likely to be observed at least at 3 units below the mainshock magnitude, the new seismic network shall achieve the goal of almost total coverage for optimal seismic-based earthquake prediction research.

physics.geo-ph↗

Storage capacity of networks with discrete synapses and sparsely encoded memories

Attractor neural networks (ANNs) are one of the leading theoretical frameworks for the formation and retrieval of memories in networks of biological neurons. In this framework, a pattern imposed by external inputs to the network is said to be learned when this pattern becomes a fixed point attractor of the network dynamics. The storage capacity is the maximum number of patterns that can be learned by the network. In this paper, we study the storage capacity of fully-connected and sparsely-connected networks with a binarized Hebbian rule, for arbitrary coding levels. Our results show that a network with discrete synapses has a similar storage capacity as the model with continuous synapses, and that this capacity tends asymptotically towards the optimal capacity, in the space of all possible binary connectivity matrices, in the sparse coding limit. We also derive finite coding level corrections for the asymptotic solution in the sparse coding limit. The result indicates the capacity of network with Hebbian learning rules converges to the optimal capacity extremely slowly when the coding level becomes small. Our results also show that in networks with sparse binary connectivity matrices, the information capacity per synapse is larger than in the fully connected case, and thus such networks store information more efficiently.

physics.bio-ph↗

The Impact of Dust on the Sizes of Galaxies in the Epoch of Reionization

We study the sizes of galaxies in the Epoch of Reionization using a sample of ~100,000 galaxies from the BlueTides cosmological hydrodynamical simulation from z=7 to 11. We measure the galaxy sizes from stellar mass and luminosity maps, defining the effective radius as the minimum radius which could enclose the pixels containing 50% of the total mass/light in the image. We find an inverse relationship between stellar mass and effective half-mass radius, suggesting that the most massive galaxies are more compact and dense than lower mass galaxies, which have flatter mass distributions. We find a mildly negative relation between intrinsic far-ultraviolet luminosity and size, while we find a positive size-luminosity relation when measured from dust-attenuated images. This suggests that dust is the predominant cause of the observed positive size-luminosity relation, with dust preferentially attenuating bright sight lines resulting in a flatter emission profile and thus larger measured effective radii. We study the size-luminosity relation across the rest-frame ultraviolet and optical, and find that the slope decreases at longer wavelengths; this is a consequence of the relation being caused by dust, which produces less attenuation at longer wavelengths. We find that the far-ultraviolet size-luminosity relation shows mild evolution from z=7 to 11, and galaxy size evolves with redshift as $R\propto(1+z)^{-m}$, where $m=0.662\pm0.009$. Finally, we investigate the sizes of z=7 quasar host galaxies, and find that while the intrinsic sizes of quasar hosts are small relative to the overall galaxy sample, they have comparable sizes when measured from dust-attenuated images.

astro-ph.GA↗

Real-Time Gaze Tracking with Event-Driven Eye Segmentation

Gaze tracking is increasingly becoming an essential component in Augmented and Virtual Reality. Modern gaze tracking al gorithms are heavyweight; they operate at most 5 Hz on mobile processors despite that near-eye cameras comfortably operate at a r eal-time rate ($>$ 30 Hz). This paper presents a real-time eye tracking algorithm that, on average, operates at 30 Hz on a mobile processor, achieves \ang{0.1}--\ang{0.5} gaze accuracies, all the while requiring only 30K parameters, one to two orders of magn itude smaller than state-of-the-art eye tracking algorithms. The crux of our algorithm is an Auto~ROI mode, which continuously pr edicts the Regions of Interest (ROIs) of near-eye images and judiciously processes only the ROIs for gaze estimation. To that end, we introduce a novel, lightweight ROI prediction algorithm by emulating an event camera. We discuss how a software emulation of events enables accurate ROI prediction without requiring special hardware. The code of our paper is available at https://github.com/horizon-research/edgaze.

cs.HC↗

Determination of building flood risk maps from LiDAR mobile mapping data

With increasing urbanization, flooding is a major challenge for many cities today. Based on forecast precipitation, topography, and pipe networks, flood simulations can provide early warnings for areas and buildings at risk of flooding. Basement windows, doors, and underground garage entrances are common places where floodwater can flow into a building. Some buildings have been prepared or designed considering the threat of flooding, but others have not. Therefore, knowing the heights of these facade openings helps to identify places that are more susceptible to water ingress. However, such data is not yet readily available in most cities. Traditional surveying of the desired targets may be used, but this is a very time-consuming and laborious process. This research presents a new process for the extraction of windows and doors from LiDAR mobile mapping data. Deep learning object detection models are trained to identify these objects. Usually, this requires to provide large amounts of manual annotations. In this paper, we mitigate this problem by leveraging a rule-based method. In a first step, the rule-based method is used to generate pseudo-labels. A semi-supervised learning strategy is then applied with three different levels of supervision. The results show that using only automatically generated pseudo-labels, the learning-based model outperforms the rule-based approach by 14.6% in terms of F1-score. After five hours of human supervision, it is possible to improve the model by another 6.2%. By comparing the detected facade openings' heights with the predicted water levels from a flood simulation model, a map can be produced which assigns per-building flood risk levels. This information can be combined with flood forecasting to provide a more targeted disaster prevention guide for the city's infrastructure and residential buildings.

cs.CV↗

IoTDataBench: Extending TPCx-IoT for Compression and Scalability

We present a record-breaking result and lessons learned in practicing TPCx-IoT benchmarking for a real-world use case. We find that more system characteristics need to be benchmarked for its application to real-world use cases. We introduce an extension to the TPCx-IoT benchmark, covering fundamental requirements of time-series data management for IoT infrastructure. We characterize them as data compression and system scalability. To evaluate these two important features of IoT databases, we propose IoTDataBench and update four aspects of TPCx-IoT, i.e., data generation, workloads, metrics and test procedures. Preliminary evaluation results show systems that fail to effectively compress data or flexibly scale can negatively affect the redesigned metrics, while systems with high compression ratios and linear scalability are rewarded in the final metrics. Such systems have the ability to scale up computing resources on demand and can thus save dollar costs.

cs.DB↗

Neutron Scattering Studies of the Breathing Pyrochlore Antiferromagnet LiGaCr$_{4}$O$_{8}$

We report neutron scattering measurements of the spinel oxide LiGaCr$_{4}$O$_{8}$, in which magnetic ions Cr$^{3+}$ form a breathing pyrochlore lattice. Our experiments reveal the coexistence of a nearly dispersionless resonance mode and dispersive spin wave excitations in the magnetically ordered state, which can be quantitatively described by a quantum spin model of hexagonal loops and linear spin wave theory with the same set of exchange parameters, respectively. Comparison to other Cr spinel oxides reveals a linear relationship between the resonance energy and lattice constant across all these materials, which is in agreement with our hexagonal loop calculations. Our results suggest a unified picture for spin resonances in Cr spinel oxides.

cond-mat.str-el↗

Massive Black Hole Mergers with Orbital Information: Predictions from the ASTRID Simulation

We examine massive black hole (MBH) mergers and their associated gravitational wave signals from the large-volume cosmological simulation Astrid. Astrid includes galaxy formation and black hole models recently updated with a MBH seed population between $3\times 10^4M_{\odot}/h$ and $3\times 10^5M_{\odot}/h$ and a sub-grid dynamical friction (DF) model to follow the MBH dynamics down to $1.5\;\text{ckpc}/h$. We calculate initial eccentricities of MBH orbits directly from the simulation at kpc-scales, and find orbital eccentricities above $0.7$ for most MBH pairs before the numerical merger. After approximating unresolved evolution on scales below ${\sim 200\,\text{pc}}$, we find that the in-simulation DF on large scales accounts for more than half of the total orbital decay time ($\sim 500\,\text{Myrs}$) due to DF. The binary hardening time is an order of magnitude longer than the DF time, especially for the seed-mass binaries ($M_\text{BH}<2M_\text{seed}$). As a result, only $\lesssim20\%$ of seed MBH pairs merge at $z>3$ after considering both unresolved DF evolution and binary hardening. These $z>3$ seed-mass mergers are hosted in a biased population of galaxies with the highest stellar masses of $>10^9\,M_\odot$. With the higher initial eccentricity prediction from Astrid, we estimate an expected merger rate of $0.3-0.7$ per year from the $z>3$ MBH population. This is a factor of $\sim 7$ higher than the prediction using the circular orbit assumption. The LISA events are expected at a similar rate, and comprise $\gtrsim 60\%$ seed-seed mergers, $\sim 30\%$ involving only one seed-mass MBH, and $\sim 10\%$ mergers of non-seed MBHs.

astro-ph.GA↗

The DESI $N$-body Simulation Project I: Testing the Robustness of Simulations for the DESI Dark Time Survey

Analysis of large galaxy surveys requires confidence in the robustness of numerical simulation methods. The simulations are used to construct mock galaxy catalogs to validate data analysis pipelines and identify potential systematics. We compare three $N$-body simulation codes, ABACUS, GADGET, and SWIFT, to investigate the regimes in which their results agree. We run $N$-body simulations at three different mass resolutions, $6.25\times10^{8}$, $2.11\times10^{9}$, and $5.00\times10^{9}~h^{-1}$M$_{\odot}$, matching phases to reduce the noise within the comparisons. We find systematic errors in the halo clustering between different codes are smaller than the DESI statistical error for $s > 20\, h^{-1}$Mpc in the correlation function in redshift space. Through the resolution comparison we find that simulations run with a mass resolution of $2.1\times10^{9}~h^{-1}$M$_{\odot}$ are sufficiently converged for systematic effects in the halo clustering to be smaller than the DESI statistical error at scales larger than $20 \, h^{-1}$Mpc. These findings show that the simulations are robust for extracting cosmological information from large scales which is the key goal of the DESI survey. Comparing matter power spectra, we find the codes agree to within 1% for $k \leq 10~h$Mpc$^{-1}$. We also run a comparison of three initial condition generation codes and find good agreement. In addition, we include a quasi-$N$-body code, FastPM, since we plan use it for certain DESI analyses. The impact of the halo definition and galaxy-halo relation will be presented in a follow up study.

astro-ph.CO↗

SAILFISH: Vetting Smart Contract State-Inconsistency Bugs in Seconds

This paper presents SAILFISH, a scalable system for automatically finding state-inconsistency bugs in smart contracts. To make the analysis tractable, we introduce a hybrid approach that includes (i) a light-weight exploration phase that dramatically reduces the number of instructions to analyze, and (ii) a precise refinement phase based on symbolic evaluation guided by our novel value-summary analysis, which generates extra constraints to over-approximate the side effects of whole-program execution, thereby ensuring the precision of the symbolic evaluation. We developed a prototype of SAILFISH and evaluated its ability to detect two state-inconsistency flaws, viz., reentrancy and transaction order dependence (TOD) in Ethereum smart contracts. Further, we present detection rules for other kinds of smart contract flaws that SAILFISH can be extended to detect. Our experiments demonstrate the efficiency of our hybrid approach as well as the benefit of the value summary analysis. In particular, we show that S SAILFISH outperforms five state-of-the-art smart contract analyzers (SECURITY, MYTHRIL, OYENTE, SEREUM and VANDAL ) in terms of performance, and precision. In total, SAILFISH discovered 47 previously unknown vulnerable smart contracts out of 89,853 smart contracts from ETHERSCAN .

cs.CR↗

Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent

Distributed Deep Learning (DDL) is essential for large-scale Deep Learning (DL) training. Synchronous Stochastic Gradient Descent (SSGD) 1 is the de facto DDL optimization method. Using a sufficiently large batch size is critical to achieving DDL runtime speedup. In a large batch setting, the learning rate must be increased to compensate for the reduced number of parameter updates. However, a large learning rate may harm convergence in SSGD and training could easily diverge. Recently, Decentralized Parallel SGD (DPSGD) has been proposed to improve distributed training speed. In this paper, we find that DPSGD not only has a system-wise run-time benefit but also a significant convergence benefit over SSGD in the large batch setting. Based on a detailed analysis of the DPSGD learning dynamics, we find that DPSGD introduces additional landscape-dependent noise that automatically adjusts the effective learning rate to improve convergence. In addition, we theoretically show that this noise smoothes the loss landscape, hence allowing a larger learning rate. We conduct extensive studies over 18 state-of-the-art DL models/tasks and demonstrate that DPSGD often converges in cases where SSGD diverges for large learning rates in the large batch setting. Our findings are consistent across two different application domains: Computer Vision (CIFAR10 and ImageNet-1K) and Automatic Speech Recognition (SWB300 and SWB2000), and two different types of neural network models: Convolutional Neural Networks and Long Short-Term Memory Recurrent Neural Networks.

cs.LG↗

Not all peaks are created equal: the early growth of Supermassive Black Holes

In this work, we use the constrained Gaussian realization technique to study the early growth of supermassive black holes (SMBHs) in cosmological hydrodynamic simulations, exploring its relationship with features of the initial density peaks on large scales, ~1 Mpc/h. Our constrained simulations of volume (20 Mpc/h)^3 successfully reconstruct the large-scale structure as well as the black hole growth for the hosts of the rare 10^9 Msun SMBHs found in the BlueTides simulation at z~7. We run a set of simulations with constrained initial conditions by imposing a 5 σ_0(R_G) peak on scale of R_G = 1 Mpc/h varying different peak features, such as the shape and compactness as well as the tidal field surrounding the peak. We find that initial density peaks with high compactness and low tidal field induce the most rapid BH growth at early epochs. This is because compact density peaks with a more spherical large scale matter distribution lead to the formation of high density gas clumps in the centers of halos, and thus boost early BH accretion. Moreover, such initially compact density peaks in low tidal field regions also lead to a more compact BH host galaxy morphology. This can explain the tight correlation between BH growth and host galaxy compactness seen in observations.

astro-ph.GA↗