Searcharxiv⌕ Search

arXiv subjects

Yu Feng

Publications and source records attributed to Yu Feng.

At least 127 records · Page 7Linked to original sources

Gap Completion in Point Cloud Scene occluded by Vehicles using SGC-Net

Recent advances in mobile mapping systems have greatly enhanced the efficiency and convenience of acquiring urban 3D data. These systems utilize LiDAR sensors mounted on vehicles to capture vast cityscapes. However, a significant challenge arises due to occlusions caused by roadside parked vehicles, leading to the loss of scene information, particularly on the roads, sidewalks, curbs, and the lower sections of buildings. In this study, we present a novel approach that leverages deep neural networks to learn a model capable of filling gaps in urban scenes that are obscured by vehicle occlusion. We have developed an innovative technique where we place virtual vehicle models along road boundaries in the gap-free scene and utilize a ray-casting algorithm to create a new scene with occluded gaps. This allows us to generate diverse and realistic urban point cloud scenes with and without vehicle occlusion, surpassing the limitations of real-world training data collection and annotation. Furthermore, we introduce the Scene Gap Completion Network (SGC-Net), an end-to-end model that can generate well-defined shape boundaries and smooth surfaces within occluded gaps. The experiment results reveal that 97.66% of the filled points fall within a range of 5 centimeters relative to the high-density ground truth point cloud scene. These findings underscore the efficacy of our proposed model in gap completion and reconstructing urban scenes affected by vehicle occlusions.

cs.CV↗

BLINK: Multimodal Large Language Models Can See but Not Perceive

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" (e.g., relative depth estimation, visual correspondence, forensics detection, and multi-view reasoning). However, we find these perception-demanding tasks cast significant challenges for current multimodal LLMs because they resist mediation through natural language. Blink reformats 14 classic computer vision tasks into 3,807 multiple-choice questions, paired with single or multiple images and visual prompting. While humans get 95.70% accuracy on average, Blink is surprisingly challenging for existing multimodal LLMs: even the best-performing GPT-4V and Gemini achieve accuracies of 51.26% and 45.72%, only 13.17% and 7.63% higher than random guessing, indicating that such perception abilities have not "emerged" yet in recent multimodal LLMs. Our analysis also highlights that specialist CV models could solve these problems much better, suggesting potential pathways for future improvements. We believe Blink will stimulate the community to help multimodal LLMs catch up with human-level visual perception.

cs.CV↗

Connection Probabilities of Multiple FK-Ising Interfaces

We find the scaling limits of a general class of boundary-to-boundary connection probabilities and multiple interfaces in the critical planar FK-Ising model, thus verifying predictions from the physics literature. We also discuss conjectural formulas using Coulomb gas integrals for the corresponding quantities in general critical planar random-cluster models with cluster-weight $q \in [1,4)$. Thus far, proofs for convergence, including ours, rely on discrete complex analysis techniques and are beyond reach for other values of $q$ than the FK-Ising model ($q=2$). Given the convergence of interfaces, the conjectural formulas for other values of $q$ could be verified similarly with relatively minor technical work. The limit interfaces are variants of $\mathrm{SLE}_κ$ curves (with $κ= 16/3$ for $q=2$). Their partition functions, that give the connection probabilities, also satisfy properties predicted for correlation functions in conformal field theory (CFT), expected to describe scaling limits of critical random-cluster models. We verify these properties for all $q \in [1,4)$, thus providing further evidence of the expected CFT description of these models.

math.PR↗

PNG-UNITsims: Halo clustering response to primordial non-Gaussianities as a function of mass

We present the largest full N-body simulation to date with local primordial non-Gaussianities (L-PNG), the \textsc{PNG-UNITsim}. It tracks the evolution of $4096^3$ particles within a periodic box with $L_{\rm box} = 1 \; h^{-1}\,{\rm Gpc}$, leading to a mass resolution of $m_{p} = 1.24\times 10^{9}\; h^{-1}\,M_\odot$. This is enough to resolve galaxies targeted by stage-IV spectroscopic surveys. The \textsc{PNG-UNIT} has \textit{Fixed} initial conditions whose phases are also \textit{Matched} to the pre-existing \textsc{UNIT} simulation. These two features in the simulations reduce our uncertainty significantly so we use 100 \textsc{FastPM} mocks to estimate this reduction. The amplitude of the non-Gaussianities used to set the initial conditions of this new simulation is $f_{\rm NL}^{\rm local} = 100$. In this first study, we use mass selected dark matter haloes from the \textsc{PNG-UNIT} simulation to constrain the local PNG parameters. PNG induce a scale dependent bias, parameterised through \bp or $p$, which might depend on the type of cosmological tracer. Those cases when $p=1$ are referred to as the {\it universality relation}. We measure $p$ as a function of the halo mass. Haloes with masses between $1\times 10^{12}$ and $2\times 10^{13} \, h^{-1} M_\odot$ are well described by the {\it universality relation}. For haloes with masses between $2\times 10^{10}$ and $1\times 10^{12} \, h^{-1} M_\odot$ we find that $p<1$ at $3σ$. Combining all the mass bins, we find $p$ consistent with a value of $0.955\pm0.013$, which is $3σ$ away from \textit{universality}, as low mass haloes are more numerous. We also study the effect of using priors on $p$ when constraining $f_{\rm NL}$. Using the values we obtain for $b_ϕ$ as priors, we forecast that a DESI-like (stage-IV) survey will be able to constrain $f_{\rm NL}$ better than if the universality relation is assumed.

astro-ph.CO↗

Estimating Earthquake Early Warning Effectiveness via Blind Zone Sizes: A Case Study of the Planned Seismic Network in Chinese Mainland

The China Earthquake Administration (CEA) has launched an ambitious nationwide earthquake early warning (EEW) system project currently under development, which will include approximately 15,000 seismic stations and be the largest EEW system in the world. The new EEW system is planned to go online at the end of 2023. In approximately 50%, 30% and 20% of Chinese mainland, the inter-station distance will soon be smaller than 50 km, 25 km and 15 km, respectively. The expected effectiveness of this EEW system can be quantified via the metric determined from the radius of the blind zone, which refers to the area near the epicenter where there is insufficient time to issue a warning before the arrival of strong S- and surface waves. This study uses a theoretical network-based method together with Monte Carlo simulation to obtain the spatial distribution of the blind zone radii and their associated uncertainties for the new seismic network based on its configuration. We find that the densified new seismic network is expected to have excellent EEW performance as the area covered by small blind zones with radius less than 30 km increases dramatically from approximately 2% to 22%, that is by 2.4 million km2 inside Chinese mainland. We also find that every 1,000,000 RMB (about 146,000 USD) invested to densify the planned network will lead to an areal increase of 3,000 km2 of small blind zones. Continuing to increase the density of stations in some key regions with blind zone radii ranging from 15 to 40 km is still necessary to control the unexpected expansion of blind zones due to possible (and common) stations failure. Our work provides insights into the expected performance of the upcoming EEW network in Chinese mainland, and our proposed evaluation approach is broadly applicable for predicting the performance of EEW systems during their planning, design, and implementation stages.

physics.geo-ph↗

BlissCam: Boosting Eye Tracking Efficiency with Learned In-Sensor Sparse Sampling

Eye tracking is becoming an increasingly important task domain in emerging computing platforms such as Augmented/Virtual Reality (AR/VR). Today's eye tracking system suffers from long end-to-end tracking latency and can easily eat up half of the power budget of a mobile VR device. Most existing optimization efforts exclusively focus on the computation pipeline by optimizing the algorithm and/or designing dedicated accelerators while largely ignoring the front-end of any eye tracking pipeline: the image sensor. This paper makes a case for co-designing the imaging system with the computing system. In particular, we propose the notion of "in-sensor sparse sampling", whereby the pixels are drastically downsampled (by 20x) within the sensor. Such in-sensor sampling enhances the overall tracking efficiency by significantly reducing 1) the power consumption of the sensor readout chain and sensor-host communication interfaces, two major power contributors, and 2) the work done on the host, which receives and operates on far fewer pixels. With careful reuse of existing pixel circuitry, our proposed BLISSCAM requires little hardware augmentation to support the in-sensor operations. Our synthesis results show up to 8.2x energy reduction and 1.4x latency reduction over existing eye tracking pipelines.

cs.AR↗

Test-Time Training on Graphs with Large Language Models (LLMs)

Graph Neural Networks have demonstrated great success in various fields of multimedia. However, the distribution shift between the training and test data challenges the effectiveness of GNNs. To mitigate this challenge, Test-Time Training (TTT) has been proposed as a promising approach. Traditional TTT methods require a demanding unsupervised training strategy to capture the information from test to benefit the main task. Inspired by the great annotation ability of Large Language Models (LLMs) on Text-Attributed Graphs (TAGs), we propose to enhance the test-time training on graphs with LLMs as annotators. In this paper, we design a novel Test-Time Training pipeline, LLMTTT, which conducts the test-time adaptation under the annotations by LLMs on a carefully-selected node set. Specifically, LLMTTT introduces a hybrid active node selection strategy that considers not only node diversity and representativeness, but also prediction signals from the pre-trained model. Given annotations from LLMs, a two-stage training strategy is designed to tailor the test-time model with the limited and noisy labels. A theoretical analysis ensures the validity of our method and extensive experiments demonstrate that the proposed LLMTTT can achieve a significant performance improvement compared to existing Out-of-Distribution (OOD) generalization methods.

cs.LG↗

Cicero: Addressing Algorithmic and Architectural Bottlenecks in Neural Rendering by Radiance Warping and Memory Optimizations

Neural Radiance Field (NeRF) is widely seen as an alternative to traditional physically-based rendering. However, NeRF has not yet seen its adoption in resource-limited mobile systems such as Virtual and Augmented Reality (VR/AR), because it is simply extremely slow. On a mobile Volta GPU, even the state-of-the-art NeRF models generally execute only at 0.8 FPS. We show that the main performance bottlenecks are both algorithmic and architectural. We introduce, CICERO, to tame both forms of inefficiencies. We first introduce two algorithms, one fundamentally reduces the amount of work any NeRF model has to execute, and the other eliminates irregular DRAM accesses. We then describe an on-chip data layout strategy that eliminates SRAM bank conflicts. A pure software implementation of CICERO offers an 8.0x speed-up and 7.9x energy saving over a mobile Volta GPU. When compared to a baseline with a dedicated DNN accelerator, our speed-up and energy reduction increase to 28.2x and 37.8x, respectively - all with minimal quality loss (less than 1.0 dB peak signal-to-noise ratio reduction).

cs.AR↗

Multiple Ising interfaces in annulus and $2N$-sided radial SLE

We consider critical planar Ising model in annulus with alternating boundary conditions on the outer boundary and free boundary conditions in the inner boundary. As the size of the inner hole goes to zero, the event that all interfaces get close to the inner hole before they meet each other is a rare event. We prove that the law of the collection of the interfaces conditional on this rare event converges in total variation distance to the so-called $2N$-sided radial SLE$_3$, introduced by~[HL21]. The proof relies crucially on an estimate for multiple chordal SLE. Suppose $(γ_1, \ldots, γ_N)$ is chordal $N$-SLE$_κ$ with $κ\in (0,4]$ in the unit disc, and we consider the probability that all $N$ curves get close to the origin. We prove that the limit $\lim_{r\to 0+}r^{-A_{2N}}\mathbb{P}[\mathrm{dist}(0,γ_j)<r, 1\le j\le N]$ exists, where $A_{2N}$ is the so-called $2N$-arm exponents and $\mathrm{dist}$ is Euclidean distance. We call the limit Green's function for chordal $N$-SLE$_κ$. This estimate is a generalization of previous conclusions with $N=1$ and $N=2$ proved in~[LR12, LR15] and~[Zha20] respectively.

math.PR↗

DiffPoint: Single and Multi-view Point Cloud Reconstruction with ViT Based Diffusion Model

As the task of 2D-to-3D reconstruction has gained significant attention in various real-world scenarios, it becomes crucial to be able to generate high-quality point clouds. Despite the recent success of deep learning models in generating point clouds, there are still challenges in producing high-fidelity results due to the disparities between images and point clouds. While vision transformers (ViT) and diffusion models have shown promise in various vision tasks, their benefits for reconstructing point clouds from images have not been demonstrated yet. In this paper, we first propose a neat and powerful architecture called DiffPoint that combines ViT and diffusion models for the task of point cloud reconstruction. At each diffusion step, we divide the noisy point clouds into irregular patches. Then, using a standard ViT backbone that treats all inputs as tokens (including time information, image embeddings, and noisy patches), we train our model to predict target points based on input images. We evaluate DiffPoint on both single-view and multi-view reconstruction tasks and achieve state-of-the-art results. Additionally, we introduce a unified and flexible feature fusion module for aggregating image features from single or multiple input images. Furthermore, our work demonstrates the feasibility of applying unified architectures across languages and images to improve 3D reconstruction tasks.

cs.CV↗

A unified Bayesian inversion approach for a class of tumor growth models with different pressure laws

In this paper, we use the Bayesian inversion approach to study the data assimilation problem for a family of tumor growth models described by porous-medium type equations. The models contain uncertain parameters and are indexed by a physical parameter $m$, which characterizes the constitutive relation between density and pressure. Based on these models, we employ the Bayesian inversion framework to infer parametric and nonparametric unknowns that affect tumor growth from noisy observations of tumor cell density. We establish the well-posedness and the stability theories for the Bayesian inversion problem and further prove the convergence of the posterior distribution in the so-called incompressible limit, $m \rightarrow \infty$. Since the posterior distribution across the index regime $m\in[2,\infty)$ can thus be treated in a unified manner, such theoretical results also guide the design of the numerical inference for the unknown. We propose a generic computational framework for such inverse problems, which consists of a typical sampling algorithm and an asymptotic preserving solver for the forward problem. With extensive numerical tests, we demonstrate that the proposed method achieves satisfactory accuracy in the Bayesian inference of the tumor growth models, which is uniform with respect to the constitutive relation.

math.NA↗

Differentiable Cosmological Simulation with Adjoint Method

Rapid advances in deep learning have brought not only myriad powerful neural networks, but also breakthroughs that benefit established scientific research. In particular, automatic differentiation (AD) tools and computational accelerators like GPUs have facilitated forward modeling of the Universe with differentiable simulations. Based on analytic or automatic backpropagation, current differentiable cosmological simulations are limited by memory, and thus are subject to a trade-off between time and space/mass resolution, usually sacrificing both. We present a new approach free of such constraints, using the adjoint method and reverse time integration. It enables larger and more accurate forward modeling at the field level, and will improve gradient based optimization and inference. We implement it in an open-source particle-mesh (PM) $N$-body library pmwd (particle-mesh with derivatives). Based on the powerful AD system JAX, pmwd is fully differentiable, and is highly performant on GPUs.

astro-ph.IM↗

JUNO: Optimizing High-Dimensional Approximate Nearest Neighbour Search with Sparsity-Aware Algorithm and Ray-Tracing Core Mapping

Approximate nearest neighbor (ANN) search is a widely applied technique in modern intelligent applications, such as recommendation systems and vector databases. Therefore, efficient and high-throughput execution of ANN search has become increasingly important. In this paper, we first characterize the state-of-the-art product quantization-based method of ANN search and identify a significant source of inefficiency in the form of unnecessary pairwise distance calculations and accumulations. To improve efficiency, we propose JUNO, an end-to-end ANN search system that adopts a carefully designed sparsity- and locality-aware search algorithm. We also present an efficient hardware mapping that utilizes ray tracing cores in modern GPUs with pipelined execution on tensor cores to execute our sparsity-aware ANN search algorithm. Our evaluations on four datasets ranging in size from 1 to 100 million search points demonstrate 2.2x-8.5x improvements in search throughput. Moreover, our algorithmic enhancements alone achieve a maximal 2.6x improvement on the hardware without the acceleration of the RT core.

cs.DC↗

Integrating 3D City Data through Knowledge Graphs

CityGML is a widely adopted standard by the Open Geospatial Consortium (OGC) for representing and exchanging 3D city models. The representation of semantic and topological properties in CityGML makes it possible to query such 3D city data to perform analysis in various applications, e.g., security management and emergency response, energy consumption and estimation, and occupancy measurement. However, the potential of querying CityGML data has not been fully exploited. The official GML/XML encoding of CityGML is only intended as an exchange format but is not suitable for query answering. The most common way of dealing with CityGML data is to store them in the 3DCityDB system as relational tables and then query them with the standard SQL query language. Nevertheless, for end users, it remains a challenging task to formulate queries over 3DCityDB directly for their ad-hoc analytical tasks, because there is a gap between the conceptual semantics of CityGML and the relational schema adopted in 3DCityDB. In fact, the semantics of CityGML itself can be modeled as a suitable ontology. The technology of Knowledge Graphs (KGs), where an ontology is at the core, is a good solution to bridge such a gap. Moreover, embracing KGs makes it easier to integrate with other spatial data sources, e.g., OpenStreetMap and existing (Geo)KGs (e.g., Wikidata, DBPedia, and GeoNames), and to perform queries combining information from multiple data sources. In this work, we describe a CityGML KG framework to populate the concepts in the CityGML ontology using declarative mappings to 3DCityDB, thus exposing the CityGML data therein as a KG. To demonstrate the feasibility of our approach, we use CityGML data from the city of Munich as test data and integrate OpenStreeMap data in the same area.

cs.DB↗

Cerberus: Query-driven Scalable Vulnerability Detection in OAuth Service Provider Implementations

OAuth protocols have been widely adopted to simplify user authentication and service authorization for third-party applications. However, little effort has been devoted to automatically checking the security of the libraries that service providers widely use. In this paper, we formalize the OAuth specifications and security best practices, and design Cerberus, an automated static analyzer, to find logical flaws and identify vulnerabilities in the implementation of OAuth service provider libraries. To efficiently detect security violations in a large codebase of service provider implementation, Cerberus employs a query-driven algorithm for answering queries about OAuth specifications. We demonstrate the effectiveness of Cerberus by evaluating it on datasets of popular OAuth libraries with millions of downloads. Among these high-profile libraries, Cerberus has identified 47 vulnerabilities from ten classes of logical flaws, 24 of which were previously unknown. We got acknowledged by the developers of eight libraries and had three accepted CVEs.

cs.CR↗

Super-sample covariance of the power spectrum, bispectrum, halos, voids, and their cross covariances

We study the effect of super-sample covariance (SSC) on the power spectrum and higher-order statistics: bispectrum, halo mass function, and void size function. We also investigate the effect of SSC on the cross covariance between the statistics. We consider both the matter and halo fields. Higher-order statistics of the large-scale structure contain additional cosmological information beyond the power spectrum and are a powerful tool to constrain cosmology. They are a promising probe for ongoing and upcoming high precision cosmological surveys such as DESI, PFS, Rubin Observatory LSST, Euclid, SPHEREx, SKA, and Roman Space Telescope. Cosmological simulations used in modeling and validating these statistics often have sizes that are much smaller than the observed Universe. Density fluctuations on scales larger than the simulation box, known as super-sample modes, are not captured by the simulations and in turn can lead to inaccuracies in the covariance matrix. We compare the covariance measured using simulation boxes containing super-sample modes to those without. We also compare with the Separate Universe approach. We find that while the power spectrum, bispectrum and halo mass function show significant scale- or mass-dependent SSC, the void size function shows relatively small SSC. We also find significant SSC contributions to the cross covariances between the different statistics, implying that future joint-analyses will need to carefully take into consideration the effect of SSC. To enable further study of SSC, our simulations have been made publicly available at https://github.com/HalfDomeSims/ssc.

astro-ph.CO↗

Efficient Reionization in a Large Hydrodynamic Galaxy Formation Simulation

Accuracy in the topology and statistics of a simulated Epoch of Reionization (EoR) are vital to draw connections between observations and physical processes. While full radiative transfer models produce the most accurate reionization models, they are highly computationally expensive, and are infeasible for the largest cosmological simulations. Instead, large simulations often include EoR models that are pre-computed via the initial density field, or post-processed where feedback effects are ignored. We introduce Astrid-ES, a resimulation of the Astrid epoch of reionisation $20 > z > 5.5$ which includes an on-the-fly excursion-set reionization algorithm. Astrid-ES produces more accurate reionization histories without significantly impacting the computational time. This model directly utilises the star particles produced in the simulation to calculate the EoR history and includes a UV background which heats the gas particles after their reionization. We contrast the reionization topology and statistics in Astrid-ES with the previously employed parametric reionisation model, finding that in Astrid-ES, ionised regions are more correlated with galaxies, and the 21cm power-spectrum shows an increase in large scale power. We calculate the relation between the size of HII regions and the UV luminosity of the brightest galaxy within them. Prior to the overlap phase, we find a power-law fit of $\mathrm{log} (R) = -0.314 M_\mathrm{UV} - 2.550 \mathrm{log}(1+z) + 7.408$ with a standard deviation $σ_R < 0.15 \mathrm{dex}$ across all mass bins. We also examine the properties of halos throughout reionization, finding that while the properties of halos in the simulation are correlated with the redshift of reionisation, they are not greatly affected by reionisation itself.

astro-ph.CO↗

The activity-weight duality in feed forward neural networks: The geometric determinants of generalization

One of the fundamental problems in machine learning is generalization. In neural network models with a large number of weights (parameters), many solutions can be found to fit the training data equally well. The key question is which solution can describe testing data not in the training set. Here, we report the discovery of an exact duality (equivalence) between changes in activities in a given layer of neurons and changes in weights that connect to the next layer of neurons in a densely connected layer in any feed forward neural network. The activity-weight (A-W) duality allows us to map variations in inputs (data) to variations of the corresponding dual weights. By using this mapping, we show that the generalization loss can be decomposed into a sum of contributions from different eigen-directions of the Hessian matrix of the loss function at the solution in weight space. The contribution from a given eigen-direction is the product of two geometric factors (determinants): the sharpness of the loss landscape and the standard deviation of the dual weights, which is found to scale with the weight norm of the solution. Our results provide an unified framework, which we used to reveal how different regularization schemes (weight decay, stochastic gradient descent with different batch sizes and learning rates, dropout), training data size, and labeling noise affect generalization performance by controlling either one or both of these two geometric determinants for generalization. These insights can be used to guide development of algorithms for finding more generalizable solutions in overparametrized neural networks.

cs.LG↗