Searcharxiv⌕ Search

arXiv subjects

Xin Fu

Publications and source records attributed to Xin Fu.

At least 55 records · Page 3Linked to original sources

DEDGAT: Dual Embedding of Directed Graph Attention Networks for Detecting Financial Risk

Graph representation plays an important role in the field of financial risk control, where the relationship among users can be constructed in a graph manner. In practical scenarios, the relationships between nodes in risk control tasks are bidirectional, e.g., merchants having both revenue and expense behaviors. Graph neural networks designed for undirected graphs usually aggregate discriminative node or edge representations with an attention strategy, but cannot fully exploit the out-degree information when used for the tasks built on directed graph, which leads to the problem of a directional bias. To tackle this problem, we propose a Directed Graph ATtention network called DGAT, which explicitly takes out-degree into attention calculation. In addition to having directional requirements, the same node might have different representations of its input and output, and thus we further propose a dual embedding of DGAT, referred to as DEDGAT. Specifically, DEDGAT assigns in-degree and out-degree representations to each node and uses these two embeddings to calculate the attention weights of in-degree and out-degree nodes, respectively. Experiments performed on the benchmark datasets show that DGAT and DEDGAT obtain better classification performance compared to undirected GAT. Also,the visualization results demonstrate that our methods can fully use both in-degree and out-degree information.

cs.LG↗

Kahler-Einstein metric near an isolated log canonical singularity

We construct Kahler-Einstein metrics with negative scalar curvature near an isolated log canonical (non-log terminal) singularity. Such metrics are complete near the singularity if the underlying space has complex dimension 2 or if the singularity is smoothable. In complex dimension 2, we show that any complete Kahler-Einstein metric of negative scalar curvature near an isolated log canonical (non-log terminal) singularity is smoothly asymptotically close to one of the model metrics constructed by Kobayashi and Nakamura arising from hyperbolic geometry.

math.DG↗

Uniqueness of Tangent Cone of Kahler Einstein Metrics on Singular Varieties with Crepant Singularities

Let $(X, L)$ be a polarized Calabi Yau variety (or canonical polarized variety) with crepant singularity. Suppose $ω_{KE} \in c_1(L)$ (or $ω_{KE} \in c_1(K_X)$) is the unique Ricci flat current (or Kahler Einstein current with negative scalar curvature) with local bounded potential constructed in [18], we show that the local tangent at any point $p \in X$ of metric $ω_{KE}$ is unique

math.DG↗

Uniform convergence for linear elastostatic systems with periodic high contrast inclusions

We consider the Lame system of linear elasticity with periodically distributed inclusions whose elastic parameters have high contrast compared to the background media. We develop a unified method based on layer potential techniques to quantify three convergence results when some parameters of the elastic inclusions are sent to extreme values. More precisely, we study the incompressible inclusions limit where the bulk modulus of the inclusions tends to infinity, the soft inclusions limit where both the bulk modulus and the shear modulus tend to zero, and the hard inclusions limit where the shear modulus tends to infinity. Our method yields convergence rates that are independent of the periodicity of the inclusions array, and are sharper than some earlier results of this type. A key ingredient of the proof is the establishment of uniform spectra gaps for the elastic Neumann-Poincare operator associated to the collection of periodic inclusions that are independent of the periodicity.

math.AP↗

Dirichlet problem of complex Monge-Ampère equation near an isolated KLT singularity

We solve the Dirichlet Problem of Monge-Ampère equation near an isolate Klt singularity, which generalizes the result of Eyssidieux-Guedj-Zeriahi \cite{EGZ}, where the Monge-Ampère equation is solved on singular varieties without boundary. As a corollary, we construct solutions to Monge-Ampère equation with isolated singularity on strongly pseudoconvex domain $Ω$ contained in $\mathbb{C}^n$.

math.CV↗

Efficient Federated Learning for AIoT Applications Using Knowledge Distillation

As a promising distributed machine learning paradigm, Federated Learning (FL) trains a central model with decentralized data without compromising user privacy, which has made it widely used by Artificial Intelligence Internet of Things (AIoT) applications. However, the traditional FL suffers from model inaccuracy since it trains local models using hard labels of data and ignores useful information of incorrect predictions with small probabilities. Although various solutions try to tackle the bottleneck of the traditional FL, most of them introduce significant communication and memory overhead, making the deployment of large-scale AIoT devices a great challenge. To address the above problem, this paper presents a novel Distillation-based Federated Learning (DFL) architecture that enables efficient and accurate FL for AIoT applications. Inspired by Knowledge Distillation (KD) that can increase the model accuracy, our approach adds the soft targets used by KD to the FL model training, which occupies negligible network resources. The soft targets are generated by local sample predictions of each AIoT device after each round of local training and used for the next round of model training. During the local training of DFL, both soft targets and hard labels are used as approximation objectives of model predictions to improve model accuracy by supplementing the knowledge of soft targets. To further improve the performance of our DFL model, we design a dynamic adjustment strategy for tuning the ratio of two loss functions used in KD, which can maximize the use of both soft targets and hard labels. Comprehensive experimental results on well-known benchmarks show that our approach can significantly improve the model accuracy of FL with both Independent and Identically Distributed (IID) and non-IID data.

cs.LG↗

Just Noticeable Difference for Deep Machine Vision

As an important perceptual characteristic of the Human Visual System (HVS), the Just Noticeable Difference (JND) has been studied for decades with image and video processing (e.g., perceptual visual signal compression). However, there is little exploration on the existence of JND for the Deep Machine Vision (DMV), although the DMV has made great strides in many machine vision tasks. In this paper, we take an initial attempt, and demonstrate that the DMV has the JND, termed as the DMV-JND. We then propose a JND model for the image classification task in the DMV. It has been discovered that the DMV can tolerate distorted images with average PSNR of only 9.56dB (the lower the better), by generating JND via unsupervised learning with the proposed DMV-JND-NET. In particular, a semantic-guided redundancy assessment strategy is designed to restrain the magnitude and spatial distribution of the DMV-JND. Experimental results on image classification demonstrate that we successfully find the JND for deep machine vision. Our DMV-JND facilitates a possible direction for DMV-oriented image and video compression, watermarking, quality assessment, deep neural network security, and so on.

cs.CV↗

Generalized Transitional Markov Chain Monte Carlo Sampling Technique for Bayesian Inversion

In the context of Bayesian inversion for scientific and engineering modeling, Markov chain Monte Carlo sampling strategies are the benchmark due to their flexibility and robustness in dealing with arbitrary posterior probability density functions (PDFs). However, these algorithms been shown to be inefficient when sampling from posterior distributions that are high-dimensional or exhibit multi-modality and/or strong parameter correlations. In such contexts, the sequential Monte Carlo technique of transitional Markov chain Monte Carlo (TMCMC) provides a more efficient alternative. Despite the recent applicability for Bayesian updating and model selection across a variety of disciplines, TMCMC may require a prohibitive number of tempering stages when the prior PDF is significantly different from the target posterior. Furthermore, the need to start with an initial set of samples from the prior distribution may present a challenge when dealing with implicit priors, e.g. based on feasible regions. Finally, TMCMC can not be used for inverse problems with improper prior PDFs that represent lack of prior knowledge on all or a subset of parameters. In this investigation, a generalization of TMCMC that alleviates such challenges and limitations is proposed, resulting in a tempering sampling strategy of enhanced robustness and computational efficiency. Convergence analysis of the proposed sequential Monte Carlo algorithm is presented, proving that the distance between the intermediate distributions and the target posterior distribution monotonically decreases as the algorithm proceeds. The enhanced efficiency associated with the proposed generalization is highlighted through a series of test inverse problems and an engineering application in the oil and gas industry.

stat.CO↗

Asymptotics of Kähler-Einstein metrics on complex hyperbolic cusps

Let $L$ be a negative holomorphic line bundle over an $(n-1)$-dimensional complex torus $D$. Let $h$ be a Hermitian metric on $L$ such that the curvature form of the dual Hermitian metric defines a flat Kähler metric on $D$. Then $h$ is unique up to scaling, and, for some closed tubular neighborhood $V$ of the zero section $D \subset L$, the form $ω_h = -(n+1)i\partial\overline\partial\log(-{\log h})$ defines a complete Kähler-Einstein metric on $V \setminus D$ with ${\rm Ric}(ω_h) = -ω_h$. In fact, $ω_h$ is complex hyperbolic, i.e., the holomorphic sectional curvature of $ω_h$ is constant, and $ω_h$ has the usual doubly-warped cusp structure familiar from complex hyperbolic geometry. In this paper, we prove that if $U$ is another closed tubular neighborhood of the zero section and if $ω$ is a complete Kähler-Einstein metric with ${\rm Ric}(ω) = -ω$ on $U \setminus D$, then there exist a Hermitian metric $h$ as above and a $δ\in \mathbb{R}^+$ such that $ω- ω_{h} = O(e^{-δ\sqrt{-{\log h}}})$ to all orders with respect to $ω_h$ as $h \to 0$. This rate is doubly exponential in the distance from a fixed point, and is sharp.

math.DG↗

Shift-BNN: Highly-Efficient Probabilistic Bayesian Neural Network Training via Memory-Friendly Pattern Retrieving

Bayesian Neural Networks (BNNs) that possess a property of uncertainty estimation have been increasingly adopted in a wide range of safety-critical AI applications which demand reliable and robust decision making, e.g., self-driving, rescue robots, medical image diagnosis. The training procedure of a probabilistic BNN model involves training an ensemble of sampled DNN models, which induces orders of magnitude larger volume of data movement than training a single DNN model. In this paper, we reveal that the root cause for BNN training inefficiency originates from the massive off-chip data transfer by Gaussian Random Variables (GRVs). To tackle this challenge, we propose a novel design that eliminates all the off-chip data transfer by GRVs through the reversed shifting of Linear Feedback Shift Registers (LFSRs) without incurring any training accuracy loss. To efficiently support our LFSR reversion strategy at the hardware level, we explore the design space of the current DNN accelerators and identify the optimal computation mapping scheme to best accommodate our strategy. By leveraging this finding, we design and prototype the first highly efficient BNN training accelerator, named Shift-BNN, that is low-cost and scalable. Extensive evaluation on five representative BNN models demonstrates that Shift-BNN achieves an average of 4.9x (up to 10.8x) boost in energy efficiency and 1.6x (up to 2.8x) speedup over the baseline DNN training accelerator.

cs.AR↗

Codimension four regularity of generalized Einstein structures

We establish codimension 4 regularity of noncollapsed sequences of metrics with bounds on natural generalizations of the Ricci tensor. We obtain a priori L2 curvature estimates on such spaces, with diffeomorphism finiteness results and rigidity theorems as corollaries.

math.DG↗

A Large-Scale Benchmark for Food Image Segmentation

Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to two reasons: (1) there is a lack of high quality food image datasets with fine-grained ingredient labels and pixel-wise location masks -- the existing datasets either carry coarse ingredient labels or are small in size; and (2) the complex appearance of food makes it difficult to localize and recognize ingredients in food images, e.g., the ingredients may overlap one another in the same image, and the identical ingredient may appear distinctly in different food images. In this work, we build a new food image dataset FoodSeg103 (and its extension FoodSeg154) containing 9,490 images. We annotate these images with 154 ingredient classes and each image has an average of 6 ingredient labels and pixel-wise masks. In addition, we propose a multi-modality pre-training approach called ReLeM that explicitly equips a segmentation model with rich and semantic food knowledge. In experiments, we use three popular semantic segmentation methods (i.e., Dilated Convolution based, Feature Pyramid based, and Vision Transformer based) as baselines, and evaluate them as well as ReLeM on our new datasets. We believe that the FoodSeg103 (and its extension FoodSeg154) and the pre-trained models using ReLeM can serve as a benchmark to facilitate future works on fine-grained food image understanding. We make all these datasets and methods public at \url{https://xiongweiwu.github.io/foodseg103.html}.

cs.CV↗

Model category structures on multicomplexes

We present a family of model structures on the category of multicomplexes. There is a cofibrantly generated model structure in which the weak equivalences are the morphisms inducing an isomorphism at a fixed stage of an associated spectral sequence. Corresponding model structures are given for truncated versions of multicomplexes, interpolating between bicomplexes and multicomplexes. For a fixed stage of the spectral sequence, the model structures on all these categories are shown to be Quillen equivalent.

math.AT↗

The homotopy classification of four-dimensional toric orbifolds

Let $X$ be a $4$-dimensional toric orbifold. If $H^3(X)$ has a non-trivial odd primary torsion, then we show that $X$ is homotopy equivalent to the wedge of a Moore space and a CW-complex. As a corollary, given two 4-dimensional toric orbifolds having no 2-torsion in the cohomology, we prove that they have the same homotopy type if and only their integral cohomology rings are isomorphic.

math.AT↗

Parallel convolution processing using an integrated photonic tensor core

With the proliferation of ultra-high-speed mobile networks and internet-connected devices, along with the rise of artificial intelligence, the world is generating exponentially increasing amounts of data - data that needs to be processed in a fast, efficient and smart way. These developments are pushing the limits of existing computing paradigms, and highly parallelized, fast and scalable hardware concepts are becoming progressively more important. Here, we demonstrate a computational specific integrated photonic tensor core - the optical analog of an ASIC-capable of operating at Tera-Multiply-Accumulate per second (TMAC/s) speeds. The photonic core achieves parallelized photonic in-memory computing using phase-change memory arrays and photonic chip-based optical frequency combs (soliton microcombs). The computation is reduced to measuring the optical transmission of reconfigurable and non-resonant passive components and can operate at a bandwidth exceeding 14 GHz, limited only by the speed of the modulators and photodetectors. Given recent advances in hybrid integration of soliton microcombs at microwave line rates, ultra-low loss silicon nitride waveguides, and high speed on-chip detectors and modulators, our approach provides a path towards full CMOS wafer-scale integration of the photonic tensor core. While we focus on convolution processing, more generally our results indicate the major potential of integrated photonics for parallel, fast, and efficient computational hardware in demanding AI applications such as autonomous driving, live video processing, and next generation cloud computing services.

physics.optics↗

Ultrafast optical circuit switching for data centers using integrated soliton microcombs

Networks inside current data centers comprise a hierarchy of power-hungry electronic packet switches interconnected via optical fibers and transceivers. As the scaling of such electrically-switched networks approaches a plateau, a power-efficient solution is to implement a flat network with optical circuit switching (OCS), without electronic switches and a reduced number of transceivers due to direct links among servers. One of the promising ways of implementing OCS is by using tunable lasers and arrayed waveguide grating routers. Such an OCS-network can offer high bandwidth and low network latency, and the possibility of photonic integration results in an energy-efficient, compact, and scalable photonic data center network. To support dynamic data center workloads efficiently, it is critical to switch between wavelengths in sub nanoseconds (ns). Here we demonstrate ultrafast photonic circuit switching based on a microcomb. Using a photonic integrated Si3N4 microcomb in conjunction with semiconductor optical amplifiers (SOAs), sub ns (< 500 ps) switching of more than 20 carriers is achieved. Moreover, the 25-Gbps non-return to zero (NRZ) and 50-Gbps four-level pulse amplitude modulation (PAM-4) burst mode transmission systems are shown. Further, on-chip Indium phosphide (InP) based SOAs and arrayed waveguide grating (AWG) are used to show sub-ns switching along with 25-Gbps NRZ burst mode transmission providing a path toward a more scalable and energy-efficient wavelength-switched network for future data centers.

physics.app-ph↗

OO-VR: NUMA Friendly Object-Oriented VR Rendering Framework For Future NUMA-Based Multi-GPU Systems

With the strong computation capability, NUMA-based multi-GPU system is a promising candidate to provide sustainable and scalable performance for Virtual Reality. However, the entire multi-GPU system is viewed as a single GPU which ignores the data locality in VR rendering during the workload distribution, leading to tremendous remote memory accesses among GPU models. By conducting comprehensive characterizations on different kinds of parallel rendering frameworks, we observe that distributing the rendering object along with its required data per GPM can reduce the inter-GPM memory accesses. However, this object-level rendering still faces two major challenges in NUMA-based multi-GPU system: (1) the large data locality between the left and right views of the same object and the data sharing among different objects and (2) the unbalanced workloads induced by the software-level distribution and composition mechanisms. To tackle these challenges, we propose object-oriented VR rendering framework (OO-VR) that conducts the software and hardware co-optimization to provide a NUMA friendly solution for VR multi-view rendering in NUMA-based multi-GPU systems. We first propose an object-oriented VR programming model to exploit the data sharing between two views of the same object and group objects into batches based on their texture sharing levels. Then, we design an object aware runtime batch distribution engine and distributed hardware composition unit to achieve the balanced workloads among GPMs. Finally, evaluations on our VR featured simulator show that OO-VR provides 1.58x overall performance improvement and 76% inter-GPM memory traffic reduction over the state-of-the-art multi-GPU systems. In addition, OO-VR provides NUMA friendly performance scalability for the future larger multi-GPU scenarios with ever increasing asymmetric bandwidth between local and remote memory.

cs.DC↗