Searcharxiv⌕ Search

arXiv subjects

Jun Cheng

Publications and source records attributed to Jun Cheng.

At least 127 records · Page 7Linked to original sources

Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification

Gait-based person re-identification (Re-ID) is valuable for safety-critical applications, and using only 3D skeleton data to extract discriminative gait features for person Re-ID is an emerging open topic. Existing methods either adopt hand-crafted features or learn gait features by traditional supervised learning paradigms. Unlike previous methods, we for the first time propose a generic gait encoding approach that can utilize unlabeled skeleton data to learn gait representations in a self-supervised manner. Specifically, we first propose to introduce self-supervision by learning to reconstruct input skeleton sequences in reverse order, which facilitates learning richer high-level semantics and better gait representations. Second, inspired by the fact that motion's continuity endows temporally adjacent skeletons with higher correlations ("locality"), we propose a locality-aware attention mechanism that encourages learning larger attention weights for temporally adjacent skeletons when reconstructing current skeleton, so as to learn locality when encoding gait. Finally, we propose Attention-based Gait Encodings (AGEs), which are built using context vectors learned by locality-aware attention, as final gait representations. AGEs are directly utilized to realize effective person Re-ID. Our approach typically improves existing skeleton-based methods by 10-20% Rank-1 accuracy, and it achieves comparable or even superior performance to multi-modal methods with extra RGB or depth information. Our codes are available at https://github.com/Kali-Hac/SGE-LA.

cs.CV↗

Solutions to the mean king's problem: higher-dimensional quantum error-correcting codes

Mean king's problem is a kind of quantum state discrimination problems. In the problem, we try to discriminate eigenstates of noncommutative observables with the help of classical delayed information. The problem has been investigated from the viewpoint of error detection and correction. We construct higher-dimensional quantum error-correcting codes against error corresponding to the noncommutative observables. Any code state of the codes provides a way to discriminate the eigenstates correctly with the classical delayed information.

quant-ph↗

Encoding Structure-Texture Relation with P-Net for Anomaly Detection in Retinal Images

Anomaly detection in retinal image refers to the identification of abnormality caused by various retinal diseases/lesions, by only leveraging normal images in training phase. Normal images from healthy subjects often have regular structures (e.g., the structured blood vessels in the fundus image, or structured anatomy in optical coherence tomography image). On the contrary, the diseases and lesions often destroy these structures. Motivated by this, we propose to leverage the relation between the image texture and structure to design a deep neural network for anomaly detection. Specifically, we first extract the structure of the retinal images, then we combine both the structure features and the last layer features extracted from original health image to reconstruct the original input healthy image. The image feature provides the texture information and guarantees the uniqueness of the image recovered from the structure. In the end, we further utilize the reconstructed image to extract the structure and measure the difference between structure extracted from original and the reconstructed image. On the one hand, minimizing the reconstruction difference behaves like a regularizer to guarantee that the image is corrected reconstructed. On the other hand, such structure difference can also be used as a metric for normality measurement. The whole network is termed as P-Net because it has a ``P'' shape. Extensive experiments on RESC dataset and iSee dataset validate the effectiveness of our approach for anomaly detection in retinal images. Further, our method also generalizes well to novel class discovery in retinal images and anomaly detection in real-world images.

eess.IV↗

Unsupervised Deformable Medical Image Registration via Pyramidal Residual Deformation Fields Estimation

Deformation field estimation is an important and challenging issue in many medical image registration applications. In recent years, deep learning technique has become a promising approach for simplifying registration problems, and has been gradually applied to medical image registration. However, most existing deep learning registrations do not consider the problem that when the receptive field cannot cover the corresponding features in the moving image and the fixed image, it cannot output accurate displacement values. In fact, due to the limitation of the receptive field, the 3 x 3 kernel has difficulty in covering the corresponding features at high/original resolution. Multi-resolution and multi-convolution techniques can improve but fail to avoid this problem. In this study, we constructed pyramidal feature sets on moving and fixed images and used the warped moving and fixed features to estimate their "residual" deformation field at each scale, called the Pyramidal Residual Deformation Field Estimation module (PRDFE-Module). The "total" deformation field at each scale was computed by upsampling and weighted summing all the "residual" deformation fields at all its previous scales, which can effectively and accurately transfer the deformation fields from low resolution to high resolution and is used for warping the moving features at each scale. Simulation and real brain data results show that our method improves the accuracy of the registration and the rationality of the deformation field.

cs.CV↗

Theoretical study of kinetics of proton coupled electron transfer in photocatalysis

Photocatalysis induced by sunlight is one of the most promising approach to environmental protection, solar energy conversion and sustainable production of fuels. The computational modeling of photocatalysis is a rapidly expending field which requires to adapt and further develop the available theoretical tools. The coupled transfer of proton and electron is an important reaction during photocatalysis. In this work, we present the first step of our methodology development in which we apply existing kinetic theory of such coupled transfer to a model system, namely, methanol photo-dissociation on rutile TiO$_2$(110) surface, with the help of high-level first-principles calculations. Moreover, we adapt the Stuchebrukhov-Hammes-Schiffer kinetic theory, where we use the Georgievskii-Stuchebrukhova vibronic coupling, to calculate the rate constant of the proton coupled electron transfer reaction for a particular pathway. In particular, we propose a modified expression to calculate the rate constant which enforces the near-resonance condition for the vibrational wavefunction during proton tunneling.

cond-mat.mtrl-sci↗

Sparse-GAN: Sparsity-constrained Generative Adversarial Network for Anomaly Detection in Retinal OCT Image

With the development of convolutional neural network, deep learning has shown its success for retinal disease detection from optical coherence tomography (OCT) images. However, deep learning often relies on large scale labelled data for training, which is oftentimes challenging especially for disease with low occurrence. Moreover, a deep learning system trained from data-set with one or a few diseases is unable to detect other unseen diseases, which limits the practical usage of the system in disease screening. To address the limitation, we propose a novel anomaly detection framework termed Sparsity-constrained Generative Adversarial Network (Sparse-GAN) for disease screening where only healthy data are available in the training set. The contributions of Sparse-GAN are two-folds: 1) The proposed Sparse-GAN predicts the anomalies in latent space rather than image-level; 2) Sparse-GAN is constrained by a novel Sparsity Regularization Net. Furthermore, in light of the role of lesions for disease screening, we present to leverage on an anomaly activation map to show the heatmap of lesions. We evaluate our proposed Sparse-GAN on a publicly available dataset, and the results show that the proposed method outperforms the state-of-the-art methods.

cs.CV↗

Universal digital filtering for denoising volumetric retinal OCT and OCT angiography in 3D shearlet domain

Retinal optical coherence tomography (OCT) and OCT angiography (OCTA) suffer from the degeneration of image quality due to speckle noise and bulk-motion noise, respectively. Because the cross-sectional retina has distinct features in OCT and OCTA B-scans, existing digital filters that can denoise OCT efficiently are unable to handle the bulk-motion noise in OCTA. In this Letter, we propose a universal digital filtering approach that is capable of minimizing both types of noise. Considering the retinal capillaries in OCTA are hard to differentiate in B-scans while having distinct curvilinear structures in 3D volumes, we decompose the volumetric OCT and OCTA data with 3D shearlets thus efficiently separate the retinal tissue and vessels from the noise in this transform domain. Compared with wavelets and curvelets, the shearlets provide better representation of the layer edges in OCT and the vasculature in OCTA. Qualitative and quantitative results show the proposed method outperforms the state-of-the-art OCT and OCTA denoising methods. Besides, the superiority of 3D denoising is demonstrated by comparing the 3D shearlet filtering with its 2D counterpart.

eess.IV↗

Digital resolution enhancement in low transverse sampling optical coherence tomography angiography using deep learning

Optical coherence tomography angiography (OCTA) requires high transverse sampling density for visualizing retinal and choroidal capillaries. Low transverse sampling causes resolution degradation, such as the angiograms in wide-field OCTA. In this paper, we propose to address this problem using deep learning. We conducted extensive experiments on converting the centrally cropped 3 x 3 mm2 field of view (FOV) of the 8 x 8 mm2 foveal OCTA images (a sampling density of 22.9 $μ$m) to the native 3 x 3 mm2 en face OCTA images (a sampling density of 12.2 $μ$m). We employed a cycle-consistent adversarial network architecture in this conversion. The quantitative analysis using the perceptual similarity measures shows the generated OCTA images are closer to the native 3 x 3 mm2 scans. Besides, the results show the proposed method could also enhance signal-to-noise ratio. We further applied our method to enhance diseased cases and calculate vascular biomarkers, which demonstrates its generalization performance and clinical perspective.

eess.IV↗

BioNet: Infusing Biomarker Prior into Global-to-Local Network for Choroid Segmentation in Optical Coherence Tomography Images

Choroid is the vascular layer of the eye, which is directly related to the incidence and severity of many ocular diseases. Optical Coherence Tomography (OCT) is capable of imaging both the cross-sectional view of retina and choroid, but the segmentation of the choroid region is challenging because of the fuzzy choroid-sclera interface (CSI). In this paper, we propose a biomarker infused global-to-local network (BioNet) for choroid segmentation, which segments the choroid with higher credibility and robustness. Firstly, our method trains a biomarker prediction network to learn the features of the biomarker. Then a global multi-layers segmentation module is applied to segment the OCT image into 12 layers. Finally, the global multi-layered result and the original OCT image are fed into a local choroid segmentation module to segment the choroid region with the biomarker infused as regularizer. We conducted comparison experiments with the state-of-the-art methods on a dataset (named AROD). The experimental results demonstrate the superiority of our method with $90.77\%$ Dice-index and 6.23 pixels Average-unsigned-surface-detection-error, etc.

cs.CV↗

Dense Dilated Network with Probability Regularized Walk for Vessel Detection

The detection of retinal vessel is of great importance in the diagnosis and treatment of many ocular diseases. Many methods have been proposed for vessel detection. However, most of the algorithms neglect the connectivity of the vessels, which plays an important role in the diagnosis. In this paper, we propose a novel method for retinal vessel detection. The proposed method includes a dense dilated network to get an initial detection of the vessels and a probability regularized walk algorithm to address the fracture issue in the initial detection. The dense dilated network integrates newly proposed dense dilated feature extraction blocks into an encoder-decoder structure to extract and accumulate features at different scales. A multiscale Dice loss function is adopted to train the network. To improve the connectivity of the segmented vessels, we also introduce a probability regularized walk algorithm to connect the broken vessels. The proposed method has been applied on three public data sets: DRIVE, STARE and CHASE_DB1. The results show that the proposed method outperforms the state-of-the-art methods in accuracy, sensitivity, specificity and also are under receiver operating characteristic curve.

eess.IV↗

Synergistic Effect of One- and Two-dimensional Connected Coral-like Li$_{6.25}$Al$_{0.25}$La$_3$Zr$_2$O$_{12}$ in PEO-Based Composite Solid State Electrolyte

As one of the most promising next-generation energy storage device, lithium metal batteries have been extensively investigated. However, the poor safety issue and undesired lithium dendrites growth hinder the development of lithium metal batteries. The application of solid state electrolytes has attracted increasing attention as they can solve the safety issue and partly inhibit the growth of lithium dendrites. Polyethylene oxide (PEO)-based electrolytes are very promising due to their enhanced safety and excellent flexibility. However, PEO-based electrolytes suffer from low ionic conductivity at room temperature and can't effectively inhibit lithium dendrites at high temperature due to the intrinsic semi-crystalline properties and poor mechanical strength. In this work, a novel coral-like Li6.25Al0.25La3Zr2O12 (LALZO) is synthesized to use as an active ceramic filler in PEO. The PEO with LALZO coral (PLC) exhibits increased ionic conductivity and mechanical strength, which leads to the uniform deposition/stripping of lithium metal. The Li symmetric cells with PLC cycle for 1500 h without short circuit at 50 centidegree. The assembled LiFePO4/PLC/Li batteries display excellent cycling stability at both 60 and 50 centidegree. This work reveals that the electrochemical properties of organic and inorganic composite electrolyte can be effectively improved by tuning the microstructure of the filler, such as the coral-like LALZO architecture.

physics.chem-ph↗

The Channel Attention based Context Encoder Network for Inner Limiting Membrane Detection

The optic disc segmentation is an important step for retinal image-based disease diagnosis such as glaucoma. The inner limiting membrane (ILM) is the first boundary in the OCT, which can help to extract the retinal pigment epithelium (RPE) through gradient edge information to locate the boundary of the optic disc. Thus, the ILM layer segmentation is of great importance for optic disc localization. In this paper, we build a new optic disc centered dataset from 20 volunteers and manually annotated the ILM boundary in each OCT scan as ground-truth. We also propose a channel attention based context encoder network modified from the CE-Net to segment the optic disc. It mainly contains three phases: the encoder module, the channel attention based context encoder module, and the decoder module. Finally, we demonstrate that our proposed method achieves state-of-the-art disc segmentation performance on our dataset mentioned above.

eess.IV↗

SkrGAN: Sketching-rendering Unconditional Generative Adversarial Networks for Medical Image Synthesis

Generative Adversarial Networks (GANs) have the capability of synthesizing images, which have been successfully applied to medical image synthesis tasks. However, most of existing methods merely consider the global contextual information and ignore the fine foreground structures, e.g., vessel, skeleton, which may contain diagnostic indicators for medical image analysis. Inspired by human painting procedure, which is composed of stroking and color rendering steps, we propose a Sketching-rendering Unconditional Generative Adversarial Network (SkrGAN) to introduce a sketch prior constraint to guide the medical image generation. In our SkrGAN, a sketch guidance module is utilized to generate a high quality structural sketch from random noise, then a color render mapping is used to embed the sketch-based representations and resemble the background appearances. Experimental results show that the proposed SkrGAN achieves the state-of-the-art results in synthesizing images for various image modalities, including retinal color fundus, X-Ray, Computed Tomography (CT) and Magnetic Resonance Imaging (MRI). In addition, we also show that the performances of medical image segmentation method have been improved by using our synthesized images as data augmentation.

cs.CV↗

Mapping Dynamical Magnetic Responses of Ultra-thin Micron-size Superconducting Films using Nitrogen-vacancy Centers in Diamond

Two-dimensional superconductors have attracted growing interest because of their scientific novelty, structural tunability, and useful properties. Studies of their magnetic responses, however, are often hampered by difficulties to grow large-size samples of high quality and uniformity. We report here an imaging method that employed NV- centers in diamond as sensor capable of mapping out the microwave magnetic field distribution on an ultrathin superconducting film of micron size. Measurements on a 33nm-thick film and a 125nm-thick bulk-like film of $Bi_2Sr_2CaCu_2O_{8+δ}$ revealed that the ac Meissner effect (or repulsion of ac magnetic field) set in at 78K and 91K, respectively; the latter was the superconducting transition temperature (Tc) of both films. The unusual ac magnetic response of the thin film presumably was due to thermally excited vortex-antivortex diffusive motion in the film. Spatial resolution of our ac magnetometer was limited by optical diffraction and the noise level was at 14 $μT/Hz^{1/2}$. The technique could be extended with better detection sensitivity to extract local ac conductivity/susceptibility of ultrathin or monolayer superconducting samples as well as ac magnetic responses of other two-dimensional exotic thin films of limited lateral size.

cond-mat.supr-con↗

CE-Net: Context Encoder Network for 2D Medical Image Segmentation

Medical image segmentation is an important step in medical image analysis. With the rapid development of convolutional neural network in image processing, deep learning has been used for medical image segmentation, such as optic disc segmentation, blood vessel detection, lung segmentation, cell segmentation, etc. Previously, U-net based approaches have been proposed. However, the consecutive pooling and strided convolutional operations lead to the loss of some spatial information. In this paper, we propose a context encoder network (referred to as CE-Net) to capture more high-level information and preserve spatial information for 2D medical image segmentation. CE-Net mainly contains three major components: a feature encoder module, a context extractor and a feature decoder module. We use pretrained ResNet block as the fixed feature extractor. The context extractor module is formed by a newly proposed dense atrous convolution (DAC) block and residual multi-kernel pooling (RMP) block. We applied the proposed CE-Net to different 2D medical image segmentation tasks. Comprehensive results show that the proposed method outperforms the original U-Net method and other state-of-the-art methods for optic disc segmentation, vessel detection, lung segmentation, cell contour segmentation and retinal optical coherence tomography layer segmentation.

cs.CV↗

A construction of UD $k$-ary multi-user codes from $(2^m(k-1)+1)$-ary codes for MAAC

In this paper, we proposed a construction of a UD $k$-ary $T$-user coding scheme for MAAC. We first give a construction of $k$-ary $T^{f+g}$-user UD code from a $k$-ary $T^{f}$-user UD code and a $k^{\pm}$-ary $T^{g}$-user difference set with its two component sets $\mathcal{D}^{+}$ and $\mathcal{D}^{-}$ {\em a priori}. Based on the $k^{\pm}$-ary $T^{g}$-user difference set constructed from a $(2k-1)$-ary UD code, we recursively construct a UD $k$-ary $T$-user codes with code length of $2^m$ from initial multi-user codes of $k$-ary, $2(k-1)+1$-ary, \dots, $(2^m(k-1)+1)$-ary. Introducing multi-user codes with higer-ary makes the total rate of generated code $\mathcal{A}$ higher than that of conventional code.

cs.IT↗

Multi-Cell Multi-Task Convolutional Neural Networks for Diabetic Retinopathy Grading

Diabetic Retinopathy (DR) is a non-negligible eye disease among patients with Diabetes Mellitus, and automatic retinal image analysis algorithm for the DR screening is in high demand. Considering the resolution of retinal image is very high, where small pathological tissues can be detected only with large resolution image and large local receptive field are required to identify those late stage disease, but directly training a neural network with very deep architecture and high resolution image is both time computational expensive and difficult because of gradient vanishing/exploding problem, we propose a \textbf{Multi-Cell} architecture which gradually increases the depth of deep neural network and the resolution of input image, which both boosts the training time but also improves the classification accuracy. Further, considering the different stages of DR actually progress gradually, which means the labels of different stages are related. To considering the relationships of images with different stages, we propose a \textbf{Multi-Task} learning strategy which predicts the label with both classification and regression. Experimental results on the Kaggle dataset show that our method achieves a Kappa of 0.841 on test set which is the 4-th rank of all state-of-the-arts methods. Further, our Multi-Cell Multi-Task Convolutional Neural Networks (M$^2$CNN) solution is a general framework, which can be readily integrated with many other deep neural network architectures.

cs.CV↗

Towards Query Efficient Black-box Attacks: An Input-free Perspective

Recent studies have highlighted that deep neural networks (DNNs) are vulnerable to adversarial attacks, even in a black-box scenario. However, most of the existing black-box attack algorithms need to make a huge amount of queries to perform attacks, which is not practical in the real world. We note one of the main reasons for the massive queries is that the adversarial example is required to be visually similar to the original image, but in many cases, how adversarial examples look like does not matter much. It inspires us to introduce a new attack called \emph{input-free} attack, under which an adversary can choose an arbitrary image to start with and is allowed to add perceptible perturbations on it. Following this approach, we propose two techniques to significantly reduce the query complexity. First, we initialize an adversarial example with a gray color image on which every pixel has roughly the same importance for the target model. Then we shrink the dimension of the attack space by perturbing a small region and tiling it to cover the input image. To make our algorithm more effective, we stabilize a projected gradient ascent algorithm with momentum, and also propose a heuristic approach for region size selection. Through extensive experiments, we show that with only 1,701 queries on average, we can perturb a gray image to any target class of ImageNet with a 100\% success rate on InceptionV3. Besides, our algorithm has successfully defeated two real-world systems, the Clarifai food detection API and the Baidu Animal Identification API.

stat.ML↗