Searcharxiv⌕ Search

arXiv subjects

Feng Cheng

Publications and source records attributed to Feng Cheng.

49 records · Page 3Linked to original sources

Boosting the Capability of Intelligent Vulnerability Detection by Training in a Human-Learning Manner

Due to its powerful automatic feature extraction, deep learning (DL) has been widely used in source code vulnerability detection. However, although it performs well on artificial datasets, its performance is not satisfactory when detecting real-world vulnerabilities due to the high complexity of real-world samples. In this paper, we propose to train DL-based vulnerability detection models in a human-learning manner, that is, start with the simplest samples and then gradually transition to difficult knowledge. Specifically, we design a novel framework (Humer) that can enhance the detection ability of DL-based vulnerability detectors. To validate the effectiveness of Humer, we select five state-of-the-art DL-based vulnerability detection models (TokenCNN, VulDeePecker, StatementGRU, ASTGRU, and Devign) to complete our evaluations. Through the results, we find that the use of Humer can increase the F1 of these models by an average of 10.5%. Moreover, Humer can make the model detect up to 16.7% more real-world vulnerabilities. Meanwhile, we also conduct a case study to uncover vulnerabilities from real-world open source products by using these enhanced DL-based vulnerability detectors. Through the results, we finally discover 281 unreported vulnerabilities in NVD, of which 98 have been silently patched by vendors in the latest version of corresponding products, but 159 still exist in the products.

cs.CR↗

Using Single-Trial Representational Similarity Analysis with EEG to track semantic similarity in emotional word processing

Electroencephalography (EEG) is a powerful non-invasive brain imaging technique with a high temporal resolution that has seen extensive use across multiple areas of cognitive science research. This thesis adapts representational similarity analysis (RSA) to single-trial EEG datasets and introduces its principles to EEG researchers unfamiliar with multivariate analyses. We have two separate aims: 1. we want to explore the effectiveness of single-trial RSA on EEG datasets; 2. we want to utilize single-trial RSA and computational semantic models to investigate the role of semantic meaning in emotional word processing. We report two primary findings: 1. single-trial RSA on EEG datasets can produce meaningful and interpretable results given a high number of trials and subjects; 2. single-trial RSA reveals that emotional processing in the 500-800ms time window is associated with additional semantic analysis.

q-bio.NC↗

Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query Workload

The use of deep learning models for forecasting the resource consumption patterns of SQL queries have recently been a popular area of study. With many companies using cloud platforms to power their data lakes for large scale analytic demands, these models form a critical part of the pipeline in managing cloud resource provisioning. While these models have demonstrated promising accuracy, training them over large scale industry workloads are expensive. Space inefficiencies of encoding techniques over large numbers of queries and excessive padding used to enforce shape consistency across diverse query plans implies 1) longer model training time and 2) the need for expensive, scaled up infrastructure to support batched training. In turn, we developed Prestroid, a tree convolution based data science pipeline that accurately predicts resource consumption patterns of query traces, but at a much lower cost. We evaluated our pipeline over 19K Presto OLAP queries from Grab, on a data lake of more than 20PB of data. Experimental results imply that our pipeline outperforms benchmarks on predictive accuracy, contributing to more precise resource prediction for large-scale workloads, yet also reduces per-batch memory footprint by 13.5x and per-epoch training time by 3.45x. We demonstrate direct cost savings of up to 13.2x for large batched model training over Microsoft Azure VMs.

cs.LG↗

Learning Directional Feature Maps for Cardiac MRI Segmentation

Cardiac MRI segmentation plays a crucial role in clinical diagnosis for evaluating personalized cardiac performance parameters. Due to the indistinct boundaries and heterogeneous intensity distributions in the cardiac MRI, most existing methods still suffer from two aspects of challenges: inter-class indistinction and intra-class inconsistency. To tackle these two problems, we propose a novel method to exploit the directional feature maps, which can simultaneously strengthen the differences between classes and the similarities within classes. Specifically, we perform cardiac segmentation and learn a direction field pointing away from the nearest cardiac tissue boundary to each pixel via a direction field (DF) module. Based on the learned direction field, we then propose a feature rectification and fusion (FRF) module to improve the original segmentation features, and obtain the final segmentation. The proposed modules are simple yet effective and can be flexibly added to any existing segmentation network without excessively increasing time and space complexity. We evaluate the proposed method on the 2017 MICCAI Automated Cardiac Diagnosis Challenge (ACDC) dataset and a large-scale self-collected dataset, showing good segmentation performance and robust generalization ability of the proposed method.

cs.CV↗

On the radius of spatial analyticity for the inviscid Boussinesq equations

In this paper, we study the problem of analyticity of smooth solutions of the inviscid Boussinesq equations. If the initial datum is real-analytic, the solution remains real-analytic on the existence interval. By an inductive method we can obtain lower bounds on the radius of spatial analyticity of the smooth solution.

math.AP↗

Probabilistic representation and inverse design of metamaterials based on a deep generative model with semi-supervised learning strategy

The research of metamaterials has achieved enormous success in the manipulation of light in an artificially prescribed manner using delicately designed sub-wavelength structures, so-called meta-atoms. Even though modern numerical methods allow to accurately calculate the optical response of complex structures, the inverse design of metamaterials is still a challenging task due to the non-intuitive and non-unique relationship between physical structures and optical responses. To better unveil this implicit relationship and thus facilitate metamaterial design, we propose to represent metamaterials and model the inverse design problem in a probabilistically generative manner. By employing an encoder-decoder configuration, our deep generative model compresses the meta-atom design and optical response into a latent space, where similar designs and similar optical responses are automatically clustered together. Therefore, by sampling in the latent space, the stochastic latent variables function as codes, from which the candidate designs are generated upon given requirements in a decoding process. With the effective latent representation of metamaterials, we can elegantly model the complex structure-performance relationship in an interpretable way, and solve the one-to-many mapping issue that is intractable in a deterministic model. Moreover, to alleviate the burden of numerical calculation in data collection, we develop a semi-supervised learning strategy that allows our model to utilize unlabeled data in addition to labeled data during training, simultaneously optimizing the generative inverse design and deterministic forward prediction in an end-to-end manner. On a data-driven basis, the proposed model can serve as a comprehensive and efficient tool that accelerates the design, characterization and even new discovery in the research domain of metamaterials and photonics in general.

physics.optics↗

Faster and Non-ergodic O(1/K) Stochastic Alternating Direction Method of Multipliers

We study stochastic convex optimization subjected to linear equality constraints. Traditional Stochastic Alternating Direction Method of Multipliers and its Nesterov's acceleration scheme can only achieve ergodic O(1/\sqrt{K}) convergence rates, where K is the number of iteration. By introducing Variance Reduction (VR) techniques, the convergence rates improve to ergodic O(1/K). In this paper, we propose a new stochastic ADMM which elaborately integrates Nesterov's extrapolation and VR techniques. We prove that our algorithm can achieve a non-ergodic O(1/K) convergence rate which is optimal for separable linearly constrained non-smooth convex problems, while the convergence rates of VR based ADMM methods are actually tight O(1/\sqrt{K}) in non-ergodic sense. To the best of our knowledge, this is the first work that achieves a truly accelerated, stochastic convergence rate for constrained convex problems. The experimental results demonstrate that our algorithm is significantly faster than the existing state-of-the-art stochastic ADMM methods.

math.OC↗

Weighted gevrey class regularity of euler equation in the whole space

In this paper we study the weighted Gevrey class regularity of Euler equation in the whole space R 3. We first establish the local existence of Euler equation in weighted Sobolev space, then obtain the weighted Gevrey regularity of Euler equation. We will use the weighted Sobolev-Gevrey space method to obtain the results of Gevrey regularity of Euler equation, and the use of the property of singular operator in the estimate of the pressure term is the improvement of our work.

math.AP↗

On the gevrey regularity of solutions to the 3d ideal mhd equations

In this paper, similar to the incompressible Euler equation, we prove the propagation of the Gevrey regularity of solutions to the three-dimensional incompressible ideal magnetohydrodynamics (MHD) equations. We also obtain an uniform estimate of Gevery radius for the solution of MHD equation.

math.AP↗

Vanishing viscosity limit of navier-stokes equations in gevrey class

In this paper we consider the inviscid limit for the periodic solutions to Navier-Stokes equation in the framework of Gevrey class. It is shown that the lifespan for the solutions to Navier-Stokes equation is independent of viscosity, and that the solutions of the Navier-Stokes equation converge to that of Euler equation in Gevrey class as the viscosity tends to zero. Moreover the convergence rate in Gevrey class is presented.

math.AP↗

Gevrey regularity with weight for incompressible Euler equation in the half plane

In this work we prove the weighted Gevrey regularity of solutions to the incompressible Euler equation with initial data decaying polynomially at infinity. This is motivated by the well-posedness problem of vertical boundary layer equation for fast rotating fluid. The method presented here is based on the basic weighted $L^2$- estimate, and the main difficulty arises from the estimate on the pressure term due to the appearance of weight function.

math.AP↗

Digital Logarithmic Airborne Gamma Ray Spectrometer

A new digital logarithmic airborne gamma ray spectrometer is designed in this study. The spectrometer adopts a high-speed and high-accuracy logarithmic amplifier (LOG114) to amplify the pulse signal logarithmically and to improve the utilization of the ADC dynamic range, because the low-energy pulse signal has a larger gain than the high-energy pulse signal. The spectrometer can clearly distinguish the photopeaks at 239, 352, 583, and 609keV in the low-energy spectral sections after the energy calibration. The photopeak energy resolution of 137Cs improves to 6.75% from the original 7.8%. Furthermore, the energy resolution of three photopeaks, namely, K, U, and Th, is maintained, and the overall stability of the energy spectrum is increased through potassium peak spectrum stabilization. Thus, effectively measuring energy from 20keV to 10MeV is possible.

physics.ins-det↗

Bose-Einstein condensation on an atom chip

We report an experiment of creating Bose-Einstein condensate (BEC) on an atom chip. The chip based Z-wire current and a homogeneous bias magnetic field create a tight magnetic trap, which allows for a fast production of BEC. After an 4.17s forced radio frequency evaporative cooling, a condensate with about 3000 atoms appears. And the transition temperature is about 300nK. This compact system is quite robust, allowing for versatile extensions and further studying of BEC.

physics.atom-ph↗