SearcharxivSearch

arXiv subjects

Shiqi Yu

Publications and source records attributed to Shiqi Yu.

44 records · Page 3Linked to original sources

Detect Faces Efficiently: A Survey and Evaluations

Face detection is to search all the possible regions for faces in images and locate the faces if there are any. Many applications including face recognition, facial expression recognition, face tracking and head-pose estimation assume that both the location and the size of faces are known in the image. In recent decades, researchers have created many typical and efficient face detectors from the Viola-Jones face detector to current CNN-based ones. However, with the tremendous increase in images and videos with variations in face scale, appearance, expression, occlusion and pose, traditional face detectors are challenged to detect various "in the wild" faces. The emergence of deep learning techniques brought remarkable breakthroughs to face detection along with the price of a considerable increase in computation. This paper introduces representative deep learning-based methods and presents a deep and thorough analysis in terms of accuracy and efficiency. We further compare and discuss the popular and challenging datasets and their evaluation metrics. A comprehensive comparison of several successful deep learning-based face detectors is conducted to uncover their efficiency using two metrics: FLOPs and latency. The paper can guide to choose appropriate face detectors for different applications and also to develop more efficient and accurate detectors.

cs.CV

A Systematic IoU-Related Method: Beyond Simplified Regression for Better Localization

Four-variable-independent-regression localization losses, such as Smooth-$\ell_1$ Loss, are used by default in modern detectors. Nevertheless, this kind of loss is oversimplified so that it is inconsistent with the final evaluation metric, intersection over union (IoU). Directly employing the standard IoU is also not infeasible, since the constant-zero plateau in the case of non-overlapping boxes and the non-zero gradient at the minimum may make it not trainable. Accordingly, we propose a systematic method to address these problems. Firstly, we propose a new metric, the extended IoU (EIoU), which is well-defined when two boxes are not overlapping and reduced to the standard IoU when overlapping. Secondly, we present the convexification technique (CT) to construct a loss on the basis of EIoU, which can guarantee the gradient at the minimum to be zero. Thirdly, we propose a steady optimization technique (SOT) to make the fractional EIoU loss approaching the minimum more steadily and smoothly. Fourthly, to fully exploit the capability of the EIoU based loss, we introduce an interrelated IoU-predicting head to further boost localization accuracy. With the proposed contributions, the new method incorporated into Faster R-CNN with ResNet50+FPN as the backbone yields \textbf{4.2 mAP} gain on VOC2007 and \textbf{2.3 mAP} gain on COCO2017 over the baseline Smooth-$\ell_1$ Loss, at almost \textbf{no training and inferencing computational cost}. Specifically, the stricter the metric is, the more notable the gain is, improving \textbf{8.2 mAP} on VOC2007 and \textbf{5.4 mAP} on COCO2017 at metric $AP_{90}$.

cs.CV

Direction Reconstruction using a CNN for GeV-Scale Neutrinos in IceCube

The IceCube Neutrino Observatory observes neutrinos interacting deep within the South Pole ice. It consists of 5,160 digital optical modules embedded within a cubic kilometer of ice, over depths of 1,450 m to 2,450 m. At the lower center of the array is the DeepCore subdetector. Its denser sensor configuration lowers the observable energy threshold to the GeV-scale, facilitating the study of atmospheric neutrino oscillations. The precise reconstruction of neutrino direction is critical in the measurements of oscillation parameters. This work presents a method to reconstruct the zenith angle of GeV-scale events in IceCube by using a convolutional neural network and compares the result to that of the current likelihood-based reconstruction algorithm.

astro-ph.HE

Dense-View GEIs Set: View Space Covering for Gait Recognition based on Dense-View GAN

Gait recognition has proven to be effective for long-distance human recognition. But view variance of gait features would change human appearance greatly and reduce its performance. Most existing gait datasets usually collect data with a dozen different angles, or even more few. Limited view angles would prevent learning better view invariant feature. It can further improve robustness of gait recognition if we collect data with various angles at 1 degree interval. But it is time consuming and labor consuming to collect this kind of dataset. In this paper, we, therefore, introduce a Dense-View GEIs Set (DV-GEIs) to deal with the challenge of limited view angles. This set can cover the whole view space, view angle from 0 degree to 180 degree with 1 degree interval. In addition, Dense-View GAN (DV-GAN) is proposed to synthesize this dense view set. DV-GAN consists of Generator, Discriminator and Monitor, where Monitor is designed to preserve human identification and view information. The proposed method is evaluated on the CASIA-B and OU-ISIR dataset. The experimental results show that DV-GEIs synthesized by DV-GAN is an effective way to learn better view invariant feature. We believe the idea of dense view generated samples will further improve the development of gait recognition.

cs.CV

iLGaCo: Incremental Learning of Gait Covariate Factors

Gait is a popular biometric pattern used for identifying people based on their way of walking. Traditionally, gait recognition approaches based on deep learning are trained using the whole training dataset. In fact, if new data (classes, view-points, walking conditions, etc.) need to be included, it is necessary to re-train again the model with old and new data samples. In this paper, we propose iLGaCo, the first incremental learning approach of covariate factors for gait recognition, where the deep model can be updated with new information without re-training it from scratch by using the whole dataset. Instead, our approach performs a shorter training process with the new data and a small subset of previous samples. This way, our model learns new information while retaining previous knowledge. We evaluate iLGaCo on CASIA-B dataset in two incremental ways: adding new view-points and adding new walking conditions. In both cases, our results are close to the classical `training-from-scratch' approach, obtaining a marginal drop in accuracy ranging from 0.2% to 1.2%, what shows the efficacy of our approach. In addition, the comparison of iLGaCo with other incremental learning methods, such as LwF and iCarl, shows a significant improvement in accuracy, between 6% and 15% depending on the experiment.

cs.CV

Electron Neutrino Energy Reconstruction in NOvA Using CNN Particle IDs

NOvA is a long-baseline neutrino oscillation experiment. It is optimized to measure $ν_e$ appearance and $ν_μ$ disappearance at the Far Detector in the $ν_μ$ beam produced by the NuMI facility at Fermilab. NOvA uses a convolutional neural network (CNN) to identify neutrino events in two functionally identical liquid scintillator detectors. A different network, called prong-CNN, has been used to classify reconstructed particles in each event as either lepton or hadron. Within each event, hits are clustered into prongs to reconstruct final-state particles and these prongs form the input to this prong-CNN classifier. Classified particle energies are then used as input to an electron neutrino energy estimator. Improving the resolution and systematic robustness of NOvA's energy estimator will improve the sensitivity of the oscillation parameters measurement. This paper describes the methods to identify particles with prong-CNN and the following approach to estimate $ν_e$ energy for signal events.

physics.ins-det

Cherenkov Light in Liquid Scintillator at the NOvA Experiment

NOvA is a long-baseline neutrino oscillation experiment with two functionally identical liquid scintillator tracking detectors, i.e. the Near Detector (ND) and the Far Detector (FD). One of NOvA's physics goal is to measure neutrino oscillation parameters by studying $ν_e$ appearance and $ν_μ$ disappearance with the Neutrinos at the Main Injector (NuMI) beam at Fermi National Accelerator Laboratory. An accurate light model is a prerequisite for precise charged particle energy estimation in the detectors. Particle energy is needed in event reconstruction and classification, both of which are critical to constraining neutrino oscillation parameters. In this paper, I will explain the details of the data-driven tuning of the Cherenkov model and the impact of the new scintillator model

physics.ins-det

PEA265: Perceptual Assessment of Video Compression Artifacts

The most widely used video encoders share a common hybrid coding framework that includes block-based motion estimation/compensation and block-based transform coding. Despite their high coding efficiency, the encoded videos often exhibit visually annoying artifacts, denoted as Perceivable Encoding Artifacts (PEAs), which significantly degrade the visual Qualityof- Experience (QoE) of end users. To monitor and improve visual QoE, it is crucial to develop subjective and objective measures that can identify and quantify various types of PEAs. In this work, we make the first attempt to build a large-scale subjectlabelled database composed of H.265/HEVC compressed videos containing various PEAs. The database, namely the PEA265 database, includes 4 types of spatial PEAs (i.e. blurring, blocking, ringing and color bleeding) and 2 types of temporal PEAs (i.e. flickering and floating). Each containing at least 60,000 image or video patches with positive and negative labels. To objectively identify these PEAs, we train Convolutional Neural Networks (CNNs) using the PEA265 database. It appears that state-of-theart ResNeXt is capable of identifying each type of PEAs with high accuracy. Furthermore, we define PEA pattern and PEA intensity measures to quantify PEA levels of compressed video sequence. We believe that the PEA265 database and our findings will benefit the future development of video quality assessment methods and perceptually motivated video encoders.

cs.CV