Searcharxiv⌕ Search

arXiv subjects

Jie Lin

Publications and source records attributed to Jie Lin.

At least 163 records · Page 9Linked to original sources

Compact Descriptors for Video Analysis: the Emerging MPEG Standard

This paper provides an overview of the on-going compact descriptors for video analysis standard (CDVA) from the ISO/IEC moving pictures experts group (MPEG). MPEG-CDVA targets at defining a standardized bitstream syntax to enable interoperability in the context of video analysis applications. During the developments of MPEGCDVA, a series of techniques aiming to reduce the descriptor size and improve the video representation ability have been proposed. This article describes the new standard that is being developed and reports the performance of these key technical contributions.

cs.CV↗

Adaptive Depth Imaging with Single-Photon Detectors

For active optical imaging, the use of single-photon detectors can greatly improve the detection sensitivity of the system. However, the traditional maximum-likelihood based imaging method needs a long acquisition time to capture clear three-dimensional (3D) image in low light-level. To tackle this problem, we present a novel imaging method for depth estimate, which can obtain the accurate 3D image in a short acquisition time. Our method combines the photon-count statistics with the temporal correlations of the reflected signal. According to the characteristics of the target surface, including the surface reflectivity, our method is capable of adaptively changing the dwell time in each pixel. The experimental results demonstrate that the proposed method can fast obtain the accurate depth image despite the existence of strong background noise.

physics.optics↗

Compression of Deep Neural Networks for Image Instance Retrieval

Image instance retrieval is the problem of retrieving images from a database which contain the same object. Convolutional Neural Network (CNN) based descriptors are becoming the dominant approach for generating {\it global image descriptors} for the instance retrieval problem. One major drawback of CNN-based {\it global descriptors} is that uncompressed deep neural network models require hundreds of megabytes of storage making them inconvenient to deploy in mobile applications or in custom hardware. In this work, we study the problem of neural network model compression focusing on the image instance retrieval task. We study quantization, coding, pruning and weight sharing techniques for reducing model size for the instance retrieval problem. We provide extensive experimental results on the trade-off between retrieval performance and model size for different types of networks on several data sets providing the most comprehensive study on this topic. We compress models to the order of a few MBs: two orders of magnitude smaller than the uncompressed models while achieving negligible loss in retrieval performance.

cs.CV↗

Cheating-Resilient Incentive Scheme for Mobile Crowdsensing Systems

Mobile Crowdsensing is a promising paradigm for ubiquitous sensing, which explores the tremendous data collected by mobile smart devices with prominent spatial-temporal coverage. As a fundamental property of Mobile Crowdsensing Systems, temporally recruited mobile users can provide agile, fine-grained, and economical sensing labors, however their self-interest cannot guarantee the quality of the sensing data, even when there is a fair return. Therefore, a mechanism is required for the system server to recruit well-behaving users for credible sensing, and to stimulate and reward more contributive users based on sensing truth discovery to further increase credible reporting. In this paper, we develop a novel Cheating-Resilient Incentive (CRI) scheme for Mobile Crowdsensing Systems, which achieves credibility-driven user recruitment and payback maximization for honest users with quality data. Via theoretical analysis, we demonstrate the correctness of our design. The performance of our scheme is evaluated based on extensive realworld trace-driven simulations. Our evaluation results show that our scheme is proven to be effective in terms of both guaranteeing sensing accuracy and resisting potential cheating behaviors, as demonstrated in practical scenarios, as well as those that are intentionally harsher.

cs.NI↗

Galois equivariance of critical values of $L$-functions for unitary groups

The goal of this paper is to provide a refinement of a formula proved by the first author which expresses some critical values of automorphic $L$-functions on unitary groups as Petersson norms of automorphic forms. Here we provide a Galois equivariant version of the formula. We also give some applications to special values of automorphic representations of $\GL_{n}\times\GL_{1}$. We show that our results are compatible with Deligne's conjecture.

math.NT↗

Period relations and special values of Rankin-Selberg $L$-functions

This is a survey of recent work on values of Rankin-Selberg $L$-functions of pairs of cohomological automorphic representations that are {\it critical} in Deligne's sense. The base field is assumed to be a CM field. Deligne's conjecture is stated in the language of motives over $\QQ$, and express the critical values, up to rational factors, as determinants of certain periods of algebraic differentials on a projective algebraic variety over homology classes. The results that can be proved by automorphic methods express certain critical values as (twisted) period integrals of automorphic forms. Using Langlands functoriality between cohomological automorphic representations of unitary groups, which can be identified with the de Rham cohomology of Shimura varieties, and cohomological automorphic representations of $GL(n)$, the automorphic periods can be interpreted as motivic periods. We report on recent results of the two authors, of the first-named author with Grobner, and of Guerberoff.

math.NT↗

Scaling Description of Non-Local Rheology

Non-locality is crucial to understand the plastic flow of an amorphous material, and has been successfully described by the fluidity, along with a cooperativity length scale ξ. We demonstrate, by applying the scaling hypothesis to the yielding transition, that non-local effects in non-uniform stress configurations can be explained within the framework of critical phenomena. From the scaling description, scaling relations between different exponents are derived, and collapses of strain rate profiles are made both in shear driven and pressure driven flow. We find that the cooperative length in non-local flow is governed by the same correlation length in finite dimensional homogeneous flow, excluding the mean field exponents. We also show that non-locality also affects the finite size scaling of the yield stress, especially the large finite size effects observed in pressure driven flow. Our theoretical results are nicely verified by the elasto-plastic model, and experimental data.

cond-mat.soft↗

Evidence for marginal stability in emulsions

We report the first measurements of the effect of pressure on vibrational modes in emulsions, which serve as a model for soft frictionless spheres at zero temperature. As a function of the applied pressure, we find that the density of states D(omega) exhibits a low-frequency cutoff omega*, which scales linearly with the number of extra contacts per particle dz. Moreover, for omega<omega*, D(omega)~ omega^2/omega*^2; a quadratic behavior whose prefactor is larger than what is expected from Debye theory. This surprising result agrees with recent theoretical findings. Finally, the degree of localization of the softest low frequency modes increases with compression, as shown by the participation ratio as well as their spatial configurations. Overall, our observations show that emulsions are marginally stable and display non-plane-wave modes up to vanishing frequencies.

cond-mat.soft↗

Nested Invariance Pooling and RBM Hashing for Image Instance Retrieval

The goal of this work is the computation of very compact binary hashes for image instance retrieval. Our approach has two novel contributions. The first one is Nested Invariance Pooling (NIP), a method inspired from i-theory, a mathematical theory for computing group invariant transformations with feed-forward neural networks. NIP is able to produce compact and well-performing descriptors with visual representations extracted from convolutional neural networks. We specifically incorporate scale, translation and rotation invariances but the scheme can be extended to any arbitrary sets of transformations. We also show that using moments of increasing order throughout nesting is important. The NIP descriptors are then hashed to the target code size (32-256 bits) with a Restricted Boltzmann Machine with a novel batch-level regularization scheme specifically designed for the purpose of hashing (RBMH). A thorough empirical evaluation with state-of-the-art shows that the results obtained both with the NIP descriptors and the NIP+RBMH hashes are consistently outstanding across a wide range of datasets.

cs.CV↗

Egocentric Activity Recognition with Multimodal Fisher Vector

With the increasing availability of wearable devices, research on egocentric activity recognition has received much attention recently. In this paper, we build a Multimodal Egocentric Activity dataset which includes egocentric videos and sensor data of 20 fine-grained and diverse activity categories. We present a novel strategy to extract temporal trajectory-like features from sensor data. We propose to apply the Fisher Kernel framework to fuse video and temporal enhanced sensor features. Experiment results show that with careful design of feature extraction and fusion algorithm, sensor data can enhance information-rich video data. We make publicly available the Multimodal Egocentric Activity dataset to facilitate future research.

cs.MM↗

Group Invariant Deep Representations for Image Instance Retrieval

Most image instance retrieval pipelines are based on comparison of vectors known as global image descriptors between a query image and the database images. Due to their success in large scale image classification, representations extracted from Convolutional Neural Networks (CNN) are quickly gaining ground on Fisher Vectors (FVs) as state-of-the-art global descriptors for image instance retrieval. While CNN-based descriptors are generally remarked for good retrieval performance at lower bitrates, they nevertheless present a number of drawbacks including the lack of robustness to common object transformations such as rotations compared with their interest point based FV counterparts. In this paper, we propose a method for computing invariant global descriptors from CNNs. Our method implements a recently proposed mathematical theory for invariance in a sensory cortex modeled as a feedforward neural network. The resulting global descriptors can be made invariant to multiple arbitrary transformation groups while retaining good discriminativeness. Based on a thorough empirical evaluation using several publicly available datasets, we show that our method is able to significantly and consistently improve retrieval results every time a new type of invariance is incorporated. We also show that our method which has few parameters is not prone to overfitting: improvements generalize well across datasets with different properties with regard to invariances. Finally, we show that our descriptors are able to compare favourably to other state-of-the-art compact descriptors in similar bitranges, exceeding the highest retrieval results reported in the literature on some datasets. A dedicated dimensionality reduction step --quantization or hashing-- may be able to further improve the competitiveness of the descriptors.

cs.CV↗

Special values of automorphic $L$-functions for $GL_{n}\times GL_{n'}$ over CM fields, factorization and functoriality of arithmetic automorphic periods

Michael HARRIS defined the arithmetic automorphic periods for certain cuspidal representations of $GL_{n}$ over quadratic imaginary fields in his Crelle paper 1997. He also showed that critical values of automorphic L-functions for $GL_{n}\times GL_{1}$ can be interpreted in terms of these arithmetic automorphic periods. In the thesis, we generalize his results in two ways. Firstly, the arithmetic automorphic periods have been defined over general CM fields. We also prove that these periods factorize as products of local periods over infinity places. Secondly, we show that critical values of automorphic $L$ functions for $GL_{n}\times GL_{n'}$ can be interpreted in terms of these automorphic periods in many situations. Consequently we show that the automorphic periods are functorial for automorphic induction and cyclic base change. We also define certain motivic periods if the motive is restricted from a CM field to the field of rational numbers. We can calculate Deligne's period for tensor product of two such motives. We see directly that our automorphic results are compatible with Deligne's conjecture for motives.

math.NT↗

Period relations for automorphic induction and applications, I

Let $K$ be a quadratic imaginary field. Let $Π$ (resp. $Π'$) be a regular algebraic cuspidal representation of $GL_{n}(K)$ (resp. $GL_{n-1}(K)$) which is moreover cohomological and conjugate self-dual. In \cite{harris97}, M. Harris has defined automorphic periods of such a representation. These periods are automorphic analogues of motivic periods. In this paper, we show that automorphic periods are functorial in the case where $Π$ is a cyclic automorphic induction of a Hecke character $χ$ over a CM field. More precisely, we prove relations between automorphic periods of $Π$ and those of $χ$. As a corollary, we refine the formula given by H. Grobner and M. Harris of critical values for the Rankin-Selberg $L$-function $L(s,Π\times Π')$ in terms of automorphic periods. This completes the proof of an automorphic version of Deligne's conjecture in certain cases.

math.NT↗

Tiny Descriptors for Image Retrieval with Unsupervised Triplet Hashing

A typical image retrieval pipeline starts with the comparison of global descriptors from a large database to find a short list of candidate matches. A good image descriptor is key to the retrieval pipeline and should reconcile two contradictory requirements: providing recall rates as high as possible and being as compact as possible for fast matching. Following the recent successes of Deep Convolutional Neural Networks (DCNN) for large scale image classification, descriptors extracted from DCNNs are increasingly used in place of the traditional hand crafted descriptors such as Fisher Vectors (FV) with better retrieval performances. Nevertheless, the dimensionality of a typical DCNN descriptor --extracted either from the visual feature pyramid or the fully-connected layers-- remains quite high at several thousands of scalar values. In this paper, we propose Unsupervised Triplet Hashing (UTH), a fully unsupervised method to compute extremely compact binary hashes --in the 32-256 bits range-- from high-dimensional global descriptors. UTH consists of two successive deep learning steps. First, Stacked Restricted Boltzmann Machines (SRBM), a type of unsupervised deep neural nets, are used to learn binary embedding functions able to bring the descriptor size down to the desired bitrate. SRBMs are typically able to ensure a very high compression rate at the expense of loosing some desirable metric properties of the original DCNN descriptor space. Then, triplet networks, a rank learning scheme based on weight sharing nets is used to fine-tune the binary embedding functions to retain as much as possible of the useful metric properties of the original space. A thorough empirical evaluation conducted on multiple publicly available dataset using DCNN descriptors shows that our method is able to significantly outperform state-of-the-art unsupervised schemes in the target bit range.

cs.IR↗

Criticality in the approach to failure in amorphous solids

Failure of amorphous solids is fundamental to various phenomena, including landslides and earthquakes. Recent experiments indicate that highly plastic regions form elongated structures that are especially apparent near the maximal shear stress $Σ_{\max}$ where failure occurs. This observation suggested that $Σ_{\max}$ acts as a critical point where the length scale of those structures diverges, possibly causing macroscopic transient shear bands. Here we argue instead that the entire solid phase ($Σ<Σ_{\max}$) is critical, that plasticity always involves system-spanning events, and that their magnitude diverges at $Σ_{\max}$ independently of the presence of shear bands. We relate the statistics and fractal properties of these rearrangements to an exponent $θ$ that captures the stability of the material, which is observed to vary continuously with stress, and we confirm our predictions in elastoplastic models.

cond-mat.soft↗

A Practical Guide to CNNs and Fisher Vectors for Image Instance Retrieval

With deep learning becoming the dominant approach in computer vision, the use of representations extracted from Convolutional Neural Nets (CNNs) is quickly gaining ground on Fisher Vectors (FVs) as favoured state-of-the-art global image descriptors for image instance retrieval. While the good performance of CNNs for image classification are unambiguously recognised, which of the two has the upper hand in the image retrieval context is not entirely clear yet. In this work, we propose a comprehensive study that systematically evaluates FVs and CNNs for image retrieval. The first part compares the performances of FVs and CNNs on multiple publicly available data sets. We investigate a number of details specific to each method. For FVs, we compare sparse descriptors based on interest point detectors with dense single-scale and multi-scale variants. For CNNs, we focus on understanding the impact of depth, architecture and training data on retrieval results. Our study shows that no descriptor is systematically better than the other and that performance gains can usually be obtained by using both types together. The second part of the study focuses on the impact of geometrical transformations such as rotations and scale changes. FVs based on interest point detectors are intrinsically resilient to such transformations while CNNs do not have a built-in mechanism to ensure such invariance. We show that performance of CNNs can quickly degrade in presence of rotations while they are far less affected by changes in scale. We then propose a number of ways to incorporate the required invariances in the CNN pipeline. Overall, our work is intended as a reference guide offering practically useful and simply implementable guidelines to anyone looking for state-of-the-art global descriptors best suited to their specific image instance retrieval problem.

cs.CV↗

Co-Regularized Deep Representations for Video Summarization

Compact keyframe-based video summaries are a popular way of generating viewership on video sharing platforms. Yet, creating relevant and compelling summaries for arbitrarily long videos with a small number of keyframes is a challenging task. We propose a comprehensive keyframe-based summarization framework combining deep convolutional neural networks and restricted Boltzmann machines. An original co-regularization scheme is used to discover meaningful subject-scene associations. The resulting multimodal representations are then used to select highly-relevant keyframes. A comprehensive user study is conducted comparing our proposed method to a variety of schemes, including the summarization currently in use by one of the most popular video sharing websites. The results show that our method consistently outperforms the baseline schemes for any given amount of keyframes both in terms of attractiveness and informativeness. The lead is even more significant for smaller summaries.

cs.CV↗

DeepHash: Getting Regularization, Depth and Fine-Tuning Right

This work focuses on representing very high-dimensional global image descriptors using very compact 64-1024 bit binary hashes for instance retrieval. We propose DeepHash: a hashing scheme based on deep networks. Key to making DeepHash work at extremely low bitrates are three important considerations -- regularization, depth and fine-tuning -- each requiring solutions specific to the hashing problem. In-depth evaluation shows that our scheme consistently outperforms state-of-the-art methods across all data sets for both Fisher Vectors and Deep Convolutional Neural Network features, by up to 20 percent over other schemes. The retrieval performance with 256-bit hashes is close to that of the uncompressed floating point features -- a remarkable 512 times compression.

cs.CV↗