SearcharxivSearch

arXiv subjects

Chengyang Hu

Publications and source records attributed to Chengyang Hu.

4 recordsLinked to original sources

BEAR: A Video Dataset For Fine-grained Behaviors Recognition Oriented with Action and Environment Factors

Behavior recognition is an important task in video representation learning. An essential aspect pertains to effective feature learning conducive to behavior recognition. Recently, researchers have started to study fine-grained behavior recognition, which provides similar behaviors and encourages the model to concern with more details of behaviors with effective features for distinction. However, previous fine-grained behaviors limited themselves to controlling partial information to be similar, leading to an unfair and not comprehensive evaluation of existing works. In this work, we develop a new video fine-grained behavior dataset, named BEAR, which provides fine-grained (i.e. similar) behaviors that uniquely focus on two primary factors defining behavior: Environment and Action. It includes two fine-grained behavior protocols including Fine-grained Behavior with Similar Environments and Fine-grained Behavior with Similar Actions as well as multiple sub-protocols as different scenarios. Furthermore, with this new dataset, we conduct multiple experiments with different behavior recognition models. Our research primarily explores the impact of input modality, a critical element in studying the environmental and action-based aspects of behavior recognition. Our experimental results yield intriguing insights that have substantial implications for further research endeavors.

cs.CV

Key frames assisted hybrid encoding for photorealistic compressive video sensing

Snapshot compressive imaging (SCI) encodes high-speed scene video into a snapshot measurement and then computationally makes reconstructions, allowing for efficient high-dimensional data acquisition. Numerous algorithms, ranging from regularization-based optimization and deep learning, are being investigated to improve reconstruction quality, but they are still limited by the ill-posed and information-deficient nature of the standard SCI paradigm. To overcome these drawbacks, we propose a new key frames assisted hybrid encoding paradigm for compressive video sensing, termed KH-CVS, that alternatively captures short-exposure key frames without coding and long-exposure encoded compressive frames to jointly reconstruct photorealistic video. With the use of optical flow and spatial warping, a deep convolutional neural network framework is constructed to integrate the benefits of these two types of frames. Extensive experiments on both simulations and real data from the prototype we developed verify the superiority of the proposed method.

eess.IV

Fourier temporal ghost imaging

Ghost imaging is a fascinating framework which constructs the image of an object by correlating measurements between received beams and reference beams, none of which carries the structure information of the object independently. Recently, by taking into account space-time duality in optics, computational temporal ghost imaging has attracted attentions. Here, we propose a novel Fourier temporal ghost imaging (FTGI) scheme to achieve single-shot non-reproducible temporal signals. By sinusoidal coded modulation, ghost images are obtained and recovered by applying Fourier transformation. For demonstration, non-repeating events are detected with single-shot exposure architecture. It's shown in results that the peak signal-to-noise ratio (PSNR) of FTGI is significantly better (13dB increase) than traditional temporal ghost imaging in the same condition. In addition, by using the obvious physical meaning of Fourier spectrum, we show some potential applications of FTGI, such as frequency division multiplexing demodulation in the visible light communications.

eess.SP

FourierCam:A camera for video spectrum acquisition in single-shot

The novel camera architecture facilitates the development of machine vision. Instead of capturing frame sequences in the temporal domain as traditional video cameras, FourierCam directly measures the pixel-wise temporal spectrum of the video in single-shot through optical coding. Compared with the classic video cameras and time-frequency transformation pipeline, this programmable frequency-domain sampling strategy has an attractive combination of characteristics for low detection bandwidth, low-light imaging, low computational burden and low data volume. Based on the various temporal filter kernel designed by FourierCam, we demonstrated a series of exciting machine vision functions, such as video compression, background subtraction, object extraction, and trajectory tracking.

eess.IV