Searcharxiv⌕ Search

arXiv subjects

Cheng Lu

Publications and source records attributed to Cheng Lu.

90 records · Page 5Linked to original sources

Virtual Multi-Modality Self-Supervised Foreground Matting for Human-Object Interaction

Most existing human matting algorithms tried to separate pure human-only foreground from the background. In this paper, we propose a Virtual Multi-modality Foreground Matting (VMFM) method to learn human-object interactive foreground (human and objects interacted with him or her) from a raw RGB image. The VMFM method requires no additional inputs, e.g. trimap or known background. We reformulate foreground matting as a self-supervised multi-modality problem: factor each input image into estimated depth map, segmentation mask, and interaction heatmap using three auto-encoders. In order to fully utilize the characteristics of each modality, we first train a dual encoder-to-decoder network to estimate the same alpha matte. Then we introduce a self-supervised method: Complementary Learning(CL) to predict deviation probability map and exchange reliable gradients across modalities without label. We conducted extensive experiments to analyze the effectiveness of each modality and the significance of different components in complementary learning. We demonstrate that our model outperforms the state-of-the-art methods.

cs.CV↗

Deep Two-Stream Video Inference for Human Body Pose and Shape Estimation

Several video-based 3D pose and shape estimation algorithms have been proposed to resolve the temporal inconsistency of single-image-based methods. However it still remains challenging to have stable and accurate reconstruction. In this paper, we propose a new framework Deep Two-Stream Video Inference for Human Body Pose and Shape Estimation (DTS-VIBE), to generate 3D human pose and mesh from RGB videos. We reformulate the task as a multi-modality problem that fuses RGB and optical flow for more reliable estimation. In order to fully utilize both sensory modalities (RGB or optical flow), we train a two-stream temporal network based on transformer to predict SMPL parameters. The supplementary modality, optical flow, helps to maintain temporal consistency by leveraging motion knowledge between two consecutive frames. The proposed algorithm is extensively evaluated on the Human3.6 and 3DPW datasets. The experimental results show that it outperforms other state-of-the-art methods by a significant margin.

cs.CV↗

Regularized Non-monotone Submodular Maximization

In this paper, we present a thorough study of maximizing a regularized non-monotone submodular function subject to various constraints, i.e., $\max \{ g(A) - \ell(A) : A \in \mathcal{F} \}$, where $g \colon 2^Ω\to \mathbb{R}_+$ is a non-monotone submodular function, $\ell \colon 2^Ω\to \mathbb{R}_+$ is a normalized modular function and $\mathcal{F}$ is the constraint set. Though the objective function $f := g - \ell$ is still submodular, the fact that $f$ could potentially take on negative values prevents the existing methods for submodular maximization from providing a constant approximation ratio for the regularized submodular maximization problem. To overcome the obstacle, we propose several algorithms which can provide a relatively weak approximation guarantee for maximizing regularized non-monotone submodular functions. More specifically, we propose a continuous greedy algorithm for the relaxation of maximizing $g - \ell$ subject to a matroid constraint. Then, the pipage rounding procedure can produce an integral solution $S$ such that $\mathbb{E} [g(S) - \ell(S)] \geq e^{-1}g(OPT) - \ell(OPT) - O(ε)$. Moreover, we present a much faster algorithm for maximizing $g - \ell$ subject to a cardinality constraint, which can output a solution $S$ with $\mathbb{E} [g(S) - \ell(S)] \geq (e^{-1} - ε) g(OPT) - \ell(OPT)$ using $O(\frac{n}{ε^2} \ln \frac 1ε)$ value oracle queries. We also consider the unconstrained maximization problem and give an algorithm which can return a solution $S$ with $\mathbb{E} [g(S) - \ell(S)] \geq e^{-1} g(OPT) - \ell(OPT)$ using $O(n)$ value oracle queries.

cs.DS↗

Implicit Normalizing Flows

Normalizing flows define a probability distribution by an explicit invertible transformation $\boldsymbol{\mathbf{z}}=f(\boldsymbol{\mathbf{x}})$. In this work, we present implicit normalizing flows (ImpFlows), which generalize normalizing flows by allowing the mapping to be implicitly defined by the roots of an equation $F(\boldsymbol{\mathbf{z}}, \boldsymbol{\mathbf{x}})= \boldsymbol{\mathbf{0}}$. ImpFlows build on residual flows (ResFlows) with a proper balance between expressiveness and tractability. Through theoretical analysis, we show that the function space of ImpFlow is strictly richer than that of ResFlows. Furthermore, for any ResFlow with a fixed number of blocks, there exists some function that ResFlow has a non-negligible approximation error. However, the function is exactly representable by a single-block ImpFlow. We propose a scalable algorithm to train and draw samples from ImpFlows. Empirically, we evaluate ImpFlow on several classification and density modeling tasks, and ImpFlow outperforms ResFlow with a comparable amount of parameters on all the benchmarks.

stat.ML↗

An Enhanced SDR based Global Algorithm for Nonconvex Complex Quadratic Programs with Signal Processing Applications

In this paper, we consider a class of nonconvex complex quadratic programming (CQP) problems, which find a broad spectrum of signal processing applications. By using the polar coordinate representations of the complex variables, we first derive a new enhanced semidefinite relaxation (SDR) for problem (CQP). Based on the newly derived SDR, we further propose an efficient branch-and-bound algorithm for solving problem (CQP). Key features of our proposed algorithm are: (1) it is guaranteed to find the global solution of the problem (within any given error tolerance); (2) it is computationally efficient because it carefully utilizes the special structure of the problem. We apply our proposed algorithm to solve the multi-input multi-output (MIMO) detection problem, the unimodular radar code design problem, and the virtual beamforming design problem. Simulation results show that our proposed enhanced SDR, when applied to the above problems, is generally much tighter than the conventional SDR and our proposed global algorithm can efficiently solve these problems. In particular, our proposed algorithm significantly outperforms the state-of-the-art sphere decode algorithm for solving the MIMO detection problem in the hard cases (where the number of inputs and outputs is equal or the SNR is low) and a state-of-the-art general-purpose global optimization solver called Baron for solving the virtual beamforming design problem.

cs.IT↗

DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the Wild

Recently, facial expression recognition (FER) in the wild has gained a lot of researchers' attention because it is a valuable topic to enable the FER techniques to move from the laboratory to the real applications. In this paper, we focus on this challenging but interesting topic and make contributions from three aspects. First, we present a new large-scale 'in-the-wild' dynamic facial expression database, DFEW (Dynamic Facial Expression in the Wild), consisting of over 16,000 video clips from thousands of movies. These video clips contain various challenging interferences in practical scenarios such as extreme illumination, occlusions, and capricious pose changes. Second, we propose a novel method called Expression-Clustered Spatiotemporal Feature Learning (EC-STFL) framework to deal with dynamic FER in the wild. Third, we conduct extensive benchmark experiments on DFEW using a lot of spatiotemporal deep feature learning methods as well as our proposed EC-STFL. Experimental results show that DFEW is a well-designed and challenging database, and the proposed EC-STFL can promisingly improve the performance of existing spatiotemporal deep neural networks in coping with the problem of dynamic FER in the wild. Our DFEW database is publicly available and can be freely downloaded from https://dfew-dataset.github.io/.

cs.CV↗

VFlow: More Expressive Generative Flows with Variational Data Augmentation

Generative flows are promising tractable models for density modeling that define probabilistic distributions with invertible transformations. However, tractability imposes architectural constraints on generative flows, making them less expressive than other types of generative models. In this work, we study a previously overlooked constraint that all the intermediate representations must have the same dimensionality with the original data due to invertibility, limiting the width of the network. We tackle this constraint by augmenting the data with some extra dimensions and jointly learning a generative flow for augmented data as well as the distribution of augmented dimensions under a variational inference framework. Our approach, VFlow, is a generalization of generative flows and therefore always performs better. Combining with existing generative flows, VFlow achieves a new state-of-the-art 2.98 bits per dimension on the CIFAR-10 dataset and is more compact than previous models to reach similar modeling quality.

stat.ML↗

Discriminative Multi-modality Speech Recognition

Vision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded to the multi-modality speech recognition (MSR). In this paper, we propose a two-stage speech recognition model. In the first stage, the target voice is separated from background noises with help from the corresponding visual information of lip movements, making the model 'listen' clearly. At the second stage, the audio modality combines visual modality again to better understand the speech by a MSR sub-network, further improving the recognition rate. There are some other key contributions: we introduce a pseudo-3D residual convolution (P3D)-based visual front-end to extract more discriminative features; we upgrade the temporal convolution block from 1D ResNet with the temporal convolutional network (TCN), which is more suitable for the temporal tasks; the MSR sub-network is built on the top of Element-wise-Attention Gated Recurrent Unit (EleAtt-GRU), which is more effective than Transformer in long sequences. We conducted extensive experiments on the LRS3-TED and the LRW datasets. Our two-stage model (audio enhanced multi-modality speech recognition, AE-MSR) consistently achieves the state-of-the-art performance by a significant margin, which demonstrates the necessity and effectiveness of AE-MSR.

cs.CV↗

Structure and Algorithm for Path of Solutions to a Class of Fused Lasso Problems

We study a class of fused lasso problems where the estimated parameters in a sequence are regressed toward their respective observed values (fidelity loss), with $\ell_1$ norm penalty (regularization loss) on the differences between successive parameters, which promotes local constancy. In many applications, there is a coefficient, often denoted as $λ$, on the regularization term, which adjusts the relative importance between the two losses. In this paper, we characterize how the optimal solution evolves with the increment of $λ$. We show that, if all fidelity loss functions are convex piecewise linear, the optimal value for \emph{each} variable changes at most $O(nq)$ times for a problem of $n$ variables and total $q$ breakpoints. On the other hand, we present an algorithm that solves the path of solutions of \emph{all} variables in $\tilde{O}(nq)$ time for all $λ\geq 0$. Interestingly, we find that the path of solutions for each variable can be divided into up to $n$ locally convex-like segments. For problems of arbitrary convex loss functions, for a given solution accuracy, one can transform the loss functions into convex piecewise linear functions and apply the above results, giving pseudo-polynomial bounds as $q$ becomes a pseudo-polynomial quantity. To our knowledge, this is the first work to solve the path of solutions for fused lasso of non-quadratic fidelity loss functions.

cs.DS↗

Prediction of mechanical properties of non-equiatomic high-entropy alloy by atomistic simulation and machine learning

High-entropy alloys (HEAs) with multiple constituent elements have been extensively studied in the past 20 years due to their promising engineering application. Previous experimental and computational studies of HEAs focused mainly on equiatomic or near equiatomic HEAs. However, there is probably far more treasure in those non-equiatomic HEAs with carefully designed composition. In this study, molecular dynamics (MD) simulation combined with machine learning (ML) methods were used to predict the mechanical properties of non-equiatomic CuFeNiCrCo HEAs. A database was established based on a tensile test of 900 HEA single-crystal samples by MD simulation. We investigated and compared eight ML models for the learning tasks, ranging from shallow models to deep models. It was found that the kernel-based extreme learning machine (KELM) model outperformed others for the prediction of yield stress and Young's modulus. The accuracy of the KELM model was further verified by the large-sized polycrystal HEA samples.

cond-mat.mtrl-sci↗

Learning to Detect Head Movement in Unconstrained Remote Gaze Estimation in the Wild

Unconstrained remote gaze estimation remains challenging mostly due to its vulnerability to the large variability in head-pose. Prior solutions struggle to maintain reliable accuracy in unconstrained remote gaze tracking. Among them, appearance-based solutions demonstrate tremendous potential in improving gaze accuracy. However, existing works still suffer from head movement and are not robust enough to handle real-world scenarios. Especially most of them study gaze estimation under controlled scenarios where the collected datasets often cover limited ranges of both head-pose and gaze which introduces further bias. In this paper, we propose novel end-to-end appearance-based gaze estimation methods that could more robustly incorporate different levels of head-pose representations into gaze estimation. Our method could generalize to real-world scenarios with low image quality, different lightings and scenarios where direct head-pose information is not available. To better demonstrate the advantage of our methods, we further propose a new benchmark dataset with the most rich distribution of head-gaze combination reflecting real-world scenarios. Extensive evaluations on several public datasets and our own dataset demonstrate that our method consistently outperforms the state-of-the-art by a significant margin.

cs.CV↗

Dually Supervised Feature Pyramid for Object Detection and Segmentation

Feature pyramid architecture has been broadly adopted in object detection and segmentation to deal with multi-scale problem. However, in this paper we show that the capacity of the architecture has not been fully explored due to the inadequate utilization of the supervision information. Such insufficient utilization is caused by the supervision signal degradation in back propagation. Thus inspired, we propose a dually supervised method, named dually supervised FPN (DSFPN), to enhance the supervision signal when training the feature pyramid network (FPN). In particular, DSFPN is constructed by attaching extra prediction (i.e., detection or segmentation) heads to the bottom-up subnet of FPN. Hence, the features can be optimized by the additional heads before being forwarded to subsequent networks. Further, the auxiliary heads can serve as a regularization term to facilitate the model training. In addition, to strengthen the capability of the detection heads in DSFPN for handling two inhomogeneous tasks, i.e., classification and regression, the originally shared hidden feature space is separated by decoupling classification and regression subnets. To demonstrate the generalizability, effectiveness, and efficiency of the proposed method, DSFPN is integrated into four representative detectors (Faster RCNN, Mask RCNN, Cascade RCNN, and Cascade Mask RCNN) and assessed on the MS COCO dataset. Promising precision improvement, state-of-the-art performance, and negligible additional computational cost are demonstrated through extensive experiments. Code will be provided.

cs.CV↗

On the Equivalence of Semidifinite Relaxations for MIMO Detection with General Constellations

The multiple-input multiple-output (MIMO) detection problem is a fundamental problem in modern digital communications. Semidefinite relaxation (SDR) based algorithms are a popular class of approaches to solving the problem because the algorithms have a polynomial-time worst-case complexity and generally can achieve a good detection error rate performance. In spite of the existence of various different SDRs for the MIMO detection problem in the literature, very little is known about the relationship between these SDRs. This paper aims to fill this theoretical gap. In particular, this paper shows that two existing SDRs for the MIMO detection problem, which take quite different forms and are proposed by using different techniques, are equivalent. As a byproduct of the equivalence result, the tightness of one of the above two SDRs under a sufficient condition can be obtained.

cs.IT↗

Tightness of a new and enhanced semidefinite relaxation for MIMO detection

In this paper, we consider a fundamental problem in modern digital communications known as multi-input multi-output (MIMO) detection, which can be formulated as a complex quadratic programming problem subject to unit-modulus and discrete argument constraints. Various semidefinite relaxation (SDR) based algorithms have been proposed to solve the problem in the literature. In this paper, we first show that the conventional SDR is generally not tight for the problem. Then, we propose a new and enhanced SDR and show its tightness under an easily checkable condition, which essentially requires the level of the noise to be below a certain threshold. The above results have answered an open question posed by So in [35]. Numerical simulation results show that our proposed SDR significantly outperforms the conventional SDR in terms of the relaxation gap.

math.OC↗

Model-based Iterative Restoration for Binary Document Image Compression with Dictionary Learning

The inherent noise in the observed (e.g., scanned) binary document image degrades the image quality and harms the compression ratio through breaking the pattern repentance and adding entropy to the document images. In this paper, we design a cost function in Bayesian framework with dictionary learning. Minimizing our cost function produces a restored image which has better quality than that of the observed noisy image, and a dictionary for representing and encoding the image. After the restoration, we use this dictionary (from the same cost function) to encode the restored image following the symbol-dictionary framework by JBIG2 standard with the lossless mode. Experimental results with a variety of document images demonstrate that our method improves the image quality compared with the observed image, and simultaneously improves the compression ratio. For the test images with synthetic noise, our method reduces the number of flipped pixels by 48.2% and improves the compression ratio by 36.36% as compared with the best encoding methods. For the test images with real noise, our method visually improves the image quality, and outperforms the cutting-edge method by 28.27% in terms of the compression ratio.

cs.CV↗

An Efficient Global Algorithm for Single-Group Multicast Beamforming

Consider the single-group multicast beamforming problem, where multiple users receive the same data stream simultaneously from a single transmitter. The problem is NP-hard and all existing algorithms for the problem either find suboptimal approximate or local stationary solutions. In this paper, we propose an efficient branch-and-bound algorithm for the problem that is guaranteed to find its global solution. To the best of our knowledge, our proposed algorithm is the first tailored global algorithm for the single-group multicast beamforming problem. Simulation results show that our proposed algorithm is computationally efficient (albeit its theoretical worst-case iteration complexity is exponential with respect to the number of receivers) and it significantly outperforms a state-of-the-art general-purpose global optimization solver called Baron. Our proposed algorithm provides an important benchmark for performance evaluation of existing algorithms for the same problem. By using it as the benchmark, we show that two state-of-the-art algorithms, semidefinite relaxation algorithm and successive linear approximation algorithm, work well when the problem dimension (i.e., the number of antennas at the transmitter and the number of receivers) is small but their performance deteriorates quickly as the problem dimension increases.

cs.IT↗

Hardness of FeB4: Density functional theory investigation

A recent experimental study reported the successful synthesis of an orthorhombic FeB4 with a high hardness of 62 GPa, which has reignited extensive interests on whether transition metal borides (TRBs) compounds will become superhard materials. However, it is contradicted with some theoretical studies suggesting transition metal boron compounds are unlikely to become superhard materials. Here, we examined structural and electronic properties of FeB4 using density functional theory. The electronic calculations show the good metallicity and covalent FeB bonding. Meanwhile, we extensively investigated stress strain relations of FeB4 under various tensile and shear loading directions. The calculated weakest tensile and shear stresses are 40 GPa and 25 GPa, respectively. Further simulations (e.g. electron localized function and bond length along the weakest loading direction) on FeB4 show the weak Fe-B bonding is responsible for this low hardness. Moreover, these results are consistent with the value of Vickers hardness (11.7 to 32.3 GPa) by employing different empirical hardness models and below the superhardness threshold of 40 GPa. Our current results suggest FeB4 is a hard material and unlikely to become superhard.

cond-mat.mtrl-sci↗

Optimal Perfect Distinguishability between Unitaries and Quantum Operations

We study optimal perfect distinguishability between a unitary and a general quantum operation. In 2-dimensional case we provide a simple sufficient and necessary condition for sequential perfect distinguishability and an analytical formula of optimal query time. We extend the sequential condition to general d-dimensional case. Meanwhile, we provide an upper bound and a lower bound for optimal sequential query time. In the process a new iterative method is given, the most notable innovation of which is its independence to auxiliary systems or entanglement. Following the idea, we further obtain an upper bound and a lower bound of (entanglement-assisted) q-maximal fidelities between a unitary and a quantum operation. Thus by the recursion in [1] an upper bound and a lower bound for optimal general perfect discrimination are achieved. Finally our lower bound result can be extended to the case of arbitrary two quantum operations.

quant-ph↗