Searcharxiv⌕ Search

arXiv subjects

Duo Li

Publications and source records attributed to Duo Li.

34 records · Page 2Linked to original sources

m-RevNet: Deep Reversible Neural Networks with Momentum

In recent years, the connections between deep residual networks and first-order Ordinary Differential Equations (ODEs) have been disclosed. In this work, we further bridge the deep neural architecture design with the second-order ODEs and propose a novel reversible neural network, termed as m-RevNet, that is characterized by inserting momentum update to residual blocks. The reversible property allows us to perform backward pass without access to activation values of the forward pass, greatly relieving the storage burden during training. Furthermore, the theoretical foundation based on second-order ODEs grants m-RevNet with stronger representational power than vanilla residual networks, which potentially explains its performance gains. For certain learning scenarios, we analytically and empirically reveal that our m-RevNet succeeds while standard ResNet fails. Comprehensive experiments on various image classification and semantic segmentation benchmarks demonstrate the superiority of our m-RevNet over ResNet, concerning both memory efficiency and recognition performance.

cs.CV↗

Learning the Superpixel in a Non-iterative and Lifelong Manner

Superpixel is generated by automatically clustering pixels in an image into hundreds of compact partitions, which is widely used to perceive the object contours for its excellent contour adherence. Although some works use the Convolution Neural Network (CNN) to generate high-quality superpixel, we challenge the design principles of these networks, specifically for their dependence on manual labels and excess computation resources, which limits their flexibility compared with the traditional unsupervised segmentation methods. We target at redefining the CNN-based superpixel segmentation as a lifelong clustering task and propose an unsupervised CNN-based method called LNS-Net. The LNS-Net can learn superpixel in a non-iterative and lifelong manner without any manual labels. Specifically, a lightweight feature embedder is proposed for LNS-Net to efficiently generate the cluster-friendly features. With those features, seed nodes can be automatically assigned to cluster pixels in a non-iterative way. Additionally, our LNS-Net can adapt the sequentially lifelong learning by rescaling the gradient of weight based on both channel and spatial context to avoid overfitting. Experiments show that the proposed LNS-Net achieves significantly better performance on three benchmarks with nearly ten times lower complexity compared with other state-of-the-art methods.

cs.CV↗

Involution: Inverting the Inherence of Convolution for Visual Recognition

Convolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatial-agnostic and channel-specific. Instead, we present a novel atomic operation for deep neural networks by inverting the aforementioned design principles of convolution, coined as involution. We additionally demystify the recent popular self-attention operator and subsume it into our involution family as an over-complicated instantiation. The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including ImageNet classification, COCO detection and segmentation, together with Cityscapes segmentation. Our involution-based models improve the performance of convolutional baselines using ResNet-50 by up to 1.6% top-1 accuracy, 2.5% and 2.4% bounding box AP, and 4.7% mean IoU absolutely while compressing the computational cost to 66%, 65%, 72%, and 57% on the above benchmarks, respectively. Code and pre-trained models for all the tasks are available at https://github.com/d-li14/involution.

cs.CV↗

Potential Convolution: Embedding Point Clouds into Potential Fields

Recently, various convolutions based on continuous or discrete kernels for point cloud processing have been widely studied, and achieve impressive performance in many applications, such as shape classification, scene segmentation and so on. However, they still suffer from some drawbacks. For continuous kernels, the inaccurate estimation of the kernel weights constitutes a bottleneck for further improving the performance; while for discrete ones, the kernels represented as the points located in the 3D space are lack of rich geometry information. In this work, rather than defining a continuous or discrete kernel, we directly embed convolutional kernels into the learnable potential fields, giving rise to potential convolution. It is convenient for us to define various potential functions for potential convolution which can generalize well to a wide range of tasks. Specifically, we provide two simple yet effective potential functions via point-wise convolution operations. Comprehensive experiments demonstrate the effectiveness of our method, which achieves superior performance on the popular 3D shape classification and scene segmentation benchmarks compared with other state-of-the-art point convolution methods.

cs.CV↗

PointFlow: Flowing Semantics Through Points for Aerial Image Segmentation

Aerial Image Segmentation is a particular semantic segmentation problem and has several challenging characteristics that general semantic segmentation does not have. There are two critical issues: The one is an extremely foreground-background imbalanced distribution, and the other is multiple small objects along with the complex background. Such problems make the recent dense affinity context modeling perform poorly even compared with baselines due to over-introduced background context. To handle these problems, we propose a point-wise affinity propagation module based on the Feature Pyramid Network (FPN) framework, named PointFlow. Rather than dense affinity learning, a sparse affinity map is generated upon selected points between the adjacent features, which reduces the noise introduced by the background while keeping efficiency. In particular, we design a dual point matcher to select points from the salient area and object boundaries, respectively. Experimental results on three different aerial segmentation datasets suggest that the proposed method is more effective and efficient than state-of-the-art general semantic segmentation methods. Especially, our methods achieve the best speed and accuracy trade-off on three aerial benchmarks. Further experiments on three general semantic segmentation datasets prove the generality of our method. Code will be provided in (https: //github.com/lxtGH/PFSegNets).

cs.CV↗

Resolution Switchable Networks for Runtime Efficient Image Recognition

We propose a general method to train a single convolutional neural network which is capable of switching image resolutions at inference. Thus the running speed can be selected to meet various computational resource limits. Networks trained with the proposed method are named Resolution Switchable Networks (RS-Nets). The basic training framework shares network parameters for handling images which differ in resolution, yet keeps separate batch normalization layers. Though it is parameter-efficient in design, it leads to inconsistent accuracy variations at different resolutions, for which we provide a detailed analysis from the aspect of the train-test recognition discrepancy. A multi-resolution ensemble distillation is further designed, where a teacher is learnt on the fly as a weighted ensemble over resolutions. Thanks to the ensemble and knowledge distillation, RS-Nets enjoy accuracy improvements at a wide range of resolutions compared with individually trained models. Extensive experiments on the ImageNet dataset are provided, and we additionally consider quantization problems. Code and models are available at https://github.com/yikaiw/RS-Nets.

cs.CV↗

A unified first order hyperbolic model for nonlinear dynamic rupture processes in diffuse fracture zones

Earthquake fault zones are more complex, both geometrically and rheologically, than an idealised infinitely thin plane embedded in linear elastic material. To incorporate nonlinear material behaviour, natural complexities, and multi-physics coupling within and outside of fault zones, here we present a first-order hyperbolic and thermodynamically compatible mathematical model for a continuum in a gravitational field which provides a unified description of nonlinear elasto-plasticity, material damage and of viscous Newtonian flows with phase transition between solid and liquid phases. The fault geometry and secondary cracks are described via a scalar function $ξ\in [0,1]$ that indicates the local level of material damage. The model also permits the representation of arbitrarily complex geometries via a diffuse interface approach based on the solid volume fraction function $α\in [0,1]$. Neither of the two scalar fields $ξ$ and $α$ needs to be mesh-aligned, allowing thus faults and cracks with complex topology and the use of adaptive Cartesian meshes (AMR). The model shares common features with phase-field approaches but substantially extends them. We show a wide range of numerical applications that are relevant for dynamic earthquake rupture in fault zones, including the co-seismic generation of secondary off-fault shear cracks, tensile rock fracture in the Brazilian disc test, as well as a natural convection problem in molten rock-like material.

physics.geo-ph↗

Learning to Learn Parameterized Classification Networks for Scalable Input Images

Convolutional Neural Networks (CNNs) do not have a predictable recognition behavior with respect to the input resolution change. This prevents the feasibility of deployment on different input image resolutions for a specific model. To achieve efficient and flexible image classification at runtime, we employ meta learners to generate convolutional weights of main networks for various input scales and maintain privatized Batch Normalization layers per scale. For improved training performance, we further utilize knowledge distillation on the fly over model predictions based on different input resolutions. The learned meta network could dynamically parameterize main networks to act on input images of arbitrary size with consistently better accuracy compared to individually trained models. Extensive experiments on the ImageNet demonstrate that our method achieves an improved accuracy-efficiency trade-off during the adaptive inference process. By switching executable input resolutions, our method could satisfy the requirement of fast adaption in different resource-constrained environments. Code and models are available at https://github.com/d-li14/SAN.

cs.CV↗

PSConv: Squeezing Feature Pyramid into One Compact Poly-Scale Convolutional Layer

Despite their strong modeling capacities, Convolutional Neural Networks (CNNs) are often scale-sensitive. For enhancing the robustness of CNNs to scale variance, multi-scale feature fusion from different layers or filters attracts great attention among existing solutions, while the more granular kernel space is overlooked. We bridge this regret by exploiting multi-scale features in a finer granularity. The proposed convolution operation, named Poly-Scale Convolution (PSConv), mixes up a spectrum of dilation rates and tactfully allocate them in the individual convolutional kernels of each filter regarding a single convolutional layer. Specifically, dilation rates vary cyclically along the axes of input and output channels of the filters, aggregating features over a wide range of scales in a neat style. PSConv could be a drop-in replacement of the vanilla convolution in many prevailing CNN backbones, allowing better representation learning without introducing additional parameters and computational complexities. Comprehensive experiments on the ImageNet and MS COCO benchmarks validate the superior performance of PSConv. Code and models are available at https://github.com/d-li14/PSConv.

cs.CV↗

HBONet: Harmonious Bottleneck on Two Orthogonal Dimensions

MobileNets, a class of top-performing convolutional neural network architectures in terms of accuracy and efficiency trade-off, are increasingly used in many resourceaware vision applications. In this paper, we present Harmonious Bottleneck on two Orthogonal dimensions (HBO), a novel architecture unit, specially tailored to boost the accuracy of extremely lightweight MobileNets at the level of less than 40 MFLOPs. Unlike existing bottleneck designs that mainly focus on exploring the interdependencies among the channels of either groupwise or depthwise convolutional features, our HBO improves bottleneck representation while maintaining similar complexity via jointly encoding the feature interdependencies across both spatial and channel dimensions. It has two reciprocal components, namely spatial contraction-expansion and channel expansion-contraction, nested in a bilaterally symmetric structure. The combination of two interdependent transformations performing on orthogonal dimensions of feature maps enhances the representation and generalization ability of our proposed module, guaranteeing compelling performance with limited computational resource and power. By replacing the original bottlenecks in MobileNetV2 backbone with HBO modules, we construct HBONets which are evaluated on ImageNet classification, PASCAL VOC object detection and Market-1501 person re-identification. Extensive experiments show that with the severe constraint of computational budget our models outperform MobileNetV2 counterparts by remarkable margins of at most 6.6%, 6.3% and 5.0% on the above benchmarks respectively. Code and pretrained models are available at https://github.com/d-li14/HBONet.

cs.CV↗

On projective varieties with strictly nef tangent bundles

In this paper, we study smooth complex projective varieties $X$ such that some exterior power $\bigwedge^r T_X$ of the tangent bundle is strictly nef. We prove that such varieties are rationally connected. We also classify the following two cases. If $T_X$ is strictly nef, then $X$ isomorphic to the projective space $\mathrm{P}^n$. If $\bigwedge^2 T_X$ is strictly nef and if $X$ has dimension at least $3$, then $X$ is either isomorphic to $\mathrm{P}^n$ or a quadric $\mathrm{Q}^n$.

math.AG↗

Characterizations of projective spaces and quadrics by strictly nef bundles

In this paper, we show that if the tangent bundle of a smooth projective variety is strictly nef, then it is isomorphic to a projective space; if a projective variety $X^n$ $(n>4)$ has strictly nef $Λ^2 TX$, then it is isomorphic to $\mathbb{P}^n$ or quadric $\mathbb{Q}^n$. We also prove that on elliptic curves, strictly nef vector bundles are ample, whereas there exist Hermitian flat and strictly nef vector bundles on any smooth curve with genus $g\geq 2$.

math.AG↗

Categorical characterization of quadrics

We give a characterization of smooth quadrics in terms of the existence of full exceptional collections of certain type, which generalizes a result of C.Vial for projective spaces.

math.AG↗

On certain K-equivalent birational maps

We study K-equivalent birational maps which are resolved by a single blowup. Examples of such maps include standard flops and twisted Mukai flops. We give a criterion for such maps to be a standard flop or a twisted Mukai flop. As an application, we classify all such birational maps up to dimension 5.

math.AG↗

Classification of two-dimensional algebraic projective semigroups

In this article, we address the classification of smooth projective algebraic surfaces over complex numbers admitting algebraic semigroup structures. We give a full description of those surfaces $S$, which has at least one non-trivial algebraic semigroup structure, when the Kodaira dimension of $S$ is $ -\infty$ and $ 0$. For the case "$ κ(S)=1$", we give a description of one special type of elliptic surfaces which admit non-trivial algebraic semigroup laws. \\ For a given surface $S$, it is an interesting problem to describe all algebraic semigroup structures on it and determine the dimension of this moduli. In this article, we solve this problem for case "$ κ(S)\ge 0$".

math.AG↗

Enhancement of shot noise due to the fluctuation of Coulomb interaction

We have developed a theoretical formalism to investigate the contribution of fluctuation of Coulomb interaction to the shot noise based on Keldysh non-equilibrium Green's function method. We have applied our theory to study the behavior of dc shot noise of atomic junctions using the method of nonequilibrium Green's function combined with the density functional theory (NEGF-DFT). In particular, for atomic carbon wire consisting 4 carbon atoms in contact with two Al(100) electrodes, first principles calculation within NEGF-DFT formalism shows a negative differential resistance (NDR) region in I-V curve at finite bias due to the effective band bottom of the Al lead. We have calculated the shot noise spectrum using the conventional gauge invariant transport theory with Coulomb interaction considered explicitly on the Hartree level along with exchange and correlation effect. Although the Fano factor is enhanced from 0.6 to 0.8 in the NDR region, the expected super-Poissonian behavior in the NDR regionis not observed. When the fluctuation of Coulomb interaction is included in the shot noise, our numerical results show that the Fano factor is greater than one in the NDR region indicating a super-Poissonian behavior.

cond-mat.mes-hall↗