Searcharxiv⌕ Search

arXiv subjects

Libo Wang

Publications and source records attributed to Libo Wang.

28 records · Page 2Linked to original sources

Transformer Meets Convolution: A Bilateral Awareness Network for Semantic Segmentation of Very Fine Resolution Urban Scene Images

Semantic segmentation from very fine resolution (VFR) urban scene images plays a significant role in several application scenarios including autonomous driving, land cover classification, and urban planning, etc. However, the tremendous details contained in the VFR image, especially the considerable variations in scale and appearance of objects, severely limit the potential of the existing deep learning approaches. Addressing such issues represents a promising research field in the remote sensing community, which paves the way for scene-level landscape pattern analysis and decision making. In this paper, we propose a Bilateral Awareness Network which contains a dependency path and a texture path to fully capture the long-range relationships and fine-grained details in VFR images. Specifically, the dependency path is conducted based on the ResT, a novel Transformer backbone with memory-efficient multi-head self-attention, while the texture path is built on the stacked convolution operation. Besides, using the linear attention mechanism, a feature aggregation module is designed to effectively fuse the dependency features and texture features. Extensive experiments conducted on the three large-scale urban scene image segmentation datasets, i.e., ISPRS Vaihingen dataset, ISPRS Potsdam dataset, and UAVid dataset, demonstrate the effectiveness of our BANet. Specifically, a 64.6% mIoU is achieved on the UAVid dataset. Code is available at https://github.com/WangLibo1995/GeoSeg.

cs.CV↗

A Novel Transformer Based Semantic Segmentation Scheme for Fine-Resolution Remote Sensing Images

The fully convolutional network (FCN) with an encoder-decoder architecture has been the standard paradigm for semantic segmentation. The encoder-decoder architecture utilizes an encoder to capture multilevel feature maps, which are incorporated into the final prediction by a decoder. As the context is crucial for precise segmentation, tremendous effort has been made to extract such information in an intelligent fashion, including employing dilated/atrous convolutions or inserting attention modules. However, these endeavors are all based on the FCN architecture with ResNet or other backbones, which cannot fully exploit the context from the theoretical concept. By contrast, we introduce the Swin Transformer as the backbone to extract the context information and design a novel decoder of densely connected feature aggregation module (DCFAM) to restore the resolution and produce the segmentation map. The experimental results on two remotely sensed semantic segmentation datasets demonstrate the effectiveness of the proposed scheme.Code is available at https://github.com/WangLibo1995/GeoSeg

cs.CV↗

Building extraction with vision transformer

As an important carrier of human productive activities, the extraction of buildings is not only essential for urban dynamic monitoring but also necessary for suburban construction inspection. Nowadays, accurate building extraction from remote sensing images remains a challenge due to the complex background and diverse appearances of buildings. The convolutional neural network (CNN) based building extraction methods, although increased the accuracy significantly, are criticized for their inability for modelling global dependencies. Thus, this paper applies the Vision Transformer for building extraction. However, the actual utilization of the Vision Transformer often comes with two limitations. First, the Vision Transformer requires more GPU memory and computational costs compared to CNNs. This limitation is further magnified when encountering large-sized inputs like fine-resolution remote sensing images. Second, spatial details are not sufficiently preserved during the feature extraction of the Vision Transformer, resulting in the inability for fine-grained building segmentation. To handle these issues, we propose a novel Vision Transformer (BuildFormer), with a dual-path structure. Specifically, we design a spatial-detailed context path to encode rich spatial details and a global context path to capture global dependencies. Besides, we develop a window-based linear multi-head self-attention to make the complexity of the multi-head self-attention linear with the window size, which strengthens the global context extraction by using large windows and greatly improves the potential of the Vision Transformer in processing large-sized remote sensing images. The proposed method yields state-of-the-art performance (75.74% IoU) on the Massachusetts building dataset. Code will be available.

cs.CV↗

Scale-aware Neural Network for Semantic Segmentation of Multi-resolution Remote Sensing Images

Assigning geospatial objects with specific categories at the pixel level is a fundamental task in remote sensing image analysis. Along with rapid development in sensor technologies, remotely sensed images can be captured at multiple spatial resolutions (MSR) with information content manifested at different scales. Extracting information from these MSR images represents huge opportunities for enhanced feature representation and characterisation. However, MSR images suffer from two critical issues: 1) increased scale variation of geo-objects and 2) loss of detailed information at coarse spatial resolutions. To bridge these gaps, in this paper, we propose a novel scale-aware neural network (SaNet) for semantic segmentation of MSR remotely sensed imagery. SaNet deploys a densely connected feature network (DCFFM) module to capture high-quality multi-scale context, such that the scale variation is handled properly and the quality of segmentation is increased for both large and small objects. A spatial feature recalibration (SFRM) module is further incorporated into the network to learn intact semantic content with enhanced spatial relationships, where the negative effects of information loss are removed. The combination of DCFFM and SFRM allows SaNet to learn scale-aware feature representation, which outperforms the existing multi-scale feature representation. Extensive experiments on three semantic segmentation datasets demonstrated the effectiveness of the proposed SaNet in cross-resolution segmentation.

cs.CV↗

Multi-Attention-Network for Semantic Segmentation of Fine Resolution Remote Sensing Images

Semantic segmentation of remote sensing images plays an important role in a wide range of applications including land resource management, biosphere monitoring and urban planning. Although the accuracy of semantic segmentation in remote sensing images has been increased significantly by deep convolutional neural networks, several limitations exist in standard models. First, for encoder-decoder architectures such as U-Net, the utilization of multi-scale features causes the underuse of information, where low-level features and high-level features are concatenated directly without any refinement. Second, long-range dependencies of feature maps are insufficiently explored, resulting in sub-optimal feature representations associated with each semantic class. Third, even though the dot-product attention mechanism has been introduced and utilized in semantic segmentation to model long-range dependencies, the large time and space demands of attention impede the actual usage of attention in application scenarios with large-scale input. This paper proposed a Multi-Attention-Network (MANet) to address these issues by extracting contextual dependencies through multiple efficient attention modules. A novel attention mechanism of kernel attention with linear complexity is proposed to alleviate the large computational demand in attention. Based on kernel attention and channel attention, we integrate local feature maps extracted by ResNeXt-101 with their corresponding global dependencies and reweight interdependent channel maps adaptively. Numerical experiments on three large-scale fine resolution remote sensing images captured by different satellite sensors demonstrate the superior performance of the proposed MANet, outperforming the DeepLab V3+, PSPNet, FastFCN, DANet, OCRNet, and other benchmark approaches.

eess.IV↗

A general construction of permutation polynomials of the form $ (x^{2^m}+x+δ)^{i(2^m-1)+1}+x$ over $\F_{2^{2m}}$

Recently, there has been a lot of work on constructions of permutation polynomials of the form $(x^{2^m}+x+δ)^{s}+x$ over the finite field $\F_{2^{2m}}$, especially in the case when $s$ is of the form $s=i(2^m-1)+1$ (Niho exponent). In this paper, we further investigate permutation polynomials with this form. Instead of seeking for sporadic constructions of the parameter $i$, we give a general sufficient condition on $i$ such that $(x^{2^m}+x+δ)^{i(2^m-1)+1}+x$ permutes $\F_{2^{2m}}$, that is, $(2^k+1)i \equiv 1 ~\textrm{or}~ 2^k~(\textrm{mod}~ 2^m+1)$, where $1 \leq k \leq m-1$ is any integer. This generalizes a recent result obtained by Gupta and Sharma who actually dealt with the case $k=2$. It turns out that most of previous constructions of the parameter $i$ are covered by our result, and it yields many new classes of permutation polynomials as well.

cs.IT↗

Bayesian Variable Selection for Skewed Heteroscedastic Response

In this article, we propose new Bayesian methods for selecting and estimating a sparse coefficient vector for skewed heteroscedastic response. Our novel Bayesian procedures effectively estimate the median and other quantile functions, accommodate non-local prior for regression effects without compromising ease of implementation via sampling based tools, and asymptotically select the true set of predictors even when the number of covariates increases in the same order of the sample size. We also extend our method to deal with some observations with very large errors. Via simulation studies and a re-analysis of a medical cost study with large number of potential predictors, we illustrate the ease of implementation and other practical advantages of our approach compared to existing methods for such studies.

stat.ME↗

More characterizations of generalized bent function in odd characteristic, their dual and the gray image

In this paper, we further investigate properties of generalized bent Boolean functions from $\Z_{p}^n$ to $\Z_{p^k}$, where $p$ is an odd prime and $k$ is a positive integer. For various kinds of representations, sufficient and necessary conditions for bent-ness of such functions are given in terms of their various kinds of component functions. Furthermore, a subclass of gbent functions corresponding to relative difference sets, which we call $\Z_{p^k}$-bent functions, are studied. It turns out that $\Z_{p^k}$-bent functions correspond to a class of vectorial bent functions, and the property of being $\Z_{p^k}$-bent is much stronger then the standard bent-ness. The dual and the generalized Gray image of gbent function are also discussed. In addition, as a further generalization, we also define and give characterizations of gbent functions from $\Z_{p^l}^n$ to $\Z_{p^k}$ for a positive integer $l$ with $l<k$.

math.NT↗

$\mathbb{Z}_q$-valued generalized bent functions in odd characteristics

In this paper, we investigate properties of functions from $\mathbb{Z}_{p}^n$ to $\mathbb{Z}_q$, where $p$ is an odd prime and $q$ is a positive integer divided by $p$. we present the sufficient and necessary conditions for bent-ness of such generalized Boolean functions in terms of classical $p$-ary bent functions, when $q=p^k$. When $q$ is divided by $p$ but not a power of it, we give an sufficient condition for weakly regular gbent functions. Some related constructions are also obtained.

math.NT↗