Searcharxiv⌕ Search

arXiv subjects

Guoping Wang

Publications and source records attributed to Guoping Wang.

54 records · Page 3Linked to original sources

CreatureShop: Interactive 3D Character Modeling and Texturing from a Single Color Drawing

Creating 3D shapes from 2D drawings is an important problem with applications in content creation for computer animation and virtual reality. We introduce a new sketch-based system, CreatureShop, that enables amateurs to create high-quality textured 3D character models from 2D drawings with ease and efficiency. CreatureShop takes an input bitmap drawing of a character (such as an animal or other creature), depicted from an arbitrary descriptive pose and viewpoint, and creates a 3D shape with plausible geometric details and textures from a small number of user annotations on the 2D drawing. Our key contributions are a novel oblique view modeling method, a set of systematic approaches for producing plausible textures on the invisible or occluded parts of the 3D character (as viewed from the direction of the input drawing), and a user-friendly interactive system. We validate our system and methods by creating numerous 3D characters from various drawings, and compare our results with related works to show the advantages of our method. We perform a user study to evaluate the usability of our system, which demonstrates that our system is a practical and efficient approach to create fully-textured 3D character models for novice users.

cs.GR↗

NeuralSound: Learning-based Modal Sound Synthesis With Acoustic Transfer

We present a novel learning-based modal sound synthesis approach that includes a mixed vibration solver for modal analysis and an end-to-end sound radiation network for acoustic transfer. Our mixed vibration solver consists of a 3D sparse convolution network and a Locally Optimal Block Preconditioned Conjugate Gradient module (LOBPCG) for iterative optimization. Moreover, we highlight the correlation between a standard modal vibration solver and our network architecture. Our radiation network predicts the Far-Field Acoustic Transfer maps (FFAT Maps) from the surface vibration of the object. The overall running time of our learning method for any new object is less than one second on a GTX 3080 Ti GPU while maintaining a high sound quality close to the ground truth that is computed using standard numerical methods. We also evaluate the numerical accuracy and perceptual accuracy of our sound synthesis approach on different objects corresponding to various materials.

cs.SD↗

6D-ViT: Category-Level 6D Object Pose Estimation via Transformer-based Instance Representation Learning

This paper presents 6D-ViT, a transformer-based instance representation learning network, which is suitable for highly accurate category-level object pose estimation on RGB-D images. Specifically, a novel two-stream encoder-decoder framework is dedicated to exploring complex and powerful instance representations from RGB images, point clouds and categorical shape priors. For this purpose, the whole framework consists of two main branches, named Pixelformer and Pointformer. The Pixelformer contains a pyramid transformer encoder with an all-MLP decoder to extract pixelwise appearance representations from RGB images, while the Pointformer relies on a cascaded transformer encoder and an all-MLP decoder to acquire the pointwise geometric characteristics from point clouds. Then, dense instance representations (i.e., correspondence matrix, deformation field) are obtained from a multi-source aggregation network with shape priors, appearance and geometric information as input. Finally, the instance 6D pose is computed by leveraging the correspondence among dense representations, shape priors, and the instance point clouds. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed 3D instance representation learning framework achieves state-of-the-art performance on both datasets, and significantly outperforms all existing methods.

cs.CV↗

AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo Network

In this paper, we present a novel recurrent multi-view stereo network based on long short-term memory (LSTM) with adaptive aggregation, namely AA-RMVSNet. We firstly introduce an intra-view aggregation module to adaptively extract image features by using context-aware convolution and multi-scale aggregation, which efficiently improves the performance on challenging regions, such as thin objects and large low-textured surfaces. To overcome the difficulty of varying occlusion in complex scenes, we propose an inter-view cost volume aggregation module for adaptive pixel-wise view aggregation, which is able to preserve better-matched pairs among all views. The two proposed adaptive aggregation modules are lightweight, effective and complementary regarding improving the accuracy and completeness of 3D reconstruction. Instead of conventional 3D CNNs, we utilize a hybrid network with recurrent structure for cost volume regularization, which allows high-resolution reconstruction and finer hypothetical plane sweep. The proposed network is trained end-to-end and achieves excellent performance on various datasets. It ranks $1^{st}$ among all submissions on Tanks and Temples benchmark and achieves competitive results on DTU dataset, which exhibits strong generalizability and robustness. Implementation of our method is available at https://github.com/QT-Zhu/AA-RMVSNet.

cs.CV↗

Deep Learning for Multi-View Stereo via Plane Sweep: A Survey

3D reconstruction has lately attracted increasing attention due to its wide application in many areas, such as autonomous driving, robotics and virtual reality. As a dominant technique in artificial intelligence, deep learning has been successfully adopted to solve various computer vision problems. However, deep learning for 3D reconstruction is still at its infancy due to its unique challenges and varying pipelines. To stimulate future research, this paper presents a review of recent progress in deep learning methods for Multi-view Stereo (MVS), which is considered as a crucial task of image-based 3D reconstruction. It also presents comparative results on several publicly available datasets, with insightful observations and inspiring future research directions.

cs.CV↗

The distance spectrum of the complements of graphs of diameter greater than three

Suppose that $G$ is a connected simple graph with the vertex set $V( G ) = \{ v_1,v_2,\cdots ,v_n \} $. Let $d( v_i,v_j ) $ be the distance between $v_i$ and $v_j$. Then the distance matrix of $G$ is $D( G ) =( d_{ij} )_{n\times n}$, where $d_{ij}=d( v_i,v_j ) $. Since $D( G )$ is a non-negative real symmetric matrix, its eigenvalues can be arranged $λ_1(G)\ge λ_2(G)\ge \cdots \ge λ_n(G)$, where eigenvalues $λ_1(G)$ and $λ_n(G)$ are called the distance spectral radius and the least distance eigenvalue of $G$, respectively. The {\it diameter} of graph $G$ is the farthest distance between all pairs of vertices. In this paper, we determine the unique graph whose distance spectral radius attains maximum and minimum among all complements of graphs of diameter greater than three, respectively. Furthermore, we also characterize the unique graph whose least distance eigenvalue attains maximum and minimum among all complements of graphs of diameter greater than three, respectively.

math.CO↗

AADS: Augmented Autonomous Driving Simulation using Data-driven Algorithms

Simulation systems have become an essential component in the development and validation of autonomous driving technologies. The prevailing state-of-the-art approach for simulation is to use game engines or high-fidelity computer graphics (CG) models to create driving scenarios. However, creating CG models and vehicle movements (e.g., the assets for simulation) remains a manual task that can be costly and time-consuming. In addition, the fidelity of CG images still lacks the richness and authenticity of real-world images and using these images for training leads to degraded performance. In this paper we present a novel approach to address these issues: Augmented Autonomous Driving Simulation (AADS). Our formulation augments real-world pictures with a simulated traffic flow to create photo-realistic simulation images and renderings. More specifically, we use LiDAR and cameras to scan street scenes. From the acquired trajectory data, we generate highly plausible traffic flows for cars and pedestrians and compose them into the background. The composite images can be re-synthesized with different viewpoints and sensor models. The resulting images are photo-realistic, fully annotated, and ready for end-to-end training and testing of autonomous driving systems from perception to planning. We explain our system design and validate our algorithms with a number of autonomous driving tasks from detection to segmentation and predictions. Compared to traditional approaches, our method offers unmatched scalability and realism. Scalability is particularly important for AD simulation and we believe the complexity and diversity of the real world cannot be realistically captured in a virtual environment. Our augmented approach combines the flexibility in a virtual environment (e.g., vehicle movements) with the richness of the real world to allow effective simulation of anywhere on earth.

cs.CV↗

Pyramid Multi-view Stereo Net with Self-adaptive View Aggregation

n this paper, we propose an effective and efficient pyramid multi-view stereo (MVS) net with self-adaptive view aggregation for accurate and complete dense point cloud reconstruction. Different from using mean square variance to generate cost volume in previous deep-learning based MVS methods, our \textbf{VA-MVSNet} incorporates the cost variances in different views with small extra memory consumption by introducing two novel self-adaptive view aggregations: pixel-wise view aggregation and voxel-wise view aggregation. To further boost the robustness and completeness of 3D point cloud reconstruction, we extend VA-MVSNet with pyramid multi-scale images input as \textbf{PVA-MVSNet}, where multi-metric constraints are leveraged to aggregate the reliable depth estimation at the coarser scale to fill in the mismatched regions at the finer scale. Experimental results show that our approach establishes a new state-of-the-art on the \textsl{\textbf{DTU}} dataset with significant improvements in the completeness and overall quality, and has strong generalization by achieving a comparable performance as the state-of-the-art methods on the \textsl{\textbf{Tanks and Temples}} benchmark. Our codebase is at \hyperlink{https://github.com/yhw-yhw/PVAMVSNet}{https://github.com/yhw-yhw/PVAMVSNet}

cs.CV↗

Dense Hybrid Recurrent Multi-view Stereo Net with Dynamic Consistency Checking

In this paper, we propose an efficient and effective dense hybrid recurrent multi-view stereo net with dynamic consistency checking, namely $D^{2}$HC-RMVSNet, for accurate dense point cloud reconstruction. Our novel hybrid recurrent multi-view stereo net consists of two core modules: 1) a light DRENet (Dense Reception Expanded) module to extract dense feature maps of original size with multi-scale context information, 2) a HU-LSTM (Hybrid U-LSTM) to regularize 3D matching volume into predicted depth map, which efficiently aggregates different scale information by coupling LSTM and U-Net architecture. To further improve the accuracy and completeness of reconstructed point clouds, we leverage a dynamic consistency checking strategy instead of prefixed parameters and strategies widely adopted in existing methods for dense point cloud reconstruction. In doing so, we dynamically aggregate geometric consistency matching error among all the views. Our method ranks \textbf{$1^{st}$} on the complex outdoor \textsl{Tanks and Temples} benchmark over all the methods. Extensive experiments on the in-door DTU dataset show our method exhibits competitive performance to the state-of-the-art method while dramatically reduces memory consumption, which costs only $19.4\%$ of R-MVSNet memory consumption. The codebase is available at \hyperlink{https://github.com/yhw-yhw/D2HC-RMVSNet}{https://github.com/yhw-yhw/D2HC-RMVSNet}.

cs.CV↗

Graph-Based Parallel Large Scale Structure from Motion

While Structure from Motion (SfM) achieves great success in 3D reconstruction, it still meets challenges on large scale scenes. In this work, large scale SfM is deemed as a graph problem, and we tackle it in a divide-and-conquer manner. Firstly, the images clustering algorithm divides images into clusters with strong connectivity, leading to robust local reconstructions. Then followed with an image expansion step, the connection and completeness of scenes are enhanced by expanding along with a maximum spanning tree. After local reconstructions, we construct a minimum spanning tree (MinST) to find accurate similarity transformations. Then the MinST is transformed into a Minimum Height Tree (MHT) to find a proper anchor node and is further utilized to prevent error accumulation. When evaluated on different kinds of datasets, our approach shows superiority over the state-of-the-art in accuracy and efficiency. Our algorithm is open-sourced at https://github.com/AIBluefisher/GraphSfM.

cs.CV↗

SegVoxelNet: Exploring Semantic Context and Depth-aware Features for 3D Vehicle Detection from Point Cloud

3D vehicle detection based on point cloud is a challenging task in real-world applications such as autonomous driving. Despite significant progress has been made, we observe two aspects to be further improved. First, the semantic context information in LiDAR is seldom explored in previous works, which may help identify ambiguous vehicles. Second, the distribution of point cloud on vehicles varies continuously with increasing depths, which may not be well modeled by a single model. In this work, we propose a unified model SegVoxelNet to address the above two problems. A semantic context encoder is proposed to leverage the free-of-charge semantic segmentation masks in the bird's eye view. Suspicious regions could be highlighted while noisy regions are suppressed by this module. To better deal with vehicles at different depths, a novel depth-aware head is designed to explicitly model the distribution differences and each part of the depth-aware head is made to focus on its own target detection range. Extensive experiments on the KITTI dataset show that the proposed method outperforms the state-of-the-art alternatives in both accuracy and efficiency with point cloud as input only.

cs.CV↗

A Variational Staggered Particle Framework for Incompressible Free-Surface Flows

Smoothed particle hydrodynamics (SPH) has been extensively studied in computer graphics to animate fluids with versatile effects. However, SPH still suffers from two numerical difficulties: the particle deficiency problem, which will deteriorate the simulation accuracy, and the particle clumping problem, which usually leads to poor stability of particle simulations. We propose to solve these two problems by developing an approximate projection method for incompressible free-surface flows under a variational staggered particle framework. After particle discretization, we first categorize all fluid particles into four subsets. Then according to the classification, we propose to solve the particle deficiency problem by analytically imposing free surface boundary conditions on both the Laplacian operator and the source term. To address the particle clumping problem, we propose to extend the Taylor-series consistent pressure gradient model with kernel function correction and semi-analytical boundary conditions. Compared to previous approximate projection method [1], our incompressibility solver is stable under both compressive and tensile stress states, no pressure clumping or iterative density correction (e.g., a density constrained pressure approach) is necessary to stabilize the solver anymore. Motivated by the Helmholtz free energy functional, we additionally introduce an iterative particle shifting algorithm to improve the accuracy. It significantly reduces particle splashes near the free surface. Therefore, high-fidelity simulations of the formation and fragmentation of liquid jets and sheets are obtained for both the two-jets and milk-crown examples.

cs.GR↗

Bundle Adjustment Revisited

3D reconstruction has been developing all these two decades, from moderate to medium size and to large scale. It's well known that bundle adjustment plays an important role in 3D reconstruction, mainly in Structure from Motion(SfM) and Simultaneously Localization and Mapping(SLAM). While bundle adjustment optimizes camera parameters and 3D points as a non-negligible final step, it suffers from memory and efficiency requirements in very large scale reconstruction. In this paper, we study the development of bundle adjustment elaborately in both conventional and distributed approaches. The detailed derivation and pseudo code are also given in this paper.

cs.CV↗

The effect of a graft transformation on distance signless Laplacian spectral radius of the graphs

Suppose that the vertex set of a connected graph $G$ is $V(G)=\{v_1,\cdots,v_n\}$. Then we denote by $Tr_{G}(v_i)$ the sum of distances between $v_i$ and all other vertices of $G$. Let $Tr(G)$ be the $n\times n$ diagonal matrix with its $(i,i)$-entry equal to $Tr_{G}(v_{i})$ and $D(G)$ be the distance matrix of $G$. Then $Q_{D}(G)=Tr(G)+D(G)$ is the distance signless Laplacian matrix of $G$. The largest eigenvalues of $Q_D(G)$ is called distance signless Laplacian spectral radius of $G$. In this paper we give some graft transformations on distance signless Laplacian spectral radius of the graphs and use them to characterize the graphs with the minimum and maximal distance signless Laplacian spectral radius among non-starlike and non-caterpillar trees.

math.CO↗

The least signless Laplacian eigenvalue of the complements of bicyclic graphs

Suppose that $G$ is a connected simple graph with the vertex set $V(G)=\{v_1, v_2,\cdots,v_n\}$. Then the adjacency matrix of $G$ is $A(G)=(a_{ij})_{n\times n}$, where $a_{ij}=1$ if $v_i$ is adjacent to $v_j$, and otherwise $a_{ij}=0$. The degree matrix $D(G)=diag(d_{G}(v_1), d_{G}(v_2), \dots, d_{G}(v_n)),$ where $d_{G}(v_i)$ denotes the degree of $v_i$ in the graph $G$ ($1\leq i\leq n$). The matrix $Q(G)=D(G)+A(G)$ is called the signless Laplacian matrix of $G$. The least eigenvalue of $Q(G)$ is also called the least signless Laplacian eigenvalue of $G$. In this paper we give two graft transformations and then use them to characterize the unique connected graph whose least signless Laplacian eigenvalue is minimum among the complements of all bicyclic graphs.

math.CO↗

Some graft transformations and their applications on distance (signless) Laplacian spectra of graphs

Suppose that the vertex set of a connected graph $G$ is $V(G)=\{v_1,\cdots,v_n\}$. Then we denote by $Tr_{G}(v_i)$ the sum of distances between $v_i$ and all other vertices of $G$. Let $Tr(G)$ be the $n\times n$ diagonal matrix with its $(i,i)$-entry equal to $Tr_{G}(v_{i})$ and $D(G)$ be the distance matrix of $G$. Then $Q_{D}(G)=Tr(G)+D(G)$ and $L_{D}(G)=Tr(G)-D(G)$ are respectively the distance signless Laplacian matrix and distance Laplacian matrix of $G$. The largest eigenvalues $ρ_{Q}(G)$ and $ρ_{L}(G)$ of $Q_D(G)$ and $L_D(G)$ are respectively called distance signless Laplacian spectral radius and distance Laplacian spectral radius of $G$. In this paper we give some graft transformations and use them to characterize the tree $T$ such that $ρ_{Q}(T)$ and $ρ_{L}(T)$ attain the maximum among all trees of order $n$ with given number of pendant vertices.

math.CO↗

X-GANs: Image Reconstruction Made Easy for Extreme Cases

Image reconstruction including image restoration and denoising is a challenging problem in the field of image computing. We present a new method, called X-GANs, for reconstruction of arbitrary corrupted resource based on a variant of conditional generative adversarial networks (conditional GANs). In our method, a novel generator and multi-scale discriminators are proposed, as well as the combined adversarial losses, which integrate a VGG perceptual loss, an adversarial perceptual loss, and an elaborate corresponding point loss together based on the analysis of image feature. Our conditional GANs have enabled a variety of applications in image reconstruction, including image denoising, image restoration from quite a sparse sampling, image inpainting, image recovery from the severely polluted block or even color-noise dominated images, which are extreme cases and haven't been addressed in the status quo. We have significantly improved the accuracy and quality of image reconstruction. Extensive perceptual experiments on datasets ranging from human faces to natural scenes demonstrate that images reconstructed by the presented approach are considerably more realistic than alternative work. Our method can also be extended to handle high-ratio image compression.

cs.CV↗

Improved magnetogram calibration of SMFT and its comparison with the HMI

In this paper, we try to improve the magnetogram calibration method of the Solar Magnetic Field Telescope (SMFT). The improved calibration process fits the observed full Stokes information, using six points on the profile of Fe ı 5324.18 Å line, and the analytical Stokes profiles under the Milne-Eddington atmosphere model, adopting the Levenberg-Marquardt least-square fitting algorithm. In Comparison with the linear calibration methods, which employs one point, there is large difference in the strength of longitudinal field $B_l$ and tranverse field $B_t$, caused by the non-linear relationship, but the discrepancy is little in the case of inclination and azimuth. We conclude that it is better to deal with the non-linear effects in the calibration of $B_l$ and $B_t$ using six points. Moreover, in comparison with SDO/HMI, SMFT has larger stray light and acquires less magnetic field strength. For vector magnetic fields in two sunspot regions, the magnetic field strength, inclination and azimuth angles between SMFT and HMI are roughly in agrement, with the linear fitted slope of 0.73/0.7, 0.95/1.04 and 0.99/1.1. In the case of pores and quiet regions ($B_l$ $<$ 600 G), the fitted slopes of the longitudinal magnetic field strength are 0.78 and 0.87 respectively.

astro-ph.SR↗