Searcharxiv⌕ Search

arXiv subjects

Lin Wang

Publications and source records attributed to Lin Wang.

At least 307 records · Page 17Linked to original sources

FMapping: Factorized Efficient Neural Field Mapping for Real-Time Dense RGB SLAM

In this paper, we introduce FMapping, an efficient neural field mapping framework that facilitates the continuous estimation of a colorized point cloud map in real-time dense RGB SLAM. To achieve this challenging goal without depth, a hurdle is how to improve efficiency and reduce the mapping uncertainty of the RGB SLAM system. To this end, we first build up a theoretical analysis by decomposing the SLAM system into tracking and mapping parts, and the mapping uncertainty is explicitly defined within the frame of neural representations. Based on the analysis, we then propose an effective factorization scheme for scene representation and introduce a sliding window strategy to reduce the uncertainty for scene reconstruction. Specifically, we leverage the factorized neural field to decompose uncertainty into a lower-dimensional space, which enhances robustness to noise and improves training efficiency. We then propose the sliding window sampler to reduce uncertainty by incorporating coherent geometric cues from observed frames during map initialization to enhance convergence. Our factorized neural mapping approach enjoys some advantages, such as low memory consumption, more efficient computation, and fast convergence during map initialization. Experiments on two benchmark datasets show that our method can update the map of high-fidelity colorized point clouds around 2 seconds in real time while requiring no customized CUDA kernels. Additionally, it utilizes x20 fewer parameters than the most concise neural implicit mapping of prior methods for SLAM, e.g., iMAP [ 31] and around x1000 fewer parameters than the state-of-the-art approach, e.g., NICE-SLAM [ 42]. For more details, please refer to our project homepage: https://vlis2022.github.io/fmap/.

cs.CV↗

Towards Language-guided Interactive 3D Generation: LLMs as Layout Interpreter with Generative Feedback

Generating and editing a 3D scene guided by natural language poses a challenge, primarily due to the complexity of specifying the positional relations and volumetric changes within the 3D space. Recent advancements in Large Language Models (LLMs) have demonstrated impressive reasoning, conversational, and zero-shot generation abilities across various domains. Surprisingly, these models also show great potential in realizing and interpreting the 3D space. In light of this, we propose a novel language-guided interactive 3D generation system, dubbed LI3D, that integrates LLMs as a 3D layout interpreter into the off-the-shelf layout-to-3D generative models, allowing users to flexibly and interactively generate visual content. Specifically, we design a versatile layout structure base on the bounding boxes and semantics to prompt the LLMs to model the spatial generation and reasoning from language. Our system also incorporates LLaVA, a large language and vision assistant, to provide generative feedback from the visual aspect for improving the visual quality of generated content. We validate the effectiveness of LI3D, primarily in 3D generation and editing through multi-round interactions, which can be flexibly extended to 2D generation and editing. Various experiments demonstrate the potential benefits of incorporating LLMs in generative AI for applications, e.g., metaverse. Moreover, we benchmark the layout reasoning performance of LLMs with neural visual artist tasks, revealing their emergent ability in the spatial layout domain.

cs.CV↗

HRDFuse: Monocular 360°Depth Estimation by Collaboratively Learning Holistic-with-Regional Depth Distributions

Depth estimation from a monocular 360° image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, \eg, OmniFusion, have applied the tangent projection (TP) to represent a 360°image and predicted depth values via patch-wise regressions, which are merged to get a depth map with equirectangular projection (ERP) format. However, these methods suffer from 1) non-trivial process of merging plenty of patches; 2) capturing less holistic-with-regional contextual information by directly regressing the depth value of each pixel. In this paper, we propose a novel framework, \textbf{HRDFuse}, that subtly combines the potential of convolutional neural networks (CNNs) and transformers by collaboratively learning the \textit{holistic} contextual information from the ERP and the \textit{regional} structural information from the TP. Firstly, we propose a spatial feature alignment (\textbf{SFA}) module that learns feature similarities between the TP and ERP to aggregate the TP features into a complete ERP feature map in a pixel-wise manner. Secondly, we propose a collaborative depth distribution classification (\textbf{CDDC}) module that learns the \textbf{holistic-with-regional} histograms capturing the ERP and TP depth distributions. As such, the final depth values can be predicted as a linear combination of histogram bin centers. Lastly, we adaptively combine the depth predictions from ERP and TP to obtain the final depth map. Extensive experiments show that our method predicts\textbf{ more smooth and accurate depth} results while achieving \textbf{favorably better} results than the SOTA methods.

cs.CV↗

SPP-CNN: An Efficient Framework for Network Robustness Prediction

This paper addresses the robustness of a network to sustain its connectivity and controllability against malicious attacks. This kind of network robustness is typically measured by the time-consuming attack simulation, which returns a sequence of values that record the remaining connectivity and controllability after a sequence of node- or edge-removal attacks. For improvement, this paper develops an efficient framework for network robustness prediction, the spatial pyramid pooling convolutional neural network (SPP-CNN). The new framework installs a spatial pyramid pooling layer between the convolutional and fully-connected layers, overcoming the common mismatch issue in the CNN-based prediction approaches and extending its generalizability. Extensive experiments are carried out by comparing SPP-CNN with three state-of-the-art robustness predictors, namely a CNN-based and two graph neural networks-based frameworks. Synthetic and real-world networks, both directed and undirected, are investigated. Experimental results demonstrate that the proposed SPP-CNN achieves better prediction performances and better generalizability to unknown datasets, with significantly lower time-consumption, than its counterparts.

cs.LG↗

Deep Reinforcement Learning Based Resource Allocation for Cloud Native Wireless Network

Cloud native technology has revolutionized 5G beyond and 6G communication networks, offering unprecedented levels of operational automation, flexibility, and adaptability. However, the vast array of cloud native services and applications presents a new challenge in resource allocation for dynamic cloud computing environments. To tackle this challenge, we investigate a cloud native wireless architecture that employs container-based virtualization to enable flexible service deployment. We then study two representative use cases: network slicing and Multi-Access Edge Computing. To optimize resource allocation in these scenarios, we leverage deep reinforcement learning techniques and introduce two model-free algorithms capable of monitoring the network state and dynamically training allocation policies. We validate the effectiveness of our algorithms in a testbed developed using Free5gc. Our findings demonstrate significant improvements in network efficiency, underscoring the potential of our proposed techniques in unlocking the full potential of cloud native wireless networks.

cs.NI↗

A nonlinear semigroup approach to Hamilton-Jacobi equations--revisited

We consider the Hamilton-Jacobi equation \[{H}(x,Du)+λ(x)u=c,\quad x\in M, \] where $M$ is a connected, closed and smooth Riemannian manifold. The functions ${H}(x,p)$ and $λ(x)$ are continuous. ${H}(x,p)$ is convex, coercive with respect to $p$, and $λ(x)$ changes the signs. The first breakthrough to this model was achieved by Jin-Yan-Zhao \cite{JYZ} under the Tonelli conditions. In this paper, we consider more detailed structure of the viscosity solution set and large time behavior of the viscosity solution on the Cauchy problem.

math.AP↗

Learning Spatial-Temporal Implicit Neural Representations for Event-Guided Video Super-Resolution

Event cameras sense the intensity changes asynchronously and produce event streams with high dynamic range and low latency. This has inspired research endeavors utilizing events to guide the challenging video superresolution (VSR) task. In this paper, we make the first attempt to address a novel problem of achieving VSR at random scales by taking advantages of the high temporal resolution property of events. This is hampered by the difficulties of representing the spatial-temporal information of events when guiding VSR. To this end, we propose a novel framework that incorporates the spatial-temporal interpolation of events to VSR in a unified framework. Our key idea is to learn implicit neural representations from queried spatial-temporal coordinates and features from both RGB frames and events. Our method contains three parts. Specifically, the Spatial-Temporal Fusion (STF) module first learns the 3D features from events and RGB frames. Then, the Temporal Filter (TF) module unlocks more explicit motion information from the events near the queried timestamp and generates the 2D features. Lastly, the SpatialTemporal Implicit Representation (STIR) module recovers the SR frame in arbitrary resolutions from the outputs of these two modules. In addition, we collect a real-world dataset with spatially aligned events and RGB frames. Extensive experiments show that our method significantly surpasses the prior-arts and achieves VSR with random scales, e.g., 6.5. Code and dataset are available at https: //vlis2022.github.io/cvpr23/egvsr.

cs.CV↗

Minimizing the fluctuation of resonance driving terms in dynamic aperture optimization

Dynamic aperture (DA) is an important nonlinear property of a storage ring lattice, which has a dominant effect on beam injection efficiency and beam lifetime. Generally, minimizing both resonance driving terms (RDTs) and amplitude dependent tune shifts is an essential condition for enlarging the DA. In this paper, we study the correlation between the fluctuation of RDTs along the ring and the DA area with double- and multi-bend achromat lattices. It is found that minimizing the RDT fluctuations is more effective than minimizing RDTs themselves in enlarging the DA, and thus can serve as a very powerful indicator in the DA optimization. Besides, it is found that minimizing lower-order RDT fluctuations can also reduce higher-order RDTs, which are not only more computationally complicated but also more numerous. The effectiveness of controlling the RDT fluctuations in enlarging the DA confirms that the local cancellation of nonlinear effects used in some diffraction-limited storage ring lattices is more effective than the global cancellation.

physics.acc-ph↗

Bloch-Lorentz magnetoresistance oscillations in delafossites

Recent measurements of the out-of-plane magnetoresistance of delafossites (PdCoO$_2$ and PtCoO$_2$) observed oscillations closely resembling the Aharonov-Bohm effect. Here, we show that the magnetoresistance oscillations are explained by the Bloch-like oscillations of the out-of-plane electron trajectories. We develop a semiclassical theory of these Bloch-Lorentz oscillations and show that they are a consequence of the ballistic motion and quasi-2D dispersion of delafossites. Our model identifies the sample wall scattering to be the most likely factor limiting the visibility of these Bloch-Lorentz oscillations in existing experiments.

cond-mat.mes-hall↗

Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic Segmentation

The ability of scene understanding has sparked active research for panoramic image semantic segmentation. However, the performance is hampered by distortion of the equirectangular projection (ERP) and a lack of pixel-wise annotations. For this reason, some works treat the ERP and pinhole images equally and transfer knowledge from the pinhole to ERP images via unsupervised domain adaptation (UDA). However, they fail to handle the domain gaps caused by: 1) the inherent differences between camera sensors and captured scenes; 2) the distinct image formats (e.g., ERP and pinhole images). In this paper, we propose a novel yet flexible dual-path UDA framework, DPPASS, taking ERP and tangent projection (TP) images as inputs. To reduce the domain gaps, we propose cross-projection and intra-projection training. The cross-projection training includes tangent-wise feature contrastive training and prediction consistency training. That is, the former formulates the features with the same projection locations as positive examples and vice versa, for the models' awareness of distortion, while the latter ensures the consistency of cross-model predictions between the ERP and TP. Moreover, adversarial intra-projection training is proposed to reduce the inherent gap, between the features of the pinhole images and those of the ERP and TP images, respectively. Importantly, the TP path can be freely removed after training, leading to no additional inference cost. Extensive experiments on two benchmarks show that our DPPASS achieves +1.06$\%$ mIoU increment than the state-of-the-art approaches.

cs.CV↗

Hierarchical Knowledge Guided Learning for Real-world Retinal Diseases Recognition

In the real world, medical datasets often exhibit a long-tailed data distribution (i.e., a few classes occupy the majority of the data, while most classes have only a limited number of samples), which results in a challenging long-tailed learning scenario. Some recently published datasets in ophthalmology AI consist of more than 40 kinds of retinal diseases with complex abnormalities and variable morbidity. Nevertheless, more than 30 conditions are rarely seen in global patient cohorts. From a modeling perspective, most deep learning models trained on these datasets may lack the ability to generalize to rare diseases where only a few available samples are presented for training. In addition, there may be more than one disease for the presence of the retina, resulting in a challenging label co-occurrence scenario, also known as \textit{multi-label}, which can cause problems when some re-sampling strategies are applied during training. To address the above two major challenges, this paper presents a novel method that enables the deep neural network to learn from a long-tailed fundus database for various retinal disease recognition. Firstly, we exploit the prior knowledge in ophthalmology to improve the feature representation using a hierarchy-aware pre-training. Secondly, we adopt an instance-wise class-balanced sampling strategy to address the label co-occurrence issue under the long-tailed medical dataset scenario. Thirdly, we introduce a novel hybrid knowledge distillation to train a less biased representation and classifier. We conducted extensive experiments on four databases, including two public datasets and two in-house databases with more than one million fundus images. The experimental results demonstrate the superiority of our proposed methods with recognition accuracy outperforming the state-of-the-art competitors, especially for these rare diseases.

cs.CV↗

Vetaverse: A Survey on the Intersection of Metaverse, Vehicles, and Transportation Systems

Since 2021, the term "Metaverse" has been the most popular one, garnering a lot of interest. Because of its contained environment and built-in computing and networking capabilities, a modern car makes an intriguing location to host its own little metaverse. Additionally, the travellers don't have much to do to pass the time while traveling, making them ideal customers for immersive services. Vetaverse (Vehicular-Metaverse), which we define as the future continuum between vehicular industries and Metaverse, is envisioned as a blended immersive realm that scales up to cities and countries, as digital twins of the intelligent Transportation Systems, referred to as "TS-Metaverse", as well as customized XR services inside each Individual Vehicle, referred to as "IV-Metaverse". The two subcategories serve fundamentally different purposes, namely long-term interconnection, maintenance, monitoring, and management on scale for large transportation systems (TS), and personalized, private, and immersive infotainment services (IV). By outlining the framework of Vetaverse and examining important enabler technologies, we reveal this impending trend. Additionally, we examine unresolved issues and potential routes for future study while highlighting some intriguing Vetaverse services.

cs.HC↗

A representation formula of the viscosity solution of the contact Hamilton-Jacobi equation and its applications

Assume $M$ is a closed, connected and smooth Riemannian manifold. We consider the evolutionary Hamilton-Jacobi equation \begin{equation*} \left\{ \begin{aligned} &\partial_t u(x,t)+H(x,u(x,t),\partial_xu(x,t))=0,\quad (x,t)\in M\times(0,+\infty), \\ &u(x,0)=φ(x), \end{aligned} \right. \end{equation*} where $φ\in C(M)$ and the stationary one \begin{equation*} H(x,u(x),\partial_x u(x))=0, \end{equation*} where $H(x,u,p)$ is continuous, convex and coercive in $p$, uniformly Lipschitz in $u$. By introducing a solution semigroup, we provide a representation formula of the viscosity solution of the evolutionary equation. As its applications, we obtain a necessary and sufficient condition for the existence of the viscosity solutions of the stationary equations. Moreover, we prove a new comparison theorem depending on the neighborhood of the projected Aubry set essentially, which is different from the one for the Hamilton-Jacobi equation independent of $u$.

math.AP↗

Topological atomic spinwave lattices by dissipative couplings

Recent experimental advance in creating dissipative couplings provides a new route for engineering exotic lattice systems and exploring topological dissipation. Using the spatial lattice of atomic spinwaves in a vacuum vapor cell, where purely dissipative couplings arise from diffusion of atoms, we experimentally realize a dissipative version of the Su-Schrieffer-Heeger (SSH) model. We construct the dissipation spectrum of the topological or trivial lattices via electromagnetically-induced-transparency (EIT) spectroscopy. The topological dissipation spectrum is found to exhibit edge modes within a dissipative gap. We validate chiral symmetry of the dissipative SSH couplings, and also probe topological features of the generalized dissipative SSH model. This work paves the way for realizing non-Hermitian topological quantum optics via dissipative couplings.

quant-ph↗

Controllability of Networked Sampled-data Systems

The controllability of networked sampled-data systems with zero-order holders on the control and transmission channels is explored, where single- and multi-rate sampling patterns are considered, respectively. The effects of sampling on the controllability of networked systems are analyzed, with some necessary and/or sufficient controllability conditions derived. Different from the sampling control of single systems, the pathological sampling of node systems could be eliminated by an appropriate design of network structure and inner couplings. While for singular topology matrices, the pathological sampling of single nodes will cause the entire system to lose controllability. Moreover, any periodic sampling will not affect the controllability of networked systems with specific node dynamics. All the results indicate that whether a networked system is under pathological sampling or not is jointly determined by mutually coupled factors.

eess.SY↗

Independence-Encouraging Subsampling for Nonparametric Additive Models

The additive model is a popular nonparametric regression method due to its ability to retain modeling flexibility while avoiding the curse of dimensionality. The backfitting algorithm is an intuitive and widely used numerical approach for fitting additive models. However, its application to large datasets may incur a high computational cost and is thus infeasible in practice. To address this problem, we propose a novel approach called independence-encouraging subsampling (IES) to select a subsample from big data for training additive models. Inspired by the minimax optimality of an orthogonal array (OA) due to its pairwise independent predictors and uniform coverage for the range of each predictor, the IES approach selects a subsample that approximates an OA to achieve the minimax optimality. Our asymptotic analyses demonstrate that an IES subsample converges to an OA and that the backfitting algorithm over the subsample converges to a unique solution even if the predictors are highly dependent in the original big data. The proposed IES method is also shown to be numerically appealing via simulations and a real data application.

stat.ME↗

A Fast Bootstrap Algorithm for Causal Inference with Large Data

Estimating causal effects from large experimental and observational data has become increasingly prevalent in both industry and research. The bootstrap is an intuitive and powerful technique used to construct standard errors and confidence intervals of estimators. Its application however can be prohibitively demanding in settings involving large data. In addition, modern causal inference estimators based on machine learning and optimization techniques exacerbate the computational burden of the bootstrap. The bag of little bootstraps has been proposed in non-causal settings for large data but has not yet been applied to evaluate the properties of estimators of causal effects. In this paper, we introduce a new bootstrap algorithm called causal bag of little bootstraps for causal inference with large data. The new algorithm significantly improves the computational efficiency of the traditional bootstrap while providing consistent estimates and desirable confidence interval coverage. We describe its properties, provide practical considerations, and evaluate the performance of the proposed algorithm in terms of bias, coverage of the true 95% confidence intervals, and computational time in a simulation study. We apply it in the evaluation of the effect of hormone therapy on the average time to coronary heart disease using a large observational data set from the Women's Health Initiative.

stat.ME↗