SearcharxivSearch

arXiv subjects

Xia Yan

Publications and source records attributed to Xia Yan.

10 recordsLinked to original sources

OutLangSplat: 3D Language Gaussian Splatting for UAV Outdoor Scenes

3D Language Gaussian Splatting embeds open-vocabulary language features into 3D Gaussian Splatting, providing an efficient explicit representation for text-driven 3D scene understanding. However, existing methods are limited to indoor or small-scale scenes, and tend to fail in Unmanned Aerial Vehicle (UAV) outdoor scenes, where severe occlusions and long distance viewpoints often lead to incorrect semantic activations and missing target responses. In this paper, we present OutLangSplat which adapts language Gaussian representations to UAV outdoor scenes by improving feature representation and aggregation reliability. For the feature representation, a 2D-3D dual-branch representation with region-based alignment and fusion is designed to improve spatial consistency, reducing incomplete target responses and background misactivations. For the feature aggregation, we introduce a training-free contribution and consistency-aware Gaussian feature aggregation strategy that leverages pixel contribution reliability and cross-view semantic consistency to suppress unreliable responses from noisy viewpoints. A new dataset is provided by manually annotating various objects on four real-world public UAV outdoor scene datasets. To the best of our knowledge, it is the first accessible dataset of open-vocabulary 3D scene understanding for UAV outdoor scenes. Quantitative evaluations and ablation studies demonstrate that OutLangSplat outperforms SOTA methods on both open-vocabulary semantic segmentation and instance localization tasks. The datasets and codes will be open-sourced.

cs.CV

TIBR4D: Tracing-Guided Iterative Boundary Refinement for Efficient 4D Gaussian Segmentation

Object-level segmentation in dynamic 4D Gaussian scenes remains challenging due to complex motion, occlusions, and ambiguous boundaries. In this paper, we present an efficient learning-free 4D Gaussian segmentation framework that lifts video segmentation masks to 4D spaces, whose core is a two-stage iterative boundary refinement, TIBR4D. The first stage is an Iterative Gaussian Instance Tracing (IGIT) at the temporal segment level. It progressively refines Gaussian-to-instance probabilities through iterative tracing, and extracts corresponding Gaussian point clouds that better handle occlusions and preserve completeness of object structures compared to existing one-shot threshold-based methods. The second stage is a frame-wise Gaussian Rendering Range Control (RCC) via suppressing highly uncertain Gaussians near object boundaries while retaining their core contributions for more accurate boundaries. Furthermore, a temporal segmentation merging strategy is proposed for IGIT to balance identity consistency and dynamic awareness. Longer segments enforce stronger multi-frame constraints for stable identities, while shorter segments allow identity changes to be captured promptly. Experiments on HyperNeRF and Neu3D demonstrate that our method produces accurate object Gaussian point clouds with clearer boundaries and higher efficiency compared to SOTA methods.

cs.CV

Results of the 2024 CommonRoad Motion Planning Competition for Autonomous Vehicles

Over the past decade, a wide range of motion planning approaches for autonomous vehicles has been developed to handle increasingly complex traffic scenarios. However, these approaches are rarely compared on standardized benchmarks, limiting the assessment of relative strengths and weaknesses. To address this gap, we present the setup and results of the 4th CommonRoad Motion Planning Competition held in 2024, conducted using the CommonRoad benchmark suite. This annual competition provides an open-source and reproducible framework for benchmarking motion planning algorithms. The benchmark scenarios span highway and urban environments with diverse traffic participants, including passenger cars, buses, and bicycles. Planner performance is evaluated along four dimensions: efficiency, safety, comfort, and compliance with selected traffic rules. This report introduces the competition format and provides a comparison of representative high-performing planners from the 2023 and 2024 editions.

cs.RO

A Discrete Neural Operator with Adaptive Sampling for Surrogate Modeling of Parametric Transient Darcy Flows in Porous Media

This study proposes a new discrete neural operator for surrogate modeling of transient Darcy flow fields in heterogeneous porous media with random parameters. The new method integrates temporal encoding, operator learning and UNet to approximate the mapping between vector spaces of random parameter and spatiotemporal flow fields. The new discrete neural operator can achieve higher prediction accuracy than the SOTA attention-residual-UNet structure. Derived from the finite volume method, the transmissibility matrices rather than permeability is adopted as the inputs of surrogates to enhance the prediction accuracy further. To increase sampling efficiency, a generative latent space adaptive sampling method is developed employing the Gaussian mixture model for density estimation of generalization error. Validation is conducted on test cases of 2D/3D single- and two-phase Darcy flow field prediction. Results reveal consistent enhancement in prediction accuracy given limited training set.

math.NA

Nonlinear Expectation Inference for Efficient Uncertainty Quantification and History Matching of Transient Darcy Flows in Porous Media with Random Parameters Under Distribution Uncertainty

The uncertainty quantification of Darcy flows using history matching is important for the evaluation and prediction of subsurface reservoir performance. Conventional methods aim to obtain the maximum a posterior or maximum likelihood estimate (MLE) using gradient-based, heuristic or ensemble-based methods. These methods can be computationally expensive for high-dimensional problems since forward simulation needs to be run iteratively as physical parameters are updated. In the current study, we propose a nonlinear expectation inference (NEI) method for efficient history matching and uncertainty quantification accounting for distribution or Knightian uncertainty. Forward simulation runs are conducted on prior realisations once, and then a range of expectations are computed in the data space based on subsets of prior realisations with no repetitive forward runs required. In NEI, no prior probability distribution for data is assumed. Instead, the probability distribution is assumed to be uncertain with prior and posterior uncertainty quantified by nonlinear expectations. The inferred result of NEI is the posterior subsets on which the expected flow rates are consistent with observation. The accuracy and efficiency of the new method are validated using single- and two-phase Darcy flows in 2D and 3D heterogeneous reservoirs.

math.NA

Gas flow and solid deformation in unconventional shale

Shale, a material that is currently at the heart of energy resource development, plays a critical role in the management of civil infrastructures. Whether it concerns geothermal energy, carbon sequestration, hydraulic fracturing, or waste storage, one is likely to encounter shale as it accounts for approximately 75\% of rocks in sedimentary basins. Despite the abundance of experimental data indicating the mechanical anisotropy of these formations, past research has often simplified the modeling process by assuming isotropy. In this study, the anisotropic elasticity model and the advanced anisotropic elastoplasticity model proposed by Semnani et al. (2016) and Zhao et al. (2018) were adopted in traditional gas production and strip footing problems, respectively. This was done to underscore the unique characteristics of unconventional shale. The first application example reveals the effects of bedding on apparent permeability and stress evolutions. In the second application example, we contrast the hydromechanical responses with a comparable case where gas is substituted by incompressible fluid. These novel findings enhance our comprehension of gas flow and solid deformation in shale.

physics.geo-ph

A Multi-Stencil Fast Marching Method with Path Correction for Efficient Reservoir Simulation and Automated History Matching

The efficiency of reservoir simulation is important for automated history matching (AHM) and production optimization, etc. The fast marching marching method (FMM) has been used for efficient reservoir simulation. FMM can be regarded as a generalised streamline method but without the need to construct streamlines. In FMM reservoir simulation, the Eikonal equation for the diffusive time-of-flight (DTOF) is solved by FMM and then the governing equations are computed on the 1D DTOF coordinate. Standard FMM solves the Eikonal equation using a 4-stencil algorithm on 2D Cartesian grids, ignoring the diagonal neighbouring cells. In the current study, we build a 8-stencil algorithm considering all neighbouring cells, and use local analytical propagation speeds. In addition, a local path-correction coefficient is introduced to further increase the accuracy of DTOF solution. Next, a discretisation scheme is built on the 1D DTOF coordinate for efficient reservoir simulation. The algorithm is validated on homogeneous and heterogeneous test cases, and its potential for efficient forward simulation in AHM is demonstrated by two examples of dimensions 2 and 6.

physics.flu-dyn

Incorporating VAD into ASR System by Multi-task Learning

When we use End-to-end automatic speech recognition (E2E-ASR) system for real-world applications, a voice activity detection (VAD) system is usually needed to improve the performance and to reduce the computational cost by discarding non-speech parts in the audio. Usually ASR and VAD systems are trained and utilized independently to each other. In this paper, we present a novel multi-task learning (MTL) framework that incorporates VAD into the ASR system. The proposed system learns ASR and VAD jointly in the training stage. With the assistance of VAD, the ASR performance improves as its connectionist temporal classification (CTC) loss function can leverage the VAD alignment information. In the inference stage, the proposed system removes non-speech parts at low computational cost and recognizes speech parts with high robustness. Experimental results on segmented speech data show that by utilizing VAD information, the proposed method outperforms the baseline ASR system on both English and Chinese datasets. On unsegmented speech data, we find that the system outperforms the ASR systems that build an extra GMM-based or DNN-based voice activity detector.

eess.AS

Fluid flow through anisotropic and deformable double porosity media with ultra-low matrix permeability: A continuum framework

Fractured porous media or double porosity media are common in nature. At the same time, accurate modeling remains a significant challenge due to bi-modal pore size distribution, anisotropy, multi-field coupling, and various flow patterns. This study aims to formulate a comprehensive coupled continuum framework that could adequately consider these critical characteristics. In our framework, fluid flow in the micro-fracture network is modeled with the generalized Darcy's law, in which the equivalent fracture permeability is upscaled from the detailed geological characterizations. The liquid in the much less permeable matrix follows a low-velocity non-Darcy flow characterized by threshold values and non-linearity. The fluid mass transfer is assumed to be a function of the shape factor, pressure difference, and (variable) interface permeability. The solid deformation relies on a thermodynamically consistent effective stress derived from the energy balance equation, and it is modeled following anisotropic poroelastic theory. The discussion revolves around generic double porosity media. Model applications reveal the capability of our framework to capture the crucial roles of coupling, poroelastic coefficients, anisotropy, and ultra-low matrix permeability in dictating the pressure and displacement fields.

physics.geo-ph

Understanding the Disharmony between Weight Normalization Family and Weight Decay: $ε-$shifted $L_2$ Regularizer

The merits of fast convergence and potentially better performance of the weight normalization family have drawn increasing attention in recent years. These methods use standardization or normalization that changes the weight $\boldsymbol{W}$ to $\boldsymbol{W}'$, which makes $\boldsymbol{W}'$ independent to the magnitude of $\boldsymbol{W}$. Surprisingly, $\boldsymbol{W}$ must be decayed during gradient descent, otherwise we will observe a severe under-fitting problem, which is very counter-intuitive since weight decay is widely known to prevent deep networks from over-fitting. In this paper, we \emph{theoretically} prove that the weight decay term $\frac{1}{2}λ||{\boldsymbol{W}}||^2$ merely modulates the effective learning rate for improving objective optimization, and has no influence on generalization when the weight normalization family is compositely employed. Furthermore, we also expose several critical problems when introducing weight decay term to weight normalization family, including the missing of global minimum and training instability. To address these problems, we propose an $ε-$shifted $L_2$ regularizer, which shifts the $L_2$ objective by a positive constant $ε$. Such a simple operation can theoretically guarantee the existence of global minimum, while preventing the network weights from being too small and thus avoiding gradient float overflow. It significantly improves the training stability and can achieve slightly better performance in our practice. The effectiveness of $ε-$shifted $L_2$ regularizer is comprehensively validated on the ImageNet, CIFAR-100, and COCO datasets. Our codes and pretrained models will be released in https://github.com/implus/PytorchInsight.

cs.LG