SearcharxivSearch

arXiv subjects

Xiao Pan

Publications and source records attributed to Xiao Pan.

At least 19 recordsLinked to original sources

Group-wise Supervision with Focal-Dice Loss for Long-Tailed Indoor Semantic Occupancy Prediction

Recently, 3D semantic occupancy prediction has garnered increasing attention for understanding the indoor scene. However, unlike structured outdoor environments, indoor scenes feature a high diversity of object categories that exhibit a severe long-tailed distribution, which has become a core bottleneck limiting the performance of existing models. To tackle this challenge, we propose a novel method, Group-UFD Occ, based on hierarchical semantic supervision and synergistic loss optimization. At the architectural level, we introduce a fine-grained semantic grouping strategy and design multi-scale, parallel ``main-expert'' prediction heads to guide the model in efficiently learning tail-class features through deep regularization. At the optimization level, we introduce the Unified Focal-Dice (UFD) loss. This synergistic loss function dynamically focuses on hard samples at the per-voxel level. Meanwhile, it simultaneously optimizes the geometric integrity of predicted objects from a region-based perspective. We conducted experiments on the large-scale EmbodiedScan dataset. The results demonstrate that our method yields a relative improvement of 11.38\% over the baseline, with substantial accuracy gains in several critical long-tailed categories.

cs.CV

Post-disaster building indoor damage and survivor detection using autonomous path planning and deep learning with unmanned aerial vehicles

Rapid response to natural disasters such as earthquakes is a crucial element in ensuring the safety of civil infrastructures and minimizing casualties. Traditional manual inspection is labour-intensive, time-consuming, and can be dangerous for inspectors and rescue workers. This paper proposed an autonomous inspection approach for structural damage inspection and survivor detection in the post-disaster building indoor scenario, which incorporates an autonomous navigation method, deep learning-based damage and survivor detection method, and a customized low-cost micro aerial vehicle (MAV) with onboard sensors. Experimental studies in a pseudo-post-disaster office building have shown the proposed methodology can achieve high accuracy in structural damage inspection and survivor detection. Overall, the proposed inspection approach shows great potential to improve the efficiency of existing manual post-disaster building inspection.

cs.CV

InsightEdit: Towards Better Instruction Following for Image Editing

In this paper, we focus on the task of instruction-based image editing. Previous works like InstructPix2Pix, InstructDiffusion, and SmartEdit have explored end-to-end editing. However, two limitations still remain: First, existing datasets suffer from low resolution, poor background consistency, and overly simplistic instructions. Second, current approaches mainly condition on the text while the rich image information is underexplored, therefore inferior in complex instruction following and maintaining background consistency. Targeting these issues, we first curated the AdvancedEdit dataset using a novel data construction pipeline, formulating a large-scale dataset with high visual quality, complex instructions, and good background consistency. Then, to further inject the rich image information, we introduce a two-stream bridging mechanism utilizing both the textual and visual features reasoned by the powerful Multimodal Large Language Models (MLLM) to guide the image editing process more precisely. Extensive results demonstrate that our approach, InsightEdit, achieves state-of-the-art performance, excelling in complex instruction following and maintaining high background consistency with the original image.

cs.CV

Research on the Spatial Data Intelligent Foundation Model

This report focuses on spatial data intelligent large models, delving into the principles, methods, and cutting-edge applications of these models. It provides an in-depth discussion on the definition, development history, current status, and trends of spatial data intelligent large models, as well as the challenges they face. The report systematically elucidates the key technologies of spatial data intelligent large models and their applications in urban environments, aerospace remote sensing, geography, transportation, and other scenarios. Additionally, it summarizes the latest application cases of spatial data intelligent large models in themes such as urban development, multimodal systems, remote sensing, smart transportation, and resource environments. Finally, the report concludes with an overview and outlook on the development prospects of spatial data intelligent large models.

cs.AI

GD^2-NeRF: Generative Detail Compensation via GAN and Diffusion for One-shot Generalizable Neural Radiance Fields

In this paper, we focus on the One-shot Novel View Synthesis (O-NVS) task which targets synthesizing photo-realistic novel views given only one reference image per scene. Previous One-shot Generalizable Neural Radiance Fields (OG-NeRF) methods solve this task in an inference-time finetuning-free manner, yet suffer the blurry issue due to the encoder-only architecture that highly relies on the limited reference image. On the other hand, recent diffusion-based image-to-3d methods show vivid plausible results via distilling pre-trained 2D diffusion models into a 3D representation, yet require tedious per-scene optimization. Targeting these issues, we propose the GD$^2$-NeRF, a Generative Detail compensation framework via GAN and Diffusion that is both inference-time finetuning-free and with vivid plausible details. In detail, following a coarse-to-fine strategy, GD$^2$-NeRF is mainly composed of a One-stage Parallel Pipeline (OPP) and a 3D-consistent Detail Enhancer (Diff3DE). At the coarse stage, OPP first efficiently inserts the GAN model into the existing OG-NeRF pipeline for primarily relieving the blurry issue with in-distribution priors captured from the training dataset, achieving a good balance between sharpness (LPIPS, FID) and fidelity (PSNR, SSIM). Then, at the fine stage, Diff3DE further leverages the pre-trained image diffusion models to complement rich out-distribution details while maintaining decent 3D consistency. Extensive experiments on both the synthetic and real-world datasets show that GD$^2$-NeRF noticeably improves the details while without per-scene finetuning.

cs.CV

Vision-aided nonlinear control framework for shake table tests

The structural response under the earthquake excitations can be simulated by scaled-down model shake table tests or full-scale model shake table tests. In this paper, adaptive control theory is used as a nonlinear shake table control algorithm which considers the inherent nonlinearity of the shake table system and the Control-Structural Interaction (CSI) effect that the linear controller cannot consider, such as the Proportional-Integral-Derivative (PID) controller. The mass of the specimen can be assumed as an unknown variation and the unknown parameter will be replaced by an estimated value in the proposed control framework. The signal generated by the control law of the adaptive control method will be implemented by a loop-shaping controller. To verify the stability and feasibility of the proposed control framework, a simulation of a bare shake table and experiments with a bare shake table with a two-story frame were carried out. This study randomly selects Earthquake recordings from the Pacific Earthquake Engineering Research Center (PEER) database. The simulation and experimental results show that the proposed control framework can be effectively used in shake table control.

eess.SY

3D vision-based structural masonry damage detection

The detection of masonry damage is essential for preventing potentially disastrous outcomes. Manual inspection can, however, take a long time and be hazardous to human inspectors. Automation of the inspection process using novel computer vision and machine learning algorithms can be a more efficient and safe solution to prevent further deterioration of the masonry structures. Most existing 2D vision-based methods are limited to qualitative damage classification, 2D localization, and in-plane quantification. In this study, we present a 3D vision-based methodology for accurate masonry damage detection, which offers a more robust solution with a greater field of view, depth of vision, and the ability to detect failures in complex environments. First, images of the masonry specimens are collected to generate a 3D point cloud. Second, 3D point clouds processing methods are developed to evaluate the masonry damage. We demonstrate the effectiveness of our approach through experiments on structural masonry components. Our experiments showed the proposed system can effectively classify damage states and localize and quantify critical damage features. The result showed the proposed method can improve the level of autonomy during the inspection of masonry structures.

cs.CV

Autonomous damage assessment of structural columns using low-cost micro aerial vehicles and multi-view computer vision

Structural columns are the crucial load-carrying components of buildings and bridges. Early detection of column damage is important for the assessment of the residual performance and the prevention of system-level collapse. This research proposes an innovative end-to-end micro aerial vehicles (MAVs)-based approach to automatically scan and inspect columns. First, an MAV-based automatic image collection method is proposed. The MAV is programmed to sense the structural columns and their surrounding environment. During the navigation, the MAV first detects and approaches the structural columns. Then, it starts to collect image data at multiple viewpoints around every detected column. Second, the collected images will be used to assess the damage types and damage locations. Third, the damage state of the structural column will be determined by fusing the evaluation outcomes from multiple camera views. In this study, reinforced concrete (RC) columns are selected to demonstrate the effectiveness of the approach. Experimental results indicate that the proposed MAV-based inspection approach can effectively collect images from multiple viewing angles, and accurately assess critical RC column damages. The approach improves the level of autonomy during the inspection. In addition, the evaluation outcomes are more comprehensive than the existing 2D vision methods. The concept of the proposed inspection approach can be extended to other structural columns such as bridge piers.

cs.CV

A reinforcement learning based construction material supply strategy using robotic crane and computer vision for building reconstruction after an earthquake

After an earthquake, it is particularly important to provide the necessary resources on site because a large number of infrastructures need to be repaired or newly constructed. Due to the complex construction environment after the disaster, there are potential safety hazards for human labors working in this environment. With the advancement of robotic technology and artificial intelligent (AI) algorithms, smart robotic technology is the potential solution to provide construction resources after an earthquake. In this paper, the robotic crane with advanced AI algorithms is proposed to provide resources for infrastructure reconstruction after an earthquake. The proximal policy optimization (PPO), a reinforcement learning (RL) algorithm, is implemented for 3D lift path planning when transporting the construction materials. The state and reward function are designed in detail for RL model training. Two models are trained through a loading task in different environments by using PPO algorithm, one considering the influence of obstacles and the other not considering obstacles. Then, the two trained models are compared and evaluated through an unloading task and a loading task in simulation environments. For each task, two different cases are considered. One is that there is no obstacle between the initial position where the construction material is lifted and the target position, and the other is that there are obstacles between the initial position and the target position. The results show that the model that considering the obstacles during training can generate proper actions for the robotic crane to execute so that the crane can automatically transport the construction materials to the desired location with swing suppression, short time consumption and collision avoidance.

cs.CV

A Real-Time Robust Ecological-Adaptive Cruise Control Strategy for Battery Electric Vehicles

This work addresses the ecological-adaptive cruise control problem for connected electric vehicles by a computationally efficient robust control strategy. The problem is formulated in the space-domain with a realistic description of the nonlinear electric powertrain model and motion dynamics to yield a convex optimal control problem (OCP). The OCP is approached by a novel robust model predictive control (RMPC) method handling various disturbances due to modelling mismatch and inaccurate leading vehicle information. The RMPC problem is solved by semi-definite programming relaxation and single linear matrix inequality (sLMI) techniques for further enhanced computational efficiency. The performance of the proposed real-time robust ecological-adaptive cruise control (REACC) method is evaluated using an experimentally collected driving cycle. Its robustness is verified by comparison with a nominal MPC which is shown to result in speed-limit constraint violations. The energy economy of the proposed method outperforms a state-of-the-art time-domain RMPC scheme, as a more precisely fitted convex powertrain model can be integrated into the space-domain scheme. The additional comparison with a traditional constant distance following strategy (CDFS) further verifies the effectiveness of the proposed REACC. Finally, it is verified that the REACC can be potentially implemented in real-time owing to the sLMI and resulting convex algorithm.

eess.SY

TransHuman: A Transformer-based Human Representation for Generalizable Neural Human Rendering

In this paper, we focus on the task of generalizable neural human rendering which trains conditional Neural Radiance Fields (NeRF) from multi-view videos of different characters. To handle the dynamic human motion, previous methods have primarily used a SparseConvNet (SPC)-based human representation to process the painted SMPL. However, such SPC-based representation i) optimizes under the volatile observation space which leads to the pose-misalignment between training and inference stages, and ii) lacks the global relationships among human parts that is critical for handling the incomplete painted SMPL. Tackling these issues, we present a brand-new framework named TransHuman, which learns the painted SMPL under the canonical space and captures the global relationships between human parts with transformers. Specifically, TransHuman is mainly composed of Transformer-based Human Encoding (TransHE), Deformable Partial Radiance Fields (DPaRF), and Fine-grained Detail Integration (FDI). TransHE first processes the painted SMPL under the canonical space via transformers for capturing the global relationships between human parts. Then, DPaRF binds each output token with a deformable radiance field for encoding the query point under the observation space. Finally, the FDI is employed to further integrate fine-grained information from reference images. Extensive experiments on ZJU-MoCap and H36M show that our TransHuman achieves a significantly new state-of-the-art performance with high efficiency. Project page: https://pansanity666.github.io/TransHuman/

cs.CV

Masked Audio Text Encoders are Effective Multi-Modal Rescorers

Masked Language Models (MLMs) have proven to be effective for second-pass rescoring in Automatic Speech Recognition (ASR) systems. In this work, we propose Masked Audio Text Encoder (MATE), a multi-modal masked language model rescorer which incorporates acoustic representations into the input space of MLM. We adopt contrastive learning for effectively aligning the modalities by learning shared representations. We show that using a multi-modal rescorer is beneficial for domain generalization of the ASR system when target domain data is unavailable. MATE reduces word error rate (WER) by 4%-16% on in-domain, and 3%-7% on out-of-domain datasets, over the text-only baseline. Additionally, with very limited amount of training data (0.8 hours), MATE achieves a WER reduction of 8%-23% over the first-pass baseline.

cs.SD

A Computationally Efficient Robust Model Predictive Control Framework for Ecological Adaptive Cruise Control Strategy of Electric Vehicles

The recent advancement in vehicular networking technology provides novel solutions for designing intelligent and sustainable vehicle motion controllers. This work addresses a car-following task, where the feedback linearisation method is combined with a robust model predictive control (RMPC) scheme to safely, optimally and efficiently control a connected electric vehicle. In particular, the nonlinear dynamics are linearised through a feedback linearisation method to maintain an efficient computational speed and to guarantee global optimality. At the same time, the inevitable model mismatch is dealt with by the RMPC design. The control objective of the RMPC is to optimise the electric energy efficiency of the ego vehicle with consideration of a bounded model mismatch disturbance subject to satisfaction of physical and safety constraints. Numerical results first verify the validity and robustness through a comparison between the proposed RMPC and a nominal MPC. Further investigation into the performance of the proposed method reveals a higher energy efficiency and passenger comfort level as compared to a recently proposed benchmark method using the space-domain modelling approach.

eess.SY

Economic Potential for Hybrid Electric Vehicles in Urban Signal-free Intersections with Decentralized MPC

The development of electric and connected vehicles as well as automated driving technologies are key towards the smart city, with convenient urban mobility and high energy economy performance. However, the global rise in electricity price provokes renewed interest on CAVs with hybrid electric powertrains rather than considering battery electric powertrains. This paper provides a decentralized coordination strategy for a group of connected and autonomous vehicles (CAVs) with a series hybrid electric (sHEV) powertrain at signal-free intersections. The problem is formulated as a convex form with suitable relaxation and approximation of the powertrain model and solved by decentralized model predictive control, which is able to ensure a rapid search and unique solution in real time. Numerical examples validate the effectiveness of the proposed methods concerning physical and safety constraints. By utilizing the petrol fuel and battery charging prices over the last year, the performance of the proposed approach is evaluated against the optimal results produced by two benchmark solutions, conventional vehicles (CVs) and battery electric vehicles (BEVs). The comparison results show that the traveling cost of sHEVs approaches and even under some circumstances reaches the same level as for BEVs, which indicates the importance of hybridization, particularly under the current rising electricity price situation.

eess.SY

Dynamic Gradient Reactivation for Backward Compatible Person Re-identification

We study the backward compatible problem for person re-identification (Re-ID), which aims to constrain the features of an updated new model to be comparable with the existing features from the old model in galleries. Most of the existing works adopt distillation-based methods, which focus on pushing new features to imitate the distribution of the old ones. However, the distillation-based methods are intrinsically sub-optimal since it forces the new feature space to imitate the inferior old feature space. To address this issue, we propose the Ranking-based Backward Compatible Learning (RBCL), which directly optimizes the ranking metric between new features and old features. Different from previous methods, RBCL only pushes the new features to find best-ranking positions in the old feature space instead of strictly alignment, and is in line with the ultimate goal of backward retrieval. However, the sharp sigmoid function used to make the ranking metric differentiable also incurs the gradient vanish issue, therefore stems the ranking refinement during the later period of training. To address this issue, we propose the Dynamic Gradient Reactivation (DGR), which can reactivate the suppressed gradients by adding dynamic computed constant during forward step. To further help targeting the best-ranking positions, we include the Neighbor Context Agents (NCAs) to approximate the entire old feature space during training. Unlike previous works which only test on the in-domain settings, we make the first attempt to introduce the cross-domain settings (including both supervised and unsupervised), which are more meaningful and difficult. The experimental results on all five settings show that the proposed RBCL outperforms previous state-of-the-art methods by large margins under all settings.

cs.CV

A Hierarchical Robust Control Strategy for Decentralized Signal-Free Intersection Management

The development of connected and automated vehicles is the key to improving urban mobility safety and efficiency. This paper focuses on cooperative vehicle management at a signal-free intersection with consideration of vehicle modeling uncertainties and sensor measurement disturbances. The problem is approached by a hierarchical robust control strategy in a decentralized traffic coordination framework where optimal control and tube-based robust model predictive control methods are designed to hierarchically solve the optimal crossing order and the velocity trajectories of a group of CAVs in terms of energy consumption and throughput. To capture the energy consumption of each vehicle, their powertrain system is modeled in line with an electric drive system. With a suitable relaxation and spatial modeling approach, the optimization problems in the proposed strategy can be formulated as convex second-order cone programs, which provide a unique and computationally efficient solution. A rigorous proof of the equivalence between the convexified and the original problems is also provided. Simulation results illustrate the effectiveness and robustness of the proposed strategy and reveal the impact of traffic density on the control solution. The study of the Pareto optimal solutions for the energy-time objective shows that a minor reduction in journey time can considerably reduce energy consumption, which emphasizes the necessity of optimizing their trade-off. Finally, the numerical comparisons carried out for different prediction horizons and sampling intervals provide insight into the control design.

math.OC

A Convex Optimal Control Framework for Autonomous Vehicle Intersection Crossing

Cooperative vehicle management emerges as a promising solution to improve road traffic safety and efficiency. This paper addresses the speed planning problem for connected and autonomous vehicles (CAVs) at an unsignalized intersection with consideration of turning maneuvers. The problem is approached by a hierarchical centralized coordination scheme that successively optimizes the crossing order and velocity trajectories of a group of vehicles so as to minimize their total energy consumption and travel time required to pass the intersection. For an accurate estimate of the energy consumption of each CAV, the vehicle modeling framework in this paper captures 1) friction losses that affect longitudinal vehicle dynamics, and 2) the powertrain of each CAV in line with a battery-electric architecture. It is shown that the underlying optimization problem subject to safety constraints for powertrain operation, cornering and collision avoidance, after convexification and relaxation in some aspects can be formulated as two second-order cone programs, which ensures a rapid solution search and a unique global optimum. Simulation case studies are provided showing the tightness of the convex relaxation bounds, the overall effectiveness of the proposed approach, and its advantages over a benchmark solution invoking the widely used first-in-first-out policy. The investigation of Pareto optimal solutions for the two objectives (travel time and energy consumption) highlights the importance of optimizing their trade-off, as small compromises in travel time could produce significant energy savings.

eess.SY

Optimal Energy Management of Series Hybrid Electric Vehicles with Engine Start-Stop System

This paper develops energy management (EM) control for series hybrid electric vehicles (HEVs) that include an engine start-stop system (SSS). The objective of the control is to optimally split the energy between the sources of the powertrain and achieve fuel consumption minimization. In contrast to existing works, a fuel penalty is used to characterize more realistically SSS engine restarts, to enable more realistic design and testing of control algorithms. The paper first derives two important analytic results: a) analytic EM optimal solutions of fundamental and commonly used series HEV frameworks, and b) proof of optimality of charge sustaining operation in series HEVs. It then proposes a novel heuristic control strategy, the hysteresis power threshold strategy (HPTS), by amalgamating simple and effective control rules extracted from the suite of derived analytic EM optimal solutions. The decision parameters of the control strategy are small in number and freely tunable. The overall control performance can be fully optimized for different HEV parameters and driving cycles by a systematic tuning process, while also targeting charge sustaining operation. The performance of HPTS is evaluated and benchmarked against existing methodologies, including dynamic programming (DP) and a recently proposed state-of-the-art heuristic strategy. The results show the effectiveness and robustness of the HPTS and also indicate its potential to be used as the benchmark strategy for high fidelity HEV models, where DP is no longer applicable due to computational complexity.

eess.SY