SearcharxivSearch

arXiv subjects

Ehsan Karimi

Publications and source records attributed to Ehsan Karimi.

9 recordsLinked to original sources

OPTNet: Ordering Point Transformer Network for Post-disaster 3D Semantic Segmentation

Post-disaster damage assessment requires rapid and accurate semantic segmentation of 3D point clouds to identify critical infrastructure such as damaged buildings and roads. Early Point Transformers (e.g., PTv1, PTv2) relied on computationally expensive neighbor searching (k-NN) and Farthest Point Sampling (FPS). To improve efficiency, recent architectures like Point Transformer V3 (PTv3) adopted static serialization methods, such as Hilbert curves or Z-order, to organize unstructured points for window-based attention. However, these fixed orderings are not optimal for capturing the complex geometry of disaster scenes. In this paper, we propose OPTNet (Ordering Point Transformer Network), which introduces a learnable Point Sorter module. OPTNet utilizes a self-supervised ordering loss to dynamically predict an optimal permutation that maximizes the locality of the attention mechanism. We evaluate our method on the 3DAeroRelief dataset, significantly outperforming state-of-the-art baselines.

cs.LG

Instruct-ICL: Instruction-Guided In-Context Learning for Post-Disaster Damage Assessment

Rapid and accurate situational awareness is essential for effective response during natural disasters, where delays in analysis can significantly hinder decision-making. Training task-specific models for post-disaster assessment is often time-consuming and computationally expensive, making such approaches impractical in time-critical scenarios. Consequently, pretrained multimodal large language models (MLLMs) have emerged as a promising alternative for post-disaster visual question answering (VQA), a task that aims to answer structured questions about visual scenes by jointly reasoning over images and text. While these models demonstrate strong multimodal reasoning capabilities, their responses can be sensitive to prompt formulation, which can limit their reliability in real-world disaster assessment scenarios. In this paper, we investigate whether structured reasoning strategies can improve the reliability of pretrained MLLMs for post-disaster VQA. Specifically, we explore multiple prompting paradigms in which one MLLM is used to generate task-specific instructions that serve as Chain-of-Thought (CoT) guidance for a second MLLM. These instructions are incorporated during answer generation with varying degrees of in-context learning (ICL), enabling the model to leverage both explicit reasoning guidance and contextual examples. We conduct our evaluation on the FloodNet dataset and compare these approaches against a zero-shot baseline. Our results demonstrate that integrating instruction-driven CoT reasoning consistently improves answer accuracy.

cs.CV

Geometric Flood Depth Estimation: Fusing Transformer-Based Segmentation with Digital Elevation Models

Post-disaster situational awareness relies heavily on understanding both the extent and the volume of floodwaters. While 2D semantic segmentation provides accurate flood masking, it lacks the vertical dimension required to assess navigability and structural risk. This paper presents a geometric "Water Surface Elevation" approach for estimating flood depth from monocular aerial imagery. Our pipeline utilizes Mask2Former, a state-of-the-art transformer-based segmentation model, to generate precise 2D flood masks. These masks are fused with Digital Elevation Models (DEMs) to identify the water-land boundary, calculate a global water surface elevation ($Z_{water}$), and compute per-pixel depth based on the principle of local hydrostatic equilibrium. We evaluate this workflow using the FloodNet and CRASAR-U-DROIDS datasets, demonstrating how high-performance segmentation can be leveraged to extract 3D volumetric data from 2D imagery without the latency of hydrodynamic simulations.

cs.CV

Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment

Timely and accurate assessment of damages following natural disasters is essential for effective emergency response and recovery. Recent AI-based frameworks have been developed to analyze large volumes of aerial imagery collected by Unmanned Aerial Vehicles, providing actionable insights rapidly. However, creating and annotating data for training these models is costly and time-consuming, resulting in datasets that are limited in size and diversity. Furthermore, most existing approaches rely on traditional classification-based frameworks with fixed answer spaces, restricting their ability to provide new information without additional data collection or model retraining. Using pre-trained generative models built on in-context learning (ICL) allows for flexible and open-ended answer spaces. However, these models often generate hallucinated outputs or produce generic responses that lack domain-specific relevance. To address these limitations, we propose ThiFAN-VQA, a two-stage reasoning-based framework for visual question answering (VQA) in disaster scenarios. ThiFAN-VQA first generates structured reasoning traces using chain-of-thought (CoT) prompting and ICL to enable interpretable reasoning under limited supervision. A subsequent answer selection module evaluates the generated responses and assigns the most coherent and contextually accurate answer, effectively improve the model performance. By integrating a custom information retrieval system, domain-specific prompting, and reasoning-guided answer selection, ThiFAN-VQA bridges the gap between zero-shot and supervised methods, combining flexibility with consistency. Experiments on FloodNet and RescueNet-VQA, UAV-based datasets from flood- and hurricane-affected regions, demonstrate that ThiFAN-VQA achieves superior accuracy, interpretability, and adaptability for real-world post-disaster damage assessment tasks.

cs.CV

3DAeroRelief: The first 3D Benchmark UAV Dataset for Post-Disaster Assessment

Timely assessment of structural damage is critical for disaster response and recovery. However, most prior work in natural disaster analysis relies on 2D imagery, which lacks depth, suffers from occlusions, and provides limited spatial context. 3D semantic segmentation offers a richer alternative, but existing 3D benchmarks focus mainly on urban or indoor scenes, with little attention to disaster-affected areas. To address this gap, we present 3DAeroRelief--the first 3D benchmark dataset specifically designed for post-disaster assessment. Collected using low-cost unmanned aerial vehicles (UAVs) over hurricane-damaged regions, the dataset features dense 3D point clouds reconstructed via Structure-from-Motion and Multi-View Stereo techniques. Semantic annotations were produced through manual 2D labeling and projected into 3D space. Unlike existing datasets, 3DAeroRelief captures 3D large-scale outdoor environments with fine-grained structural damage in real-world disaster contexts. UAVs enable affordable, flexible, and safe data collection in hazardous areas, making them particularly well-suited for emergency scenarios. To demonstrate the utility of 3DAeroRelief, we evaluate several state-of-the-art 3D segmentation models on the dataset to highlight both the challenges and opportunities of 3D scene understanding in disaster response. Our dataset serves as a valuable resource for advancing robust 3D vision systems in real-world applications for post-disaster scenarios.

cs.CV

ZeShot-VQA: Zero-Shot Visual Question Answering Framework with Answer Mapping for Natural Disaster Damage Assessment

Natural disasters usually affect vast areas and devastate infrastructures. Performing a timely and efficient response is crucial to minimize the impact on affected communities, and data-driven approaches are the best choice. Visual question answering (VQA) models help management teams to achieve in-depth understanding of damages. However, recently published models do not possess the ability to answer open-ended questions and only select the best answer among a predefined list of answers. If we want to ask questions with new additional possible answers that do not exist in the predefined list, the model needs to be fin-tuned/retrained on a new collected and annotated dataset, which is a time-consuming procedure. In recent years, large-scale Vision-Language Models (VLMs) have earned significant attention. These models are trained on extensive datasets and demonstrate strong performance on both unimodal and multimodal vision/language downstream tasks, often without the need for fine-tuning. In this paper, we propose a VLM-based zero-shot VQA (ZeShot-VQA) method, and investigate the performance of on post-disaster FloodNet dataset. Since the proposed method takes advantage of zero-shot learning, it can be applied on new datasets without fine-tuning. In addition, ZeShot-VQA is able to process and generate answers that has been not seen during the training procedure, which demonstrates its flexibility.

cs.CV

Automated Segmentation of Large Image Datasets using Artificial Intelligence for Microstructure Characterisation, Damage Analysis and High-Throughput Modelling Input

Many properties of commonly used materials are driven by their microstructure, which can be influenced by the composition and manufacturing processes. To optimise future materials, understanding the microstructure is critically important. Here, we present two novel approaches based on artificial intelligence that allow the segmentation of the phases of a microstructure for which simple numerical approaches, such as thresholding, are not applicable: One is based on the nnU-Net neural network, and the other on generative adversarial networks (GAN). Using large panoramic scanning electron microscopy images of dual-phase steels as a case study, we demonstrate how both methods effectively segment intricate microstructural details, including martensite, ferrite, and damage sites, for subsequent analysis. Either method shows substantial generalizability across a range of image sizes and conditions, including heat-treated microstructures with different phase configurations. The nnU-Net excels in mapping large image areas. Conversely, the GAN-based method performs reliably on smaller images, providing greater step-by-step control and flexibility over the segmentation process. This study highlights the benefits of segmented microstructural data for various purposes, such as calculating phase fractions, modelling material behaviour through finite element simulation, and conducting geometrical analyses of damage sites and the local properties of their surrounding microstructure.

cond-mat.mtrl-sci

Rate dependence of damage formation in metallic-intermetallic Mg-Al-Ca composites

We study a cast Mg-4.65Al-2.82Ca alloy with a microstructure containing $α$-Mg matrix reinforced with a C36 Laves phase skeleton. Such ternary alloys are targeted for elevated temperature applications in automotive engines since they possess excellent creep properties. However, in application, the alloy may be subjected to a wide range of strain rates and in material development, accelerated testing is often of essence. It is therefore crucial to understand the effect of such rate variations. Here, we focus on their impact on damage formation. Due to the locally highly variable skeleton forming the reinforcement in this alloy, we employ an analysis based on high resolution panoramic imaging by scanning electron microscopy coupled with automated damage analysis by deep learning-based object detection and classification convolutional neural network algorithm (YOLOV5). We find that with decreasing strain rate the dominant damage mechanism for a given strain level changes: at a strain rate of $5\cdot10^{-4}/s$ the evolution of microcracks in the C36 Laves phase governs damage formation. However , when the strain rate is decreased to $5\cdot10^{-6}/s$, interface decohesion at the $α$-Mg/Laves phase interfaces becomes equally important. We also observe a change in crack orientation indicating an increasing influence of plastic co-deformation of the α-Mg matrix and Laves phase. We attribute this transition in leading damage mechanism to thermally activated processes at the interface.

cond-mat.mtrl-sci

Three-Dimensional Damage Characterisation in Dual Phase Steel using Deep Learning

High performance sheet metals with a multi-phase microstructure suffer from deformation induced damage formation during forming in the constituent phases but importantly also where these intersect. To capture damage in terms of the physical processes in three dimensions (3D) and its stochastic nature during deformation, two challenges remain to be tackled: First, bridging high resolution analysis towards large scales to consider statistical data and, second, characterising in 3D with a resolution appropriate for sub-micron sized voids at a large scale. Here, we present how this can be achieved using panoramic scanning electron microscopy (SEM), metallographic serial sectioning, and deep-learning assisted automatic image analysis. This brings together the 3D evolution of active damage mechanisms with volumetric and environmental information for thousands of individual damage sites. We also assess potential surface preparation artefacts in 2D analyses. Overall, we find that for the material considered here, a dual phase (DP800) steel, martensite cracking is the dominant but not sole origin of deformation induced damage and that for a quantitative comparison of damage density, metallographic preparation can induce additional surface damage density far exceeding what is commonly induced between uniaxial straining steps.

cond-mat.mtrl-sci