Searcharxiv⌕ Search

arXiv subjects

Reza Hoseinnezhad

Publications and source records attributed to Reza Hoseinnezhad.

At least 19 recordsLinked to original sources

RMR-P: Road Metadata-Aware Restoration for Pavement Inspection

Road-surface images captured by vehicle-mounted cameras are often degraded by motion blur, defocus, poor illumination, and noise due to vehicle motion, camera limitations, and varying environmental conditions. These degradations can obscure thin cracks and pothole boundaries that are critical for accurate road-defect detection. This paper presents RMR-P, a restoration network designed to recover defect-relevant information from degraded road images. It estimates degradation characteristics from the input image and can optionally incorporate external degradation parameters to guide restoration. To evaluate whether the recovered information improves downstream detection, a clean-trained YOLO11s detector is applied to degraded and restored images without further modification. Experiments on the IVCNZ and PCM datasets, with known synthetic degradation parameters provided as conditioning information, demonstrate that RMR-P achieves the highest mAP50 in seven of eight held-out degradation conditions, including improvements from 0.140 to 0.427 under IVCNZ motion blur and from 0.060 to 0.233 under PCM defocus. Moreover, our ablation studies show that preserving fine pavement details (detail-preserving pathway) provides the largest contribution to defect-detection improvement, while degradation conditioning and task-guided training offer complementary benefits.

cs.CV↗

Distributed Multi-Sensor Control for Multi-Target Tracking Using Adaptive Complementary Fusion for LMB Densities

Tracking multiple targets in dynamic environments using distributed sensor networks is a fundamental problem in statistical signal processing. In such scenarios, the network of mobile sensors must coordinate their actions to accurately estimate the locations and trajectories of multiple targets, balancing limited computation and communication resources with multi-target tracking accuracy. Multi-sensor control methods can improve the performance of these networks by enabling efficient utilization of resources and enhancing the accuracy of the estimated target states. This paper proposes a novel multi-sensor control method that utilizes multi-agent coordinate descent to address this problem, ensuring distributed consensus of optimal sensor actions throughout the sensor network. To achieve this, a novel adaptive complementary fusion approach that prioritizes information from the most informative sensors is developed. Our method improves computational tractability and enables fully distributed control, ensuring the scalability and flexibility necessary for large-scale real-time sensing systems. Experimental results on several challenging multi-target tracking scenarios demonstrate that our approach significantly improves both multi-target tracking accuracy and computation efficiency over competing methods.

eess.SP↗

Enhanced Multi-Target Tracking in Dynamic Environments: Distributed Flooding Control in the Random Finite Set Framework

Tracking multiple targets in dynamic environments using distributed sensor networks is a challenging problem for situational awareness in connected autonomous vehicles (CAVs). In such scenarios, the network of mobile sensors must coordinate their actions to accurately estimate the locations and trajectories of multiple targets, balancing limited computation and communication resources with multi-target tracking accuracy. Multi-sensor control methods can improve the performance of these networks by enabling efficient utilization of resources and enhancing the accuracy of the estimated target states. This paper proposes a novel multi-sensor control method that utilizes flooding-based communication to address this problem, ensuring distributed consensus of optimal sensor actions throughout the sensor network. Our method improves computational tractability and enables fully distributed control, ensuring the scalability and flexibility necessary for real-time CAV applications. Experimental results on several challenging multi-target tracking scenarios demonstrate that our approach significantly improves both multi-target tracking accuracy and computation time over competing methods.

eess.SY↗

Automated Keypoint Estimation for Self-Piercing Rivet Joints Using micro-CT Imaging and Transfer Learning

The structural integrity of self-piercing rivet (SPR) joints is critical in automotive industries, yet its evaluation poses challenges due to the limitations of traditional destructive methods. This research introduces an innovative approach for non-destructive evaluation using micro-CT imaging, Micro-Computed Tomography, combined with machine vision and deep learning techniques, specifically focusing on automated keypoint estimation to assess joint quality. Recognizing the scarcity of real micro-CT data, this study utilizes synthetic data for initial model training, followed by transfer learning to adapt the model for real-world conditions. A UNet-based architecture is employed to localize three keypoints with precision, enabling the measurement of critical parameters such as head height, interlock, and bottom layer thickness. Extensive validation demonstrates that pre-training on synthetic data, complemented by fine-tuning with limited real data, bridges domain gaps and enhances predictive accuracy. The proposed framework not only offers a scalable and cost-efficient solution for evaluating SPR joints but also establishes a foundation for broader applications of machine vision and non-destructive testing in manufacturing processes. By addressing data scarcity and leveraging advanced machine learning techniques, this work represents a significant step toward automated quality control in engineering contexts.

cs.CE↗

Single Domain Generalization via Normalised Cross-correlation Based Convolutions

Deep learning techniques often perform poorly in the presence of domain shift, where the test data follows a different distribution than the training data. The most practically desirable approach to address this issue is Single Domain Generalization (S-DG), which aims to train robust models using data from a single source. Prior work on S-DG has primarily focused on using data augmentation techniques to generate diverse training data. In this paper, we explore an alternative approach by investigating the robustness of linear operators, such as convolution and dense layers commonly used in deep learning. We propose a novel operator called XCNorm that computes the normalized cross-correlation between weights and an input feature patch. This approach is invariant to both affine shifts and changes in energy within a local feature patch and eliminates the need for commonly used non-linear activation functions. We show that deep neural networks composed of this operator are robust to common semantic distribution shifts. Furthermore, our empirical results on single-domain generalization benchmarks demonstrate that our proposed technique performs comparably to the state-of-the-art methods.

cs.CV↗

IT-RUDA: Information Theory Assisted Robust Unsupervised Domain Adaptation

Distribution shift between train (source) and test (target) datasets is a common problem encountered in machine learning applications. One approach to resolve this issue is to use the Unsupervised Domain Adaptation (UDA) technique that carries out knowledge transfer from a label-rich source domain to an unlabeled target domain. Outliers that exist in either source or target datasets can introduce additional challenges when using UDA in practice. In this paper, $α$-divergence is used as a measure to minimize the discrepancy between the source and target distributions while inheriting robustness, adjustable with a single parameter $α$, as the prominent feature of this measure. Here, it is shown that the other well-known divergence-based UDA techniques can be derived as special cases of the proposed method. Furthermore, a theoretical upper bound is derived for the loss in the target domain in terms of the source loss and the initial $α$-divergence between the two domains. The robustness of the proposed method is validated through testing on several benchmarked datasets in open-set and partial UDA setups where extra classes existing in target and source datasets are considered as outliers.

cs.LG↗

Interaction-Aware Labeled Multi-Bernoulli Filter

Tracking multiple objects through time is an important part of an intelligent transportation system. Random finite set (RFS)-based filters are one of the emerging techniques for tracking multiple objects. In multi-object tracking (MOT), a common assumption is that each object is moving independent of its surroundings. But in many real-world applications, target objects interact with one another and the environment. Such interactions, when considered for tracking, are usually modeled by an interactive motion model which is application specific. In this paper, we present a novel approach to incorporate target interactions within the prediction step of an RFS-based multi-target filter, i.e. labeled multi-Bernoulli (LMB) filter. The method has been developed for two practical applications of tracking a coordinated swarm and vehicles. The method has been tested for a complex vehicle tracking dataset and compared with the LMB filter through the OSPA and OSPA$^{(2)}$ metrics. The results demonstrate that the proposed interaction-aware method depicts considerable performance enhancement over the LMB filter in terms of the selected metrics.

eess.SP↗

ITSA: An Information-Theoretic Approach to Automatic Shortcut Avoidance and Domain Generalization in Stereo Matching Networks

State-of-the-art stereo matching networks trained only on synthetic data often fail to generalize to more challenging real data domains. In this paper, we attempt to unfold an important factor that hinders the networks from generalizing across domains: through the lens of shortcut learning. We demonstrate that the learning of feature representations in stereo matching networks is heavily influenced by synthetic data artefacts (shortcut attributes). To mitigate this issue, we propose an Information-Theoretic Shortcut Avoidance~(ITSA) approach to automatically restrict shortcut-related information from being encoded into the feature representations. As a result, our proposed method learns robust and shortcut-invariant features by minimizing the sensitivity of latent features to input variations. To avoid the prohibitive computational cost of direct input sensitivity optimization, we propose an effective yet feasible algorithm to achieve robustness. We show that using this method, state-of-the-art stereo matching networks that are trained purely on synthetic data can effectively generalize to challenging and previously unseen real data scenarios. Importantly, the proposed method enhances the robustness of the synthetic trained networks to the point that they outperform their fine-tuned counterparts (on real data) for challenging out-of-domain stereo datasets.

cs.CV↗

Anomaly Detection of Defect using Energy of Point Pattern Features within Random Finite Set Framework

In this paper, we propose an efficient approach for industrial defect detection that is modeled based on anomaly detection using point pattern data. Most recent works use \textit{global features} for feature extraction to summarize image content. However, global features are not robust against lighting and viewpoint changes and do not describe the image's geometrical information to be fully utilized in the manufacturing industry. To the best of our knowledge, we are the first to propose using transfer learning of local/point pattern features to overcome these limitations and capture geometrical information of the image regions. We model these local/point pattern features as a random finite set (RFS). In addition we propose RFS energy, in contrast to RFS likelihood as anomaly score. The similarity distribution of point pattern features of the normal sample has been modeled as a multivariate Gaussian. Parameters learning of the proposed RFS energy does not require any heavy computation. We evaluate the proposed approach on the MVTec AD dataset, a multi-object defect detection dataset. Experimental results show the outstanding performance of our proposed approach compared to the state-of-the-art methods, and the proposed RFS energy outperforms the state-of-the-art in the few shot learning settings.

cs.CV↗

Robust Pooling through the Data Mode

The task of learning from point cloud data is always challenging due to the often occurrence of noise and outliers in the data. Such data inaccuracies can significantly influence the performance of state-of-the-art deep learning networks and their ability to classify or segment objects. While there are some robust deep learning approaches, they are computationally too expensive for real-time applications. This paper proposes a deep learning solution that includes a novel robust pooling layer which greatly enhances network robustness and performs significantly faster than state-of-the-art approaches. The proposed pooling layer looks for data a mode/cluster using two methods, RANSAC, and histogram, as clusters are indicative of models. We tested the pooling layer into frameworks such as Point-based and graph-based neural networks, and the tests showed enhanced robustness as compared to robust state-of-the-art methods.

cs.CV↗

Evaluation of Point Pattern Features for Anomaly Detection of Defect within Random Finite Set Framework

Defect detection in the manufacturing industry is of utmost importance for product quality inspection. Recently, optical defect detection has been investigated as an anomaly detection using different deep learning methods. However, the recent works do not explore the use of point pattern features, such as SIFT for anomaly detection using the recently developed set-based methods. In this paper, we present an evaluation of different point pattern feature detectors and descriptors for defect detection application. The evaluation is performed within the random finite set framework. Handcrafted point pattern features, such as SIFT as well as deep features are used in this evaluation. Random finite set-based defect detection is compared with state-of-the-arts anomaly detection methods. The results show that using point pattern features, such as SIFT as data points for random finite set-based anomaly detection achieves the most consistent defect detection accuracy on the MVTec-AD dataset.

cs.CV↗

Multiple Instance-Based Video Anomaly Detection using Deep Temporal Encoding-Decoding

In this paper, we propose a weakly supervised deep temporal encoding-decoding solution for anomaly detection in surveillance videos using multiple instance learning. The proposed approach uses both abnormal and normal video clips during the training phase which is developed in the multiple instance framework where we treat video as a bag and video clips as instances in the bag. Our main contribution lies in the proposed novel approach to consider temporal relations between video instances. We deal with video instances (clips) as a sequential visual data rather than independent instances. We employ a deep temporal and encoder network that is designed to capture spatial-temporal evolution of video instances over time. We also propose a new loss function that is smoother than similar loss functions recently presented in the computer vision literature, and therefore; enjoys faster convergence and improved tolerance to local minima during the training phase. The proposed temporal encoding-decoding approach with modified loss is benchmarked against the state-of-the-art in simulation studies. The results show that the proposed method performs similar to or better than the state-of-the-art solutions for anomaly detection in video surveillance applications.

cs.CV↗

Adjusting Bias in Long Range Stereo Matching: A semantics guided approach

Stereo vision generally involves the computation of pixel correspondences and estimation of disparities between rectified image pairs. In many applications, including simultaneous localization and mapping (SLAM) and 3D object detection, the disparities are primarily needed to calculate depth values and the accuracy of depth estimation is often more compelling than disparity estimation. The accuracy of disparity estimation, however, does not directly translate to the accuracy of depth estimation, especially for faraway objects. In the context of learning-based stereo systems, this is largely due to biases imposed by the choices of the disparity-based loss function and the training data. Consequently, the learning algorithms often produce unreliable depth estimates of foreground objects, particularly at large distances~($>50$m). To resolve this issue, we first analyze the effect of those biases and then propose a pair of novel depth-based loss functions for foreground and background, separately. These loss functions are tunable and can balance the inherent bias of the stereo learning algorithms. The efficacy of our solution is demonstrated by an extensive set of experiments, which are benchmarked against state of the art. We show on KITTI~2015 benchmark that our proposed solution yields substantial improvements in disparity and depth estimation, particularly for objects located at distances beyond 50 meters, outperforming the previous state of the art by $10\%$.

cs.CV↗

Robust Object Classification Approach using Spherical Harmonics

In this paper, we present a robust spherical harmonics approach for the classification of point cloud-based objects. Spherical harmonics have been used for classification over the years, with several frameworks existing in the literature. These approaches use variety of spherical harmonics based descriptors to classify objects. We first investigated these frameworks robustness against data augmentation, such as outliers and noise, as it has not been studied before. Then we propose a spherical convolution neural network framework for robust object classification. The proposed framework uses the voxel grid of concentric spheres to learn features over the unit ball. Our proposed model learn features that are less sensitive to data augmentation due to the selected sampling strategy and the designed convolution operation. We tested our proposed model against several types of data augmentation, such as noise and outliers. Our results show that the proposed model outperforms the state of art networks in terms of robustness to data augmentation.

cs.CV↗

Computationally Efficient Distributed Multi-sensor Fusion with Multi-Bernoulli Filter

This paper proposes a computationally efficient algorithm for distributed fusion in a sensor network in which multi-Bernoulli (MB) filters are locally running in every sensor node for multi-target tracking. The generalized Covariance Intersection (GCI) fusion rule is employed to fuse multiple MB random finite set densities. The fused density comprises a set of fusion hypotheses that grow exponentially with the number of Bernoulli components. Thus, GCI fusion with MB filters can become computationally intractable in practical applications that involve tracking of even a moderate number of objects. In order to accelerate the multi-sensor fusion procedure, we derive a theoretically sound approximation to the fused density. The number of fusion hypotheses in the resulting density is significantly smaller than the original fused density. It also has a parallelizable structure that allows multiple clusters of Bernoulli components to be fused independently. By carefully clustering Bernoulli components into isolated clusters using the GCI divergence as the distance metric, we propose an alternative to build exactly the approximated density without exhaustively computing all the fusion hypotheses. The combination of the proposed approximation technique and the fast clustering algorithm can enable a novel and fast GCIMB fusion implementation. Our analysis shows that the proposed fusion method can dramatically reduce the computational and memory requirements with small bounded L1-error. The Gaussian mixture implementation of the proposed method is also presented. In various numerical experiments, including a challenging scenario with up to forty objects, the efficacy of the proposed fusion method is demonstrated.

eess.SY↗

Multi-object Tracking for Generic Observation Model Using Labeled Random Finite Sets

This paper presents an exact Bayesian filtering solution for the multi-object tracking problem with the generic observation model. The proposed solution is designed in the labeled random finite set framework, using the product styled representation of labeled multi-object densities, with the standard multi-object transition kernel and no particular simplifying assumptions on the multi-object likelihood. Computationally tractable solutions are also devised by applying a principled approximation involving the replacement of the full multi-object density with a labeled multi-Bernoulli density that minimizes the Kullback-Leibler divergence and preserves the first-order moment. To achieve the fast performance, a dynamic grouping procedure based implementation is presented with a step-by-step algorithm. The performance of the proposed filter and its tractable implementations are verified and compared with the state-of-the-art in numerical experiments.

eess.SY↗

Robust Distributed Fusion with Labeled Random Finite Sets

This paper considers the problem of the distributed fusion of multi-object posteriors in the labeled random finite set filtering framework, using Generalized Covariance Intersection (GCI) method. Our analysis shows that GCI fusion with labeled multi-object densities strongly relies on label consistencies between local multi-object posteriors at different sensor nodes, and hence suffers from a severe performance degradation when perfect label consistencies are violated. Moreover, we mathematically analyze this phenomenon from the perspective of Principle of Minimum Discrimination Information and the so called yes-object probability. Inspired by the analysis, we propose a novel and general solution for the distributed fusion with labeled multi-object densities that is robust to label inconsistencies between sensors. Specifically, the labeled multi-object posteriors are firstly marginalized to their unlabeled posteriors which are then fused using GCI method. We also introduce a principled method to construct the labeled fused density and produce tracks formally. Based on the developed theoretical framework, we present tractable algorithms for the family of generalized labeled multi-Bernoulli (GLMB) filters including $δ$-GLMB, marginalized $δ$-GLMB and labeled multi-Bernoulli filters. The robustness and efficiency of the proposed distributed fusion algorithm are demonstrated in challenging tracking scenarios via numerical experiments.

eess.SY↗

Effective Sampling: Fast Segmentation Using Robust Geometric Model Fitting

Identifying the underlying models in a set of data points contaminated by noise and outliers, leads to a highly complex multi-model fitting problem. This problem can be posed as a clustering problem by the projection of higher order affinities between data points into a graph, which can then be clustered using spectral clustering. Calculating all possible higher order affinities is computationally expensive. Hence in most cases only a subset is used. In this paper, we propose an effective sampling method to obtain a highly accurate approximation of the full graph required to solve multi-structural model fitting problems in computer vision. The proposed method is based on the observation that the usefulness of a graph for segmentation improves as the distribution of hypotheses (used to build the graph) approaches the distribution of actual parameters for the given data. In this paper, we approximate this actual parameter distribution using a k-th order statistics based cost function and the samples are generated using a greedy algorithm coupled with a data sub-sampling strategy. The experimental analysis shows that the proposed method is both accurate and computationally efficient compared to the state-of-the-art robust multi-model fitting techniques. The code is publicly available from https://github.com/RuwanT/model-fitting-cbs.

cs.CV↗