SearcharxivSearch

arXiv subjects

Angel Bueno Rodriguez

Publications and source records attributed to Angel Bueno Rodriguez.

6 recordsLinked to original sources

From Simulation to Real Scans: Anomaly Detection in Maritime Cargo with Muon Scattering Tomography

Maritime cargo inspection requires imaging technologies capable of detecting concealed threats within dense, sealed containers, a role for which Muon Scattering Tomography (MST) is well suited: it images their interior through the density-dependent deflection of naturally occurring cosmic muons. However, MST remains constrained by the scarcity of labeled scans and by a cosmic muon flux that is both low and stochastic. Anomaly detection algorithms must therefore be trained on simulations, yet operate on measured scans acquired under different conditions, a sim-to-real gap that remains a central obstacle to operational deployment. We present the first end-to-end anomaly detection framework for maritime MST, from physically consistent simulations to validation on real container scans from the SilentBorder demonstration campaign. The task is cast as an out-of-distribution problem: the framework learns the spatial configurations of benign cargo and flags threats as deviations in the reconstruction error space, remaining agnostic to threat type and geometry. An attention U-Net, trained exclusively on benign synthetic scenes, preserves small-scale scattering signatures through its skip connections, and contraband consequently persists in the pixel-wise reconstruction error instead of being absorbed into the reconstructed background. A scoring function, the Homogeneity Index (HI), suppresses spatially uniform cosmic-ray statistical noise while amplifying coherent anomaly signatures: where pixel-level metrics collapse under a change of cargo configuration, HI retains its discriminative power. We evaluate three training strategies across two distinct cargo configurations under operational one-hour scan times, and test the best model on real muon cargo scans. The results for the studied scenarios indicate that the sim-to-real gap can be bridged.

physics.ins-det

Synthetic-to-Real Domain Bridging for Single-View 3D Reconstruction of Ships for Maritime Monitoring

Three-dimensional (3D) reconstruction of ships is an important part of maritime monitoring, allowing improved visualization, inspection, and decision-making in real-world monitoring environments. However, most state-ofthe-art 3D reconstruction methods require multi-view supervision, annotated 3D ground truth, or are computationally intensive, making them impractical for real-time maritime deployment. In this work, we present an efficient pipeline for single-view 3D reconstruction of real ships by training entirely on synthetic data and requiring only a single view at inference. Our approach uses the Splatter Image network, which represents objects as sparse sets of 3D Gaussians for rapid and accurate reconstruction from single images. The model is first fine-tuned on synthetic ShapeNet vessels and further refined with a diverse custom dataset of 3D ships, bridging the domain gap between synthetic and real-world imagery. We integrate a state-of-the-art segmentation module based on YOLOv8 and custom preprocessing to ensure compatibility with the reconstruction network. Postprocessing steps include real-world scaling, centering, and orientation alignment, followed by georeferenced placement on an interactive web map using AIS metadata and homography-based mapping. Quantitative evaluation on synthetic validation data demonstrates strong reconstruction fidelity, while qualitative results on real maritime images from the ShipSG dataset confirm the potential for transfer to operational maritime settings. The final system provides interactive 3D inspection of real ships without requiring real-world 3D annotations. This pipeline provides an efficient, scalable solution for maritime monitoring and highlights a path toward real-time 3D ship visualization in practical applications. Interactive demo: https://dlr-mi.github.io/ship3d-demo/.

cs.CV

UTrack: Multi-Object Tracking with Uncertain Detections

The tracking-by-detection paradigm is the mainstream in multi-object tracking, associating tracks to the predictions of an object detector. Although exhibiting uncertainty through a confidence score, these predictions do not capture the entire variability of the inference process. For safety and security critical applications like autonomous driving, surveillance, etc., knowing this predictive uncertainty is essential though. Therefore, we introduce, for the first time, a fast way to obtain the empirical predictive distribution during object detection and incorporate that knowledge in multi-object tracking. Our mechanism can easily be integrated into state-of-the-art trackers, enabling them to fully exploit the uncertainty in the detections. Additionally, novel association methods are introduced that leverage the proposed mechanism. We demonstrate the effectiveness of our contribution on a variety of benchmarks, such as MOT17, MOT20, DanceTrack, and KITTI.

cs.CV

The 2nd Workshop on Maritime Computer Vision (MaCVi) 2024

The 2nd Workshop on Maritime Computer Vision (MaCVi) 2024 addresses maritime computer vision for Unmanned Aerial Vehicles (UAV) and Unmanned Surface Vehicles (USV). Three challenges categories are considered: (i) UAV-based Maritime Object Tracking with Re-identification, (ii) USV-based Maritime Obstacle Segmentation and Detection, (iii) USV-based Maritime Boat Tracking. The USV-based Maritime Obstacle Segmentation and Detection features three sub-challenges, including a new embedded challenge addressing efficicent inference on real-world embedded devices. This report offers a comprehensive overview of the findings from the challenges. We provide both statistical and qualitative analyses, evaluating trends from over 195 submissions. All datasets, evaluation code, and the leaderboard are available to the public at https://macvi.org/workshop/macvi24.

cs.CV

B2G4: A synthetic data pipeline for the integration of Blender models in Geant4 simulation toolkit

The correctness and precision of particle physics simulation software, such as Geant4, is expected to yield results that closely align with real-world observations or well-established theoretical predictions. Notably, the accuracy of these simulated outcomes is contingent upon the software's capacity to encapsulate detailed attributes, including its prowess in generating or incorporating complex geometrical constructs. While the imperatives of precision and accuracy are essential in these simulations, the need to manually code highly detailed geometries emerges as a salient bottleneck in developing software-driven physics simulations. This research proposes Blender-to-Geant4 (B2G4), a modular data workflow that utilizes Blender to create 3D scenes, which can be exported as geometry input for Geant4. B2G4 offers a range of tools to streamline the creation of simulation scenes with multiple complex geometries and realistic material properties. Here, we demonstrate the use of B2G4 in a muon scattering tomography application to image the interior of a sealed steel structure. The modularity of B2G4 paves the way for the designed scenes and tools to be embedded not only in Geant4, but in other scientific applications or simulation software.

physics.comp-ph

Look ATME: The Discriminator Mean Entropy Needs Attention

Generative adversarial networks (GANs) are successfully used for image synthesis but are known to face instability during training. In contrast, probabilistic diffusion models (DMs) are stable and generate high-quality images, at the cost of an expensive sampling procedure. In this paper, we introduce a simple method to allow GANs to stably converge to their theoretical optimum, while bringing in the denoising machinery from DMs. These models are combined into a simpler model (ATME) that only requires a forward pass during inference, making predictions cheaper and more accurate than DMs and popular GANs. ATME breaks an information asymmetry existing in most GAN models in which the discriminator has spatial knowledge of where the generator is failing. To restore the information symmetry, the generator is endowed with knowledge of the entropic state of the discriminator, which is leveraged to allow the adversarial game to converge towards equilibrium. We demonstrate the power of our method in several image-to-image translation tasks, showing superior performance than state-of-the-art methods at a lesser cost. Code is available at https://github.com/DLR-MI/atme

cs.CV