SearcharxivSearch

arXiv subjects

Mohamed Youssef

Publications and source records attributed to Mohamed Youssef.

13 recordsLinked to original sources

Vision-Reasoning-Guided Occlusion Removal from Light Fields

Occlusion-robust scene recovery remains a major challenge in computational imaging, particularly where dense vegetation severely limits visibility. We propose a visionreasoning-guided light field occlusion removal framework combining light field integration (LFI) with vision-language model (VLM) semantic reasoning. Multi-view observations are first integrated via LFI to suppress foreground occlusions, producing an initial visibility-enhanced representation, a VLM then acts as a conditional semantic prior to restore degraded structures and fine details. A multi-sample fusion strategy aggregates multiple generated hypotheses to improve consistency and reduce hallucination. Experimental results on synthetic and real-world datasets show state-of-the-art performance, achieving the highest average SSIM across four synthetic benchmark scenes (4-Syn) and strong generalization across structured and unstructured acquisition settings, with applicability to search-and-rescue and exploratory robotic navigation.

cs.CV

Through-Foliage Surface-Temperature Reconstruction for Early Wildfire Detection

We present a method to reconstruct surface temperatures through forest vegetation by combining signal processing and machine learning, enabling fully automated aerial wildfire monitoring with drones for early fire detection. Synthetic aperture (SA) sensing reduces canopy occlusion but introduces thermal blur. To overcome this, we train a visual state space model to recover subtle thermal signals of partially occluded soil and fire hotspots from blurred data. To address limited real-world training data, we generate realistic surface temperature simulations using a latent diffusion model, temperature augmentation, and procedural thermal forest modeling. On simulated datasets, our method reduces RMSE by 2-2.5 versus conventional thermal and uncorrected SA imaging; in field experiments on hotspots, RMSE improved by 12.8-fold and 2.6-fold, respectively. Our approach also generalizes to other thermal signals, including human signatures, capturing morphology and extent -- critical where simple thresholding fails -- while conventional imaging struggles with partial occlusion.

cs.CV

Origins of suppressed self-diffusion of nanoscale constituents of a complex liquid

Understanding and ultimately controlling the transformations and properties of nanoscale systems, from proteins to synthetic nanomaterial assemblies, is limited by the inability to uncover their dynamics on their characteristic length and time scales. Here, we nevertheless demonstrate this ability using MHz X-ray photon correlation spectroscopy (XPCS) -- directly elucidating the characteristic microsecond-dynamics of density fluctuations of semiconductor nanocrystals (NCs), not only in a colloidal dispersion but also in a liquid phase consisting of densely packed, yet mobile, NCs with no long-range order. We find the wavevector-dependent fluctuation rates in the liquid phase are suppressed relative to those in the colloidal phase and relative to observations of densely packed repulsive particles. We show that the suppressed rates are due to a substantial decrease in the self-diffusion of NCs, which we attribute to explicit attractive interactions. Using coarse-grained simulations, we find that the extracted shape and strength of the interparticle potential explains the stability of the liquid phase, in contrast to the gelation observed via XPCS in many other charged colloidal systems. This work opens the door to elucidating fast, condensed phase dynamics in complex fluids and other nanoscale soft matter, such as densely packed proteins and non-equilibrium self-assembly processes, in addition to designing microscopic strategies to avert gelation.

cond-mat.soft

Ontology-Guided Diffusion for Zero-Shot Visual Sim2Real Transfer

Bridging the simulation-to-reality (sim2real) gap remains challenging as labelled real-world data is scarce. Existing diffusion-based approaches rely on unstructured prompts or statistical alignment, which do not capture the structured factors that make images look real. We introduce Ontology- Guided Diffusion (OGD), a neuro-symbolic zero-shot sim2real image translation framework that represents realism as structured knowledge. OGD decomposes realism into an ontology of interpretable traits -- such as lighting and material properties -- and encodes their relationships in a knowledge graph. From a synthetic image, OGD infers trait activations and uses a graph neural network to produce a global embedding. In parallel, a symbolic planner uses the ontology traits to compute a consistent sequence of visual edits needed to narrow the realism gap. The graph embedding conditions a pretrained instruction-guided diffusion model via cross-attention, while the planned edits are converted into a structured instruction prompt. Across benchmarks, our graph-based embeddings better distinguish real from synthetic imagery than baselines, and OGD outperforms state-of-the-art diffusion methods in sim2real image translations. Overall, OGD shows that explicitly encoding realism structure enables interpretable, data-efficient, and generalisable zero-shot sim2real transfer.

cs.CV

Single-shot sorting of Mössbauer time-domain data at X-ray free-electron lasers

Mössbauer spectroscopy is widely used to study structure and dynamics of matter with remarkably high energy resolution, provided by the narrow nuclear resonance line widths. However, the narrow width implies low count rates, such that experiments commonly average over extended measurement times or many x-ray pulses (``shots''). This averaging impedes the study of non-equilibrium phenomena. It has been suggested that X-ray free-electron lasers (XFELs) could enable Mössbauer single-shot measurements without averaging, and a proof-of-principle demonstration has been reported. However, so far, only a tiny fraction of all shots resulted in signal-photon numbers which are sufficiently high for a single-shot analysis. Here, we demonstrate coherent nuclear-forward-scattering of self-seeded XFEL radiation, with up to 900 signal-photons per shot. We develop a sorting approach which allows us to include all data on a single-shot level, independent of the signal content of the individual shots. It utilizes the presence of different dynamics classes, i.e. different nuclear evolutions after each excitation. Each shot is assigned to one of the classes, which can then be analyzed separately. Our approach determines the classes from the data without requiring theory modeling nor prior knowledge on the dynamics, making it also applicable to unknown phenomena. We envision that our approach opens up new grounds for Mössbauer science, enabling the study of out-of-equilibrium transient dynamics of the nuclei or their environment.

quant-ph

How Robot Dogs See the Unseeable: Improving Visual Interpretability via Peering for Exploratory Robots

In vegetated environments, such as forests, exploratory robots play a vital role in navigating complex, cluttered environments where human access is limited and traditional equipment struggles. Visual occlusion from obstacles, such as foliage, can severely obstruct a robot's sensors, impairing scene understanding. We show that "peering", a characteristic side-to-side movement used by insects to overcome their visual limitations, can also allow robots to markedly improve visual reasoning under partial occlusion. This is accomplished by applying core signal processing principles, specifically optical synthetic aperture sensing, together with the vision reasoning capabilities of modern large multimodal models. Peering enables real-time, high-resolution, and wavelength-independent perception, which is crucial for vision-based scene understanding across a wide range of applications. The approach is low-cost and immediately deployable on any camera-equipped robot. We investigated different peering motions and occlusion masking strategies, demonstrating that, unlike peering, state-of-the-art multi-view 3D vision techniques fail in these conditions due to their high susceptibility to occlusion. Our experiments were carried out on an industrial-grade quadrupedal robot. However, the ability to peer is not limited to such platforms, but potentially also applicable to bipedal, hexapod, wheeled, or crawling platforms. Robots that can effectively see through partial occlusion will gain superior perception abilities - including enhanced scene understanding, situational awareness, camouflage breaking, and advanced navigation in complex environments.

cs.RO

Coherent X-rays reveal anomalous molecular diffusion and cage effects in crowded protein solutions

Understanding protein motion within the cell is crucial for predicting reaction rates and macromolecular transport in the cytoplasm. A key question is how crowded environments affect protein dynamics through hydrodynamic and direct interactions at molecular length scales. Using megahertz X-ray Photon Correlation Spectroscopy (MHz-XPCS) at the European X-ray Free Electron Laser (EuXFEL), we investigate ferritin diffusion at microsecond time scales. Our results reveal anomalous diffusion, indicated by the non-exponential decay of the intensity autocorrelation function $g_2(q,t)$ at high concentrations. This behavior is consistent with the presence of cage-trapping in between the short- and long-time protein diffusion regimes. Modeling with the $δγ$-theory of hydrodynamically interacting colloidal spheres successfully reproduces the experimental data by including a scaling factor linked to the protein direct interactions. These findings offer new insights into the complex molecular motion in crowded protein solutions, with potential applications for optimizing ferritin-based drug delivery, where protein diffusion is the rate-limiting step.

cond-mat.soft

Depletion-Induced Interactions Modulate Nanoscale Protein Diffusion in Polymeric Crowder Solutions

Macromolecular crowding plays a crucial role in modulating protein dynamics in cellular and in vitro environments. Polymeric crowders such as dextran and Ficoll are known to induce entropic forces, including depletion interactions, that promote structural organization, but the nanoscale consequences for protein dynamics remain less well understood. Here, we employ megahertz X-ray photon correlation spectroscopy (MHz-XPCS) at the European XFEL to probe the dynamics of the protein ferritin in solutions containing sucrose, dextran, and Ficoll. We find that depletion-driven short-range attractions combined with long-range repulsions give rise to intermediate-range order (IRO) once the polysaccharide overlap concentration $c^*$ is exceeded. These IRO features fluctuate on microsecond to millisecond timescales, strongly modulating the collective dynamics of ferritin. The magnitude of these effects depends sensitively on crowder type, concentration, and molecular weight. Normalizing the crowder concentration by $c^*$ reveals scaling behavior in ferritin self-diffusion with a crossover near 2$c^*$, marking a transition from depletion-enhanced mobility to viscosity-dominated slowing. Our results demonstrate that bulk properties alone cannot account for protein dynamics in crowded solutions, underscoring the need to include polymer-specific interactions and depletion theory in models of crowded environments.

cond-mat.soft

DeepForest: Sensing Into Self-Occluding Volumes of Vegetation With Aerial Imaging

Access to below-canopy volumetric vegetation data is crucial for understanding ecosystem dynamics. We address the long-standing limitation of remote sensing to penetrate deep into dense canopy layers. LiDAR and radar are currently considered the primary options for measuring 3D vegetation structures, while cameras can only extract the reflectance and depth of top layers. Using conventional, high-resolution aerial images, our approach allows sensing deep into self-occluding vegetation volumes, such as forests. It is similar in spirit to the imaging process of wide-field microscopy, but can handle much larger scales and strong occlusion. We scan focal stacks by synthetic-aperture imaging with drones and reduce out-of-focus signal contributions using pre-trained 3D convolutional neural networks with mean squared error (MSE) as the loss function. The resulting volumetric reflectance stacks contain low-frequency representations of the vegetation volume. Combining multiple reflectance stacks from various spectral channels provides insights into plant health, growth, and environmental conditions throughout the entire vegetation volume. Compared with simulated ground truth, our correction leads to ~x7 average improvements (min: ~x2, max: ~x12) for forest densities of 220 trees/ha - 1680 trees/ha. In our field experiment, we achieved an MSE of 0.05 when comparing with the top-vegetation layer that was measured with classical multispectral aerial imaging.

cs.CV

An aerial color image anomaly dataset for search missions in complex forested terrain

After a family murder in rural Germany, authorities failed to locate the suspect in a vast forest despite a massive search. To aid the search, a research aircraft captured high-resolution aerial imagery. Due to dense vegetation obscuring small clues, automated analysis was ineffective, prompting a crowd-search initiative. This effort produced a unique dataset of labeled, hard-to-detect anomalies under occluded, real-world conditions. It can serve as a benchmark for improving anomaly detection approaches in complex forest environments, supporting manhunts and rescue operations. Initial benchmark tests showed existing methods performed poorly, highlighting the need for context-aware approaches. The dataset is openly accessible for offline processing. An additional interactive web interface supports online viewing and dynamic growth by allowing users to annotate and submit new findings.

cs.CV

Softness and Hydrodynamic Interactions Regulate Lipoprotein Transport in Crowded Yolk Environments

Low-density lipoproteins (LDLs) serve as nutrient reservoirs in egg yolk for embryonic development and as promising drug carriers. Both roles critically depend on their mobility in densely crowded biological environments. Under these crowded conditions, diffusion is hindered by transient confinement within dynamic cages formed by neighboring particles, driven by solvent-mediated hydrodynamic interactions and memory effects -- phenomena that have remained challenging to characterize computationally and experimentally. Here, we employ megahertz X-ray photon correlation spectroscopy to directly probe the cage dynamics of LDLs in yolk-plasma across various concentrations. We find that LDLs undergo anomalous diffusion, experiencing $\approx$ 100-fold reduction in self-diffusion at high concentrations compared to dilute solutions. This drastic slowing-down is attributed to a combination of hydrodynamic interactions, direct particle-particle interactions, and the inherent softness of LDL particles. Despite reduced dynamics, yolk-plasma remains as a liquid, yet sluggish, balancing dense packing, structural stability, and fluidity essential for controlled lipid release during embryogenesis.

cond-mat.soft

Dark-Field X-ray Microscopy for 2D and 3D imaging of Microstructural Dynamics at the European X-ray Free Electron Laser

Dark field X-ray microscopy (DXFM) can visualize microstructural distortions in bulk crystals. Using the femtosecond X-ray pulses generated by X-ray free-electron lasers (XFEL), DFXM can achieve sub-μm spatial resolution and <100 fs time resolution simultaneously. In this paper, we demonstrate ultrafast DFXM measurements at the European XFEL to visualize an optically-driven longitudinal strain wave propagating through a diamond single crystal. We also present two DFXM scanning modalities that are new to the XFEL sources: spatially 3D and 2D axial-strain scans with sub-μm spatial resolution. With this progress in XFEL-based DFXM, we discuss new opportunities to study multi-timescale spatio-temporal dynamics of microstructures.

cond-mat.mes-hall

Fusion of Single and Integral Multispectral Aerial Images

An adequate fusion of the most significant salient information from multiple input channels is essential for many aerial imaging tasks. While multispectral recordings reveal features in various spectral ranges, synthetic aperture sensing makes occluded features visible. We present a first and hybrid (model- and learning-based) architecture for fusing the most significant features from conventional aerial images with the ones from integral aerial images that are the result of synthetic aperture sensing for removing occlusion. It combines the environment's spatial references with features of unoccluded targets that would normally be hidden by dense vegetation. Our method outperforms state-of-the-art two-channel and multi-channel fusion approaches visually and quantitatively in common metrics, such as mutual information, visual information fidelity, and peak signal-to-noise ratio. The proposed model does not require manually tuned parameters, can be extended to an arbitrary number and arbitrary combinations of spectral channels, and is reconfigurable for addressing different use cases. We demonstrate examples for search and rescue, wildfire detection, and wildlife observation.

eess.IV