Searcharxiv⌕ Search

arXiv subjects

Mohamed Abouagour

Publications and source records attributed to Mohamed Abouagour.

5 recordsLinked to original sources

Audit Before You Commit: Locating Belief Failures in Active Identification for One-Shot Manipulation

A robot that probes a few times before one irreversible action, such as tapping a surface before inserting a peg, must decide when the evidence is enough to commit. We argue that this decision rests on two conditions that existing methods do not separate: the belief must still cover the truth in the coordinate that decides the action, and the failure model that scores actions must track realized failure. We audit both conditions separately, offline and with ground truth, on a deployed probe-then-commit pipeline: a particle belief, a scenario failure score, and one commit. On simulated insertion, more taps sharpen the belief while the truth leaves its support on 16.9% of episodes and the failure score turns optimistic by 0.31. Conformal calibration restores coverage but not the decision: confidently wrong instances still pass a confidence gate. The audit's signatures instead point at the observation model, where a hand scan finds a 2.1 mm error in the tap boundary. Correcting that one number cuts failure from 0.354 to 0.112 on untouched instances and transfers unrefitted to a second engine, while in a third engine the same audit suggests an execution-model mismatch instead. Across seven task families in three engines, a few probes at a fixed executor reduce miss or failure. On a physical arm inserting a tool into a rigid pocket by touch, the gain and the audit's two conditions reproduce, and replaying the recorded taps under an injected model error shows the audit's signature on real data. Additional materials are available at https://sites.google.com/view/auditbeforeyoucommit.

cs.RO↗

ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans

We introduce ResPlan, a dataset of 17,000 residential floor plans with vector geometry, room-connectivity graphs, and metric-scale coordinates. Each plan annotates walls, doors, windows, and functional spaces (kitchens, bedrooms, bathrooms, balconies, and others) under a 17-class taxonomy, with polygons in pixel and meter coordinates. Four typed edges (via_door, adjacency, direct, via_window) accompany every plan, supporting graph-based generation and spatial reasoning. Compared with RPLAN (Wu et al., 2019), which is raster-only with about 6.7 rooms per plan and an observed maximum of 8 functional rooms in our converted split, and MSD (van Engelenburg et al., 2024), which is floor-plate-level and requires extraction, ResPlan provides self-contained unit-level layouts averaging 8.1 functional rooms and 9.2 graph nodes, spanning apartments, villas, and multi-wing residences. The release includes the dataset, loading and post-processing code, a canonical split, and baselines for three benchmark tasks: semantic room labeling, constrained floor-plan generation, and plan-to-graph extraction. The dataset and code are publicly available.

cs.CV↗

GFLAN: Generative Functional Layouts

Automated floor plan generation lies at the intersection of combinatorial search, geometric constraint satisfaction, and functional design requirements -- a confluence that has historically resisted a unified computational treatment. While recent deep learning approaches have improved the state of the art, they often struggle to capture architectural reasoning: the precedence of topological relationships over geometric instantiation, the propagation of functional constraints through adjacency networks, and the emergence of circulation patterns from local connectivity decisions. To address these fundamental challenges, this paper introduces GFLAN, a generative framework that restructures floor plan synthesis through explicit factorization into topological planning and geometric realization. Given a single exterior boundary and a front-door location, our approach departs from direct pixel-to-pixel or wall-tracing generation in favor of a principled two-stage decomposition. Stage A employs a specialized convolutional architecture with dual encoders -- separating invariant spatial context from evolving layout state -- to sequentially allocate room centroids within the building envelope via discrete probability maps over feasible placements. Stage B constructs a heterogeneous graph linking room nodes to boundary vertices, then applies a Transformer-augmented graph neural network (GNN) that jointly regresses room boundaries.

cs.CV↗

PRISM: Differentiable Analysis-by-Synthesis for Fixel Recovery in Diffusion MRI

Diffusion MRI microstructure fitting is nonconvex and often performed voxelwise, which limits fiber peak recovery in narrow crossings. This work introduces PRISM, a differentiable analysis-by-synthesis framework that fits an explicit multi-compartment forward model end-to-end over spatial patches. The model combines cerebrospinal fluid (CSF), gray matter, up to K white-matter fiber compartments (stick-and-zeppelin), and a restricted compartment, with explicit fiber directions and soft model selection via repulsion and sparsity priors. PRISM supports a fast MSE objective and a Rician negative log-likelihood (NLL) that jointly learns sigma without oracle information. A lightweight nuisance calibration module (smooth bias field and per-measurement scale/offset) is included for robustness and regularized to identity in clean-data tests. On synthetic crossing-fiber data (SNR=30; five methods, 16 crossing angles), PRISM achieves 3.5 degrees best-match angular error with 95% recall, which is 1.9x lower than the best baseline (MSMT-CSD, 6.8 degrees, 83% recall); in NLL mode with learned sigma, error drops to 2.3 degrees with 99% recall, resolving crossings down to 20 degrees. On the DiSCo1 phantom (NLL mode), PRISM improves connectivity correlation over CSD baselines at all four tracking angles (best r=.934 at 25 degrees vs. .920 for MSMT-CSD). Whole-brain HCP fitting (~741k voxels, MSE mode) completes in ~12 min on a single GPU with near-identical results across random seeds.

cs.CV↗

Spherical Hermite Maps

Spherical functions appear throughout computer graphics, from spherical harmonic lighting and precomputed radiance transfer to neural radiance fields and procedural planet rendering. Efficient evaluation is critical for real-time applications, yet existing approaches face a quality-performance trade-off: bilinear LUT sampling is fast but produces faceting, while bicubic filtering requires 16 texture samples. Most implementations use finite differences for normals, requiring extra samples and introducing noise. This paper presents Spherical Hermite Maps, a derivative-augmented LUT representation that resolves this trade-off. By storing function values alongside scaled partial derivatives at each texel of a padded cubemap, bicubic-Hermite reconstruction is enabled from only four texture samples (a 2x2 footprint) while providing continuous gradients from the same samples. The key insight is that Hermite interpolation reconstructs smooth derivatives as a byproduct of value reconstruction, making surface normals effectively free. In controlled experiments, Spherical Hermite Maps improve PSNR by 8-41 dB over bilinear interpolation and match 16-tap bicubic quality at one-quarter the cost. Analytic normals reduce mean angular error by 9-13% on complex surfaces while yielding stable specular highlights. Three applications demonstrate versatility: spherical harmonic glyph visualization, radial depth-map impostors for mesh level-of-detail, and procedural planet/asteroid rendering with spherical heightfields.

cs.GR↗