SearcharxivSearch

arXiv subjects

Francesco Capuano

Publications and source records attributed to Francesco Capuano.

15 recordsLinked to original sources

ORCA: A Platform for Open-Source Dexterity Research

Robotics manipulation research increasingly focuses on two-finger parallel grippers for their effectiveness, affordability, and ease of teleoperation. Grippers are nonetheless limited by their form factor, often requiring bimanual setups even for simple reorientation tasks. Anthropomorphic hands are a more natural platform for dexterous robot learning -- closer to the human hand, and capable of learning from human video -- yet they remain hard to use in learning research: even where open and accessible hand hardware exists, the software for control, simulation, teleoperation, and retargeting is scattered in one-off code bases, and largely disconnected from the robot-learning ecosystem. In this work, we introduce the \orca~learning stack, an open-source research stack for dexterity as a first-class robot learning domain. Our \orca~stack unifies low-level control, simulation, teleoperation from a range of consumer platforms, and hand retargeting, behind a single interface, and integrates natively with popular robot-learning frameworks such as \lerobot, so dexterous hand researchers can leverage the same data, training, and evaluation pipelines used for non-dexterous robot learning. We demonstrate a complete end-to-end workflow, collecting expert demonstrations of an in-hand reorientation task by teleoperation with a consumer-grade VR headset, training an autonomous policy with \lerobot, and evaluating the learned policy in a fully reproducible and observable setup. We open-source the entire stack as a shared, reproducible foundation for dexterous-manipulation research.

cs.RO

stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation

World models are central to building agents that can reason, plan, and generalize beyond their training data. However, research on world models is currently fragmented, with disparate codebases, data pipelines, and evaluation protocols hindering reproducibility and fair comparison. Current practice is further limited by three key bottlenecks: fragile one-off codebases, slow video data loading, and the lack of standardized generalization benchmarks. We present stable-worldmodel (swm), an open-source platform for standardized and reproducible world modeling research and evaluation. It delivers (1) a high-performance Lance-based data layer with native support and conversion tools for MP4, HDF5, and LeRobot datasets, (2) clean, well-tested implementations of modern world model baselines and planning solvers, and (3) a broad suite of environments and tasks extended with controllable visual, geometric, and physical factors of variation for systematic in-silico evaluation of dynamics understanding, control performance, representation quality, and out-of-distribution generalization. By unifying the full pipeline under a single, scalable framework, \texttt{swm} dramatically reduces research overhead and accelerates trustworthy progress toward reliable world models.

cs.LG

LeRobot: An Open-Source Library for End-to-End Robot Learning

Robotics is undergoing a significant transformation powered by advances in high-level control techniques based on machine learning, giving rise to the field of robot learning. Recent progress in robot learning has been accelerated by the increasing availability of affordable teleoperation systems, large-scale openly available datasets, and scalable learning-based methods. However, development in the field of robot learning is often slowed by fragmented, closed-source tools designed to only address specific sub-components within the robotics stack. In this paper, we present \texttt{lerobot}, an open-source library that integrates across the entire robot learning stack, from low-level middleware communication for motor controls to large-scale dataset collection, storage and streaming. The library is designed with a strong focus on real-world robotics, supporting accessible hardware platforms while remaining extensible to new embodiments. It also supports efficient implementations for various state-of-the-art robot learning algorithms from multiple prominent paradigms, as well as a generalized asynchronous inference stack. Unlike traditional pipelines which heavily rely on hand-crafted techniques, \texttt{lerobot} emphasizes scalable learning approaches that improve directly with more data and compute. Designed for accessibility, scalability, and openness, \texttt{lerobot} lowers the barrier to entry for researchers and practitioners to robotics while providing a platform for reproducible, state-of-the-art robot learning.

cs.RO

Robot Learning: A Tutorial

Robot learning is at an inflection point, driven by rapid advancements in machine learning and the growing availability of large-scale robotics data. This shift from classical, model-based methods to data-driven, learning-based paradigms is unlocking unprecedented capabilities in autonomous systems. This tutorial navigates the landscape of modern robot learning, charting a course from the foundational principles of Reinforcement Learning and Behavioral Cloning to generalist, language-conditioned models capable of operating across diverse tasks and even robot embodiments. This work is intended as a guide for researchers and practitioners, and our goal is to equip the reader with the conceptual understanding and practical tools necessary to contribute to developments in robot learning, with ready-to-use examples implemented in $\texttt{lerobot}$.

cs.RO

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Vision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive--often with billions of parameters--leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10x larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic benchmarks and release all code, pretrained models, and training data.

cs.LG

A theoretical framework for flow-compatible reconstruction of heart motion

Accurate three-dimensional (3D) reconstruction of cardiac chamber motion from time-resolved medical imaging modalities is of growing interest in both the clinical and biomechanical fields. Despite recent advancement, the cardiac motion reconstruction process remains complex and prone to uncertainties. Moreover, traditional assessments often focus on static comparisons, lacking assurances of dynamic consistency and physical relevance. This work introduces a novel paradigm of flow-compatible motion reconstruction, integrating anatomical imaging with flow data to ensure adherence to fundamental physical principles, such as mass and momentum conservation. The approach is demonstrated in the context of right ventricular motion, utilizing diffeomorphic mappings and multi-slice MRI to achieve dynamically consistent and physically robust reconstructions. Results show that enforcing flow compatibility within the reconstruction process is feasible and enhances the physical realism of the resulting kinematics.

physics.med-ph

Searching on a Budget: HW-NAS with 10 Latency Probes

Existing hardware-aware NAS (HW-NAS) methods typically assume access to precise information circa the target device, either via analytical approximations of the post-compilation latency model, or through learned latency predictors. Such approximate approaches risk introducing estimation errors that may prove detrimental in risk-sensitive applications. In this work, we propose a two-stage HW-NAS framework, in which we first learn an architecture controller on a distribution of synthetic devices, and then directly deploy the controller on a target device. At test-time, our network controller deploys directly to the target device without relying on any pre-collected information, and only exploits direct interactions. In particular, the pre-training phase on synthetic devices enables the controller to design an architecture for the target device by interacting with it through a small number of high-fidelity latency measurements. To guarantee accessibility of our method, we only train our controller with training-free accuracy proxies, allowing us to scale the meta-training phase without incurring the overhead of full network training. We benchmark on HW-NATS-Bench, demonstrating that our method generalizes to unseen devices and searches for latency-efficient architectures by in-context adaptation using only a few real-world latency evaluations at test-time.

cs.LG

Shaping Laser Pulses with Reinforcement Learning

High Power Laser (HPL) systems operate in the attoseconds regime -- the shortest timescale ever created by humanity. HPL systems are instrumental in high-energy physics, leveraging ultra-short impulse durations to yield extremely high intensities, which are essential for both practical applications and theoretical advancements in light-matter interactions. Traditionally, the parameters regulating HPL optical performance have been manually tuned by human experts, or optimized using black-box methods that can be computationally demanding. Critically, black box methods rely on stationarity assumptions overlooking complex dynamics in high-energy physics and day-to-day changes in real-world experimental settings, and thus need to be often restarted. Deep Reinforcement Learning (DRL) offers a promising alternative by enabling sequential decision making in non-static settings. This work explores the feasibility of applying DRL to HPL systems, extending the current research by (1) learning a control policy relying solely on non-destructive image observations obtained from readily available diagnostic devices, and (2) retaining performance when the underlying dynamics vary. We evaluate our method across various test dynamics, and observe that DRL effectively enables cross-domain adaptability, coping with dynamics' fluctuations while achieving 90\% of the target intensity in test environments.

cs.LG

Linear stability analysis of wall-bounded high-pressure transcritical fluids

Mixing and heat transfer rates are typically enhanced when operating at high-pressure transcritical turbulent flow regimes. The rapid variation of thermophysical properties in the vicinity of the pseudo-boiling region can be leveraged to significantly increase the Reynolds numbers and destabilize the flow. The underlying physical mechanism responsible for this destabilization is the presence of a baroclinic torque mainly driven by large localized density gradients across the pseudo-boiling line. As a result, the enstrophy levels are enhanced compared to equivalent low-pressure cases, and the flow physics behavior deviates from standard wall turbulence characteristics. In this work, the nature of this instability is carefully analyzed and characterized by means of linear stability theory. It is found that, at isothermal wall-bounded transcritical conditions, the non-linear thermodynamics exhibited near the pseudo-boiling region propitiates the laminar-to-turbulent transition with respect to sub- and super-critical thermodynamic states. This transition is further exacerbated for non-isothermal flows even at low Brinkman numbers. Particularly, neutral curve sensitivity to Brinkman numbers and perturbation profiles of dynamic and thermodynamic unstable modes based on modal and non-modal analysis, which trigger the early flow destabilization, confirm this phenomenon. Nonetheless, a non-isothermal setup is a necessary condition for transition when operating at low-Mach/Reynolds-number regimes. In detail, on equal Brinkman number, turbulence transition is accelerated and algebraic growth enhanced in comparison to isothermal cases. Consequently, high-pressure transcritical setups result in larger kinetic energy budgets due to larger production rates and lower viscous dissipation.

physics.flu-dyn

On the performances of standard and kinetic energy preserving time-integration methods for incompressible-flow simulations

The effects of kinetic-energy preservation errors due to Runge-Kutta (RK) temporal integrators have been analyzed for the case of large-eddy simulations of incompressible turbulent channel flow. Simulations have been run using the open-source solver Xcompact3D with an implicit spectral vanishing viscosity model and a variety of temporal Runge-Kutta integrators. Explicit pseudo-symplectic schemes, with improved energy preservation properties, have been compared to standard RK methods. The results show a marked decrease in the temporal error for higher-order pseudo-symplectic methods; on the other hand, an analysis of the energy spectra indicates that the dissipation introduced by the commonly used three-stage RK scheme can lead to significant distortion of the energy distribution within the inertial range. A cost-vs-accuracy analysis suggests that pseudo-symplectic schemes could be used to attain results comparable to traditional methods at a reduced computational cost.

physics.flu-dyn

TempoRL: laser pulse temporal shape optimization with Deep Reinforcement Learning

High Power Laser's (HPL) optimal performance is essential for the success of a wide variety of experimental tasks related to light-matter interactions. Traditionally, HPL parameters are optimised in an automated fashion relying on black-box numerical methods. However, these can be demanding in terms of computational resources and usually disregard transient and complex dynamics. Model-free Deep Reinforcement Learning (DRL) offers a promising alternative framework for optimising HPL performance since it allows to tune the control parameters as a function of system states subject to nonlinear temporal dynamics without requiring an explicit dynamics model of those. Furthermore, DRL aims to find an optimal control policy rather than a static parameter configuration, particularly suitable for dynamic processes involving sequential decision-making. This is particularly relevant as laser systems are typically characterised by dynamic rather than static traits. Hence the need for a strategy to choose the control applied based on the current context instead of one single optimal control configuration. This paper investigates the potential of DRL in improving the efficiency and safety of HPL control systems. We apply this technique to optimise the temporal profile of laser pulses in the L1 pump laser hosted at the ELI Beamlines facility. We show how to adapt DRL to the setting of spectral phase control by solely tuning dispersion coefficients of the spectral phase and reaching pulses similar to transform limited with full-width at half-maximum (FWHM) of ca1.6 ps.

physics.optics

Laser Pulse Duration Optimization With Numerical Methods

In this study we explore the optimization of laser pulse duration to obtain the shortest possible pulse. We do this by employing a feedback loop between a pulse shaper and pulse duration measurements. We apply to this problem several iterative algorithms including gradient descent, Bayesian Optimization and genetic algorithms, using a simulation of the actual laser represented via a semi-physical model of the laser based on the process of linear and non-linear phase accumulation.

physics.comp-ph

Computational Study of Pulmonary Flow Patterns after Repair of Transposition of Great Arteries

Patients that undergo the arterial switch operation (ASO) to repair transposition of great arteries (TGA) can develop abnormal pulmonary trunk morphology with significant long-term complications. In this study, cardiovascular magnetic resonance was combined with computational fluid dynamics to investigate the impact of the post-operative layout on the pulmonary flow patterns. Three ASO patients were analyzed and compared to a normal control. Results showed the presence of anomalous shear layer instabilities, vortical and helical structures, and turbulent-like states in all patients, particularly as a consequence of the unnatural curvature of the pulmonary bifurcation. Streamlined, mostly laminar flow was instead found in the healthy subject. These findings shed light on the correlation between the post-ASO anatomy and the presence of altered flow features, and may be useful to improve surgical planning as well as the long-term care of TGA patients.

physics.med-ph

Numerically stable formulations of convective terms for turbulent compressible flows

A systematic analysis of the discrete conservation properties of non-dissipative, central-difference approximations of the compressible Navier-Stokes equations is reported. A general triple splitting of the nonlinear convective terms is considered, and energy-preserving formulations are fully characterized by deriving a two-parameter family of split forms. Previously developed formulations reported in literature are shown to be particular members of this family; novel splittings are introduced and discussed as well. Furthermore, the conservation properties yielded by different choices for the energy equation (i.e. total and internal energy, entropy) are analyzed thoroughly. It is shown that additional preserved quantities can be obtained through a suitable adaptive selection of the split form within the derived family. Local conservation of primary invariants, which is a fundamental property to build high-fidelity shock-capturing methods, is also discussed in the paper. Numerical tests performed for the Taylor-Green Vortex at zero viscosity fully confirm the theoretical findings, and show that a careful choice of both the splitting and the energy formulation can provide remarkably robust and accurate results.

physics.flu-dyn

Effects of discrete energy and helicity conservation in numerical simulations of helical turbulence

Helicity is the scalar product between velocity and vorticity and, just like energy, its integral is an in-viscid invariant of the three-dimensional incompressible Navier-Stokes equations. However, space-and time-discretization methods typically corrupt this property, leading to violation of the inviscid conservation principles. This work investigates the discrete helicity conservation properties of spectral and finite-differencing methods, in relation to the form employed for the convective term. Effects due to Runge-Kutta time-advancement schemes are also taken into consideration in the analysis. The theoretical results are proved against inviscid numerical simulations, while a scale-dependent analysis of energy, helicity and their non-linear transfers is performed to further characterize the discretization errors of the different forms in forced helical turbulence simulations.

math.NA