SearcharxivSearch

arXiv subjects

Mihir Kulkarni

Publications and source records attributed to Mihir Kulkarni.

At least 19 recordsLinked to original sources

Abundant Heavy Black Hole Seeds from Moderate Lyman-Werner Radiation

The existence of high-redshift quasars may indicate that massive black hole seeds formed via supermassive Population III stars in atomic-cooling halos with large gas inflow rates; however, the dependence of this process on halo assembly rate and radiative background remains poorly constrained. We present a large suite of 65 high-resolution cosmological zoom-in simulations of 15 pristine halos spanning a wide range of Lyman-Werner radiation backgrounds and halo assembly histories. We introduce a novel method to estimate the final Population III stellar mass from radial gas infall profiles at the onset of runaway collapse and validate it against simulations from the literature that explicitly follow protostellar accretion with sink particles, reproducing protostellar masses to within a factor of $\sim2$. We find a clear transition in gas inflow rates between halos exposed to $J_{\rm 21} \lesssim 1$ and $J_{\rm 21} \gtrsim 10$, with the latter frequently sustaining inflow rates above the adopted threshold for supermassive star formation and producing estimated stellar masses up to $10^{5} \, M_{\odot}$. In contrast, the halo assembly timescale, $M_{\rm Halo}$/$\dot{M}_{\rm Halo}$, shows no statistically significant correlation with predicted stellar mass, despite halo assembly rates spanning $0.01$-$7 \, M_{\rm \odot} \, {\rm yr}^{-1}$. The Lyman-Werner radiation field therefore is a stronger predictor of sustained high accretion within our parameter space. Finally, a semi-analytic model applied to cosmological volumes shows that halos exposed to intermediate Lyman-Werner backgrounds ($1 \lesssim J_{\rm 21} < 10$) are orders of magnitude more common than those in the high-$J_{\rm 21}$ tail. If sustained high accretion extends into this intermediate regime, heavy black hole seeds may form in substantially more common environments than required by classical direct-collapse scenarios.

astro-ph.GA

Embodiment-conditioned Generalist Control for Multirotor Aerial Robots

We present a generalist position control policy capable of controlling arbitrary multirotor configurations of a certain rotor count (e.g., hexarotors or quadrotors) with a single set of network weights. The policy is conditioned on a physics-grounded embodiment descriptor: a mass and inertia-normalized control allocation matrix that captures how mass-normalized motor thrusts generate linear and angular accelerations in the body-frame. To train the policy, we sample from a broad distribution of arbitrary multirotor configurations, including non-planar and asymmetric systems, and optimize a single, compact network using Proximal Policy Optimization. Training requires only five minutes on an RTX 3090 GPU using a custom NVIDIA Warp-based dynamics simulator. Through extensive simulation experiments, we show that embodiment conditioning enables robust generalist control across arbitrary morphologies. We demonstrate zero-shot real-world transfer of this generalist policy on three diverse hexarotor systems, including a planar robot, a partially symmetric non-planar system, and a random asymmetric, non-planar configuration.

cs.RO

The Unified Autonomy Stack: Toward a Blueprint for Generalizable Robot Autonomy

We introduce and open-source the Unified Autonomy Stack, a system-level solution that enables resilient autonomy across diverse aerial and ground robot morphologies. The architecture centers on three synergistic modules -- multi-modal perception, multi-behavior planning, and multi-layered safe navigation -- that together deliver comprehensive mission autonomy. The stack fuses data from LiDAR, radar, vision, and inertial sensing, enabling (a) robust localization and mapping through factor graph-based fusion, (b) semantic scene understanding, (c) motion and informative path planning through sampling-based techniques adaptive across spatial scales, as well as (d) multi-layered safe navigation both through planning on the online reconstructed map and deep learning-driven exteroceptive policies alongside last-resort safety filters using control barrier functions. The resulting behaviors include safe GNSS-denied navigation into unknown and perceptually-degraded regions, exploration of complex environments, object discovery, and efficient inspection planning. The stack has been field-tested and validated on both aerial (rotorcraft) and ground (legged) robots operating in a host of demanding environments, including self-similar and smoke-filled settings, with complex geometries and high obstacle clutter. These tests demonstrate resilient performance in challenging conditions. To facilitate ease of adoption, we open-source the implementation alongside supporting documentation, validation, and evaluation datasets https://github.com/ntnu-arl/unified_autonomy_stack. A video giving the overview of the paper and the field experiments is available at https://youtu.be/l8Su8OXsM-E.

cs.RO

Efficient Knowledge Transfer for Jump-Starting Control Policy Learning of Multirotors through Physics-Aware Neural Architectures

Efficiently training control policies for robots is a major challenge that can greatly benefit from utilizing knowledge gained from training similar systems through cross-embodiment knowledge transfer. In this work, we focus on accelerating policy training using a library-based initialization scheme that enables effective knowledge transfer across multirotor configurations. By leveraging a physics-aware neural control architecture that combines a reinforcement learning-based controller and a supervised control allocation network, we enable the reuse of previously trained policies. To this end, we utilize a policy evaluation-based similarity measure that identifies suitable policies for initialization from a library. We demonstrate that this measure correlates with the reduction in environment interactions needed to reach target performance and is therefore suited for initialization. Extensive simulation and real-world experiments confirm that our control architecture achieves state-of-the-art control performance, and that our initialization scheme saves on average up to $73.5\%$ of environment interactions (compared to training a policy from scratch) across diverse quadrotor and hexarotor designs, paving the way for efficient cross-embodiment transfer in reinforcement learning.

cs.RO

Reinforcement Learning for Active Perception in Autonomous Navigation

This paper addresses the challenge of active perception within autonomous navigation in complex, unknown environments. Revisiting the foundational principles of active perception, we introduce an end-to-end reinforcement learning framework in which a robot must not only reach a goal while avoiding obstacles, but also actively control its onboard camera to enhance situational awareness. The policy receives observations comprising the robot state, the current depth frame, and a particularly local geometry representation built from a short history of depth readings. To couple collision-free motion planning with information-driven active camera control, we augment the navigation reward with a voxel-based information metric. This enables an aerial robot to learn a robust policy that balances goal-directed motion with exploratory sensing. Extensive evaluation demonstrates that our strategy achieves safer flight compared to using fixed, non-actuated camera baselines while also inducing intrinsic exploratory behaviors.

cs.RO

Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning

We present Isaac Lab, the natural successor to Isaac Gym, which extends the paradigm of GPU-native robotics simulation into the era of large-scale multi-modal learning. Isaac Lab combines high-fidelity GPU parallel physics, photorealistic rendering, and a modular, composable architecture for designing environments and training robot policies. Beyond physics and rendering, the framework integrates actuator models, multi-frequency sensor simulation, data collection pipelines, and domain randomization tools, unifying best practices for reinforcement and imitation learning at scale within a single extensible platform. We highlight its application to a diverse set of challenges, including whole-body control, cross-embodiment mobility, contact-rich and dexterous manipulation, and the integration of human demonstrations for skill acquisition. Finally, we discuss upcoming integration with the differentiable, GPU-accelerated Newton physics engine, which promises new opportunities for scalable, data-efficient, and gradient-based approaches to robot learning. We believe Isaac Lab's combination of advanced simulation capabilities, rich sensing, and data-center scale execution will help unlock the next generation of breakthroughs in robotics research.

cs.RO

From Primordial Stars to Early Galaxies: A Semi-Analytic Model Calibrated with Aeos and Renaissance

We present an extension of our semi-analytic model that follows the formation of Population III stars and their metal-enriched descendants, incorporating dark matter halo merger trees from cosmological $N$-body simulations and feedback from reionization. Our extended model is calibrated using two complementary cosmological hydrodynamical simulations: Aeos, which resolves individual Population III and II stars to $z\sim14.6$, and Renaissance, which is lower resolution but follows large-scale metal-enriched star formation to $z \sim 11$. With a combined calibration, we capture small-scale physics of primordial star formation over a large range in halo mass. We find good agreement between our calibrated model and Aeos, reproducing the evolution in number of star-forming halos and total stellar mass. Achieving this agreement requires increasing the normalization of, flattening the redshift dependence of, and adding scatter to the commonly used critical mass threshold $M_{\mathrm{crit}}$. Our treatment of the delay between Pop III stellar death and subsequent Pop II star formation emphasizes the need to account for halos that have yet to transition to Pop II, since incomplete sampling of this delay in simulations limits physically motivated calibrations. Finally, we apply our model to larger-volume dark matter only simulations and predict $\sim10$ active Pop III sources at $z = 10$ lie within the area strongly lensed by galaxy cluster MACS J0416 with a magnification exceeding $\mu > 30$. These results demonstrate that semi-analytic approaches, when calibrated to hydrodynamical simulations, can provide accurate, computationally efficient predictions for the earliest stages of cosmic star formation.

astro-ph.GA

Performance-guided Task-specific Optimization for Multirotor Design

This paper introduces a methodology for task-specific design optimization of multirotor Micro Aerial Vehicles. By leveraging reinforcement learning, Bayesian optimization, and covariance matrix adaptation evolution strategy, we optimize aerial robot designs guided exclusively by their closed-loop performance in a considered task. Our approach systematically explores the design space of motor pose configurations while ensuring manufacturability constraints and minimal aerodynamic interference. Results demonstrate that optimized designs achieve superior performance compared to conventional multirotor configurations in agile waypoint navigation tasks, including against fully actuated designs from the literature. We build and test one of the optimized designs in the real world to validate the sim2real transferability of our approach.

cs.RO

UniPilot: Enabling GPS-Denied Autonomy Across Embodiments

This paper presents UniPilot, a compact hardware-software autonomy payload that can be integrated across diverse robot embodiments to enable autonomous operation in GPS-denied environments. The system integrates a multi-modal sensing suite including LiDAR, radar, vision, and inertial sensing for robust operation in conditions where uni-modal approaches may fail. UniPilot runs a complete autonomy software comprising multi-modal perception, exploration and inspection path planning, and learning-based navigation policies. The payload provides robust localization, mapping, planning, and safety and control capabilities in a single unit that can be deployed across a wide range of platforms. A large number of experiments are conducted across diverse environments and on a variety of robot platforms to validate the mapping, planning, and safe navigation capabilities enabled by the payload.

cs.RO

Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability

Vision-Language Models (VLMs) have shown remarkable performance in integrating visual and textual information for tasks such as image captioning and visual question answering. However, these models struggle with compositional generalization and object binding, which limit their ability to handle novel combinations of objects and their attributes. Our work explores the root causes of these failures using mechanistic interpretability techniques. We show evidence that individual neurons in the MLP layers of CLIP's vision encoder represent multiple features, and this "superposition" directly hinders its compositional feature representation which consequently affects compositional reasoning and object binding capabilities. We hope this study will serve as an initial step toward uncovering the mechanistic roots of compositional failures in VLMs. The code and supporting results can be found https://github.com/Mystic-Slice/Do-VLMs-Have-Bad-Eyes.

cs.CV

Evaluating the Accuracy of Reionization Prescriptions in Semi-analytic Models of the First Stars and Galaxies

Semi-analytic models are a valuable tool to study the first stars and galaxies. Their numerical efficiency makes it possible to survey broad regions of astrophysical parameter space across large volumes and redshift ranges. Following reionization in these models is necessary since star formation is suppressed in ionized regions due to photoheating of the gas. Here we evaluate the accuracy of three semi-analytic reionization prescriptions (two previously developed and one new model) by comparing their three-dimensional distribution of ionized bubbles to the Renaissance hydrodynamical cosmological radiative transfer simulations. We find that the previously existing models accurately determine the distribution of the larger bubbles within our ${\sim}6$ comoving Mpc simulation box, but that these models fail to take into account self-shielded neutral gas in dense filaments. Thus, these prescriptions overestimate the fraction of halos in HII regions impacted by reionization feedback by up to an order of magnitude (depending on halo mass and redshift). This leads to an unrealistically large effect of reionization feedback on Pop III stars and low-mass metal-enriched galaxies. Our newly developed model takes into account the density structure of the cosmic web, leading to good agreement with Renaissance in the fraction of halos found in ionized regions.

astro-ph.GA

Semantically-driven Deep Reinforcement Learning for Inspection Path Planning

This paper introduces a novel semantics-aware inspection planning policy derived through deep reinforcement learning. Reflecting the fact that within autonomous informative path planning missions in unknown environments, it is often only a sparse set of objects of interest that need to be inspected, the method contributes an end-to-end policy that simultaneously performs semantic object visual inspection combined with collision-free navigation. Assuming access only to the instantaneous depth map, the associated segmentation image, the ego-centric local occupancy, and the history of past positions in the robot's neighborhood, the method demonstrates robust generalizability and successful crossing of the sim2real gap. Beyond simulations and extensive comparison studies, the approach is verified in experimental evaluations onboard a flying robot deployed in novel environments with previously unseen semantics and overall geometric configurations.

cs.RO

A Neural Network Mode for PX4 on Embedded Flight Controllers

This paper contributes an open-sourced implementation of a neural-network based controller framework within the PX4 stack. We develop a custom module for inference on the microcontroller while retaining all of the functionality of the PX4 autopilot. Policies trained in the Aerial Gym Simulator are converted to the TensorFlow Lite format and then built together with PX4 and flashed to the flight controller. The policies substitute the control-cascade within PX4 to offer an end-to-end position-setpoint tracking controller directly providing normalized motor RPM setpoints. Experiments conducted in simulation and the real-world show similar tracking performance. We thus provide a flight-ready pipeline for testing neural control policies in the real world. The pipeline simplifies the deployment of neural networks on embedded flight controller hardware thereby accelerating research on learning-based control. Both the Aerial Gym Simulator and the PX4 module are open-sourced at https://github.com/ntnu-arl/aerial_gym_simulator and https://github.com/SindreMHegre/PX4-Autopilot-public/tree/for_paper. Video: https://youtu.be/lY1OKz_UOqM?si=VtzL243BAY3lblTJ.

cs.RO

MapQA: Open-domain Geospatial Question Answering on Map Data

Geospatial question answering (QA) is a fundamental task in navigation and point of interest (POI) searches. While existing geospatial QA datasets exist, they are limited in both scale and diversity, often relying solely on textual descriptions of geo-entities without considering their geometries. A major challenge in scaling geospatial QA datasets for reasoning lies in the complexity of geospatial relationships, which require integrating spatial structures, topological dependencies, and multi-hop reasoning capabilities that most text-based QA datasets lack. To address these limitations, we introduce MapQA, a novel dataset that not only provides question-answer pairs but also includes the geometries of geo-entities referenced in the questions. MapQA is constructed using SQL query templates to extract question-answer pairs from OpenStreetMap (OSM) for two study regions: Southern California and Illinois. It consists of 3,154 QA pairs spanning nine question types that require geospatial reasoning, such as neighborhood inference and geo-entity type identification. Compared to existing datasets, MapQA expands both the number and diversity of geospatial question types. We explore two approaches to tackle this challenge: (1) a retrieval-based language model that ranks candidate geo-entities by embedding similarity, and (2) a large language model (LLM) that generates SQL queries from natural language questions and geo-entity attributes, which are then executed against an OSM database. Our findings indicate that retrieval-based methods effectively capture concepts like closeness and direction but struggle with questions that require explicit computations (e.g., distance calculations). LLMs (e.g., GPT and Gemini) excel at generating SQL queries for one-hop reasoning but face challenges with multi-hop reasoning, highlighting a key bottleneck in advancing geospatial QA systems.

cs.CL

Coordinated Ramp Metering Control based on Scalable Nonlinear Traffic Dynamics Model Discovery in a Large Network

This study proposes a coordinated ramp metering control framework in large networks based on scalable nonlinear traffic dynamics model discovery. Existing coordinated ramp metering control methods often require accurate traffic dynamics models in real time, however, for large-scale highway networks, since these models are always nonlinear, they are extremely challenging to obtain. To overcome this limitation, this study utilizes the Sparse Identification of Nonlinear Dynamics with Control (SINDYc) to derive the accurate nonlinear traffic dynamics model from observed data. The discovered dynamics model is then integrated into a Model Predictive Control (MPC) coordinated ramp metering controller, enabling optimized control actions that enhance traffic flow and efficiency. The proposed framework is tested on a large-scale highway network that includes three intersecting highways and eight on-ramps, which outperforms the existing approaches, demonstrating its effectiveness and potential for real-time application. This framework can offer a scalable and robust solution for improving real-time traffic management in complex urban environments.

eess.SY

Aerial Gym Simulator: A Framework for Highly Parallelized Simulation of Aerial Robots

This paper contributes the Aerial Gym Simulator, a highly parallelized, modular framework for simulation and rendering of arbitrary multirotor platforms based on NVIDIA Isaac Gym. Aerial Gym supports the simulation of under-, fully- and over-actuated multirotors offering parallelized geometric controllers, alongside a custom GPU-accelerated rendering framework for ray-casting capable of capturing depth, segmentation and vertex-level annotations from the environment. Multiple examples for key tasks, such as depth-based navigation through reinforcement learning are provided. The comprehensive set of tools developed within the framework makes it a powerful resource for research on learning for control, planning, and navigation using state information as well as exteroceptive sensor observations. Extensive simulation studies are conducted and successful sim2real transfer of trained policies is demonstrated. The Aerial Gym Simulator is open-sourced at: https://github.com/ntnu-arl/aerial_gym_simulator.

cs.RO

Radiative Transfer Simulations of Ly$\alpha$ Intensity Mapping During Cosmic Reionization Including Sources from Galaxies and the Intergalactic Medium

We present new simulations of Lyman-$\alpha$ (Ly$\alpha$) intensity maps that include Ly$\alpha$ radiative transfer in the intergalactic medium (IGM) and all significant sources of Ly$\alpha$ photons. The sources considered include Ly$\alpha$ directly from galaxies, cooling at the edges of ionized bubbles, recombinations within these bubbles, and reprocessing of galaxy continuum emission in the IGM. We also vary astrophysical parameters including the average neutral fraction of the IGM, the dust absorption of Ly$\alpha$ in galaxies, and the ionizing escape fraction. Previous work has suggested that Ly$\alpha$ intensity mapping can be used to constrain the neutral fraction of the IGM when accounting for radiative transfer in the IGM. When radiative transfer is ignored, direct Ly$\alpha$ emission from galaxies has the highest amplitude of power on all scales. When we include radiative transfer in our simulations, we find continuum emission reprocessed as Ly$\alpha$ is comparable to the Ly$\alpha$ emission directly from galaxies on large scales. For high neutral fraction in the IGM, emission from recombinations is comparable to galaxies on large scales. We find that the slope of the power spectrum is sensitive to the neutral fraction of the IGM when radiative transfer is included, suggesting that this may be useful for placing constraints on cosmic reionization. In addition, we find the power of galaxies is decreased across all scales due to dust absorption. We also find the escape fraction must be large for recombinations and bubble edges to contribute significantly to the power. We find the cross power is observable between SPHEREx Ly$\alpha$ intensity maps and a hypothetical galaxy survey is observable with a total signal-to-noise of 4 from $k = 0.035$ Mpc$^{-1}$ to $k = 1$ Mpc$^{-1}$.

astro-ph.GA

Can supermassive stars form in protogalaxies due to internal Lyman-Werner feedback?

Population III stars are possible precursors to early massive and supermassive black holes (BHs). The presence of soft UV Lyman Werner (LW) background radiation can suppress Population III star formation in minihalos and allow them to form in pristine atomic cooling halos. In the absence of molecular hydrogen ($\rm H_2$) cooling, atomic-cooling halos enable rapid collapse with suppressed fragmentation. High background LW fluxes from preceding star-formation have been proposed to dissociate $\rm H_2$. This flux can be supplemented by LW radiation from one or more Population III star(s) in the same halo, reducing the necessary background level. Here we consider atomic-cooling halos in which multiple protostellar cores form close to one another nearly simultaneously. We assess whether the first star's LW radiation can dissociate nearby $\rm H_2$, enabling the prompt formation of a second, supermassive star (SMS) from warm, atomically-cooled gas. We use a set of hydrodynamical simulations with the code ENZO, with identical LW backgrounds centered on a halo with two adjacent collapsing gas clumps. When an additional large local LW flux is introduced, we observe immediate reductions in both the accretion rates and the stellar masses that form within these clumps. While the LW flux reduces the $\text{H}_2$ fraction and increases the gas temperature, the halo core's potential well is too shallow to promptly heat the gas to $\gtrsim$ 1000 K and increase the accretion rate onto the second protostar. We conclude that internal LW feedback inside atomic-cooling halos is unlikely to facilitate the formation of SMSs or massive BH seeds.

astro-ph.GA