SearcharxivSearch

arXiv subjects

Rohit Gupta

Publications and source records attributed to Rohit Gupta.

At least 19 recordsLinked to original sources

Geometric Fixed-Time Sliding Mode Control for Constrained Attitude Tracking on $\mathrm{SO}(3)$

This paper studies constrained spacecraft attitude tracking on the Riemannian configuration manifold $\mathrm{SO}(3)$ in the presence of multiple attitude pointing constraints and matched external disturbances. To address this, an attitude potential function is proposed intrinsically on $\mathrm{SO}(3)$, and its key properties are established using intrinsic geometric analysis. Under mild conditions, the potential function is shown to admit a unique nondegenerate minimum at the desired attitude over the admissible subset of $\mathrm{SO}(3)$, defined by excluding the forbidden attitude regions as well as a measure-zero set, thereby ensuring a well-posed constrained attitude tracking problem. A Riemannian Hessian analysis shows that the Hessian of the potential function is locally uniform positive definite in an open neighborhood of the desired attitude, thereby establishing local strong convexity. A nonsingular fixed-time geometric sliding manifold is proposed using the Riemannian gradient of the potential function, leading to a geometric fixed-time sliding-mode-based constrained attitude control law. It is shown that, for every initial attitude in the admissible subset, the closed-loop state trajectory evolves on $\mathrm{SO}(3)\times\mathbb{R}^3$, with the attitude remaining in the admissible subset throughout the maneuver, while the state converges to a sufficiently small compact neighborhood of the desired equilibrium in a prescribed fixed time. Numerical simulations validate the proposed control approach and illustrate the theoretical results.

eess.SY

A Cosmic Muon Tomography System with Machine Learning based Momentum Measurement for Multi-Object Reconstruction and Material Characterization

Cosmic muon tomography is a powerful non-destructive imaging technique for inspecting dense and shielded materials through multiple Coulomb scattering. In this work, we present the design, simulation, and performance evaluation of a complete muon tomography system comprising six scintillator-strip tracking stations for trajectory reconstruction and a four-station magnetic spectrometer for muon momentum estimation. The detector geometry is implemented in the GEANT4 framework and optimized for object localization and material characterization. The reconstructed momentum is combined with the scattering angle to define the scattering density $\rho_s = {(\theta p)^2}/{L_{\mathrm{eff}}}$, which enhances sensitivity to material-dependent scattering. Point-of-Closest-Approach (PoCA) reconstruction is used to estimate scattering locations within the imaging volume. To detect and separate multiple unknown objects, Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) is applied to the reconstructed PoCA cloud. Cluster-level scattering and geometric features are then extracted for object characterization. The proposed framework enables object detection, localization, volume estimation, shape reconstruction, and material ranking within a unified analysis pipeline. Simulation studies with multiple objects of different compositions demonstrate accurate reconstruction of object positions and geometries, while providing reliable material discrimination based on scattering density. The developed system offers a scalable approach for next-generation cosmic muon tomography applications in security screening, nuclear waste characterization, and non-destructive inspection.

physics.ins-det

A Passive Daytime Colored Radiative Cooler with DBR-Engineered Color Selectivity

Passive daytime colored radiative cooling (PDCRC) offers an attractive approach to thermal management while enabling structural coloration. However, achieving both strong daytime cooling and vivid, stable colors remains challenging. Here, we propose a lithography-free PDCRC comprising a polydimethylsiloxane (PDMS) emitter, SiC/SiO2 distributed Bragg reflector (DBR), MgF2 spacer, and Ag back reflector. The DBR and MgF2 spacer enable selective visible-spectrum reflection for cyan, magenta, and yellow (CMY) color generation, while the PDMS layer provides strong thermal emission within the atmospheric transparency window (8-14 micrometers). Finite-difference time-domain (FDTD) simulations, validated using the transfer matrix method (TMM), yield an average emissivity of 92% in the atmospheric window and an average reflectivity of 93.3% over 0.75-6 micrometers. CIE 1931 chromaticity analysis confirms high-contrast CMY colors with low color differences from standard subtractive primary colors. Under daytime conditions with a heat-transfer coefficient of 6 W m-2 K-1, the proposed cooler achieves cooling powers of approximately 140-145 W m-2 and a temperature reduction of about 12 degrees C below ambient. The color characteristics remain stable over a broad range of incidence angles, with only a gradual reduction in color intensity at larger angles, while the spectral response remains nearly insensitive to TE and TM polarizations. Simulations using meteorological conditions from different global cities further demonstrate robust cooling performance under diverse environmental conditions. The combination of high cooling performance, vivid coloration, simple multilayer architecture, and fabrication feasibility makes the proposed PDCRC promising for energy-efficient thermal management in buildings, vehicle coatings, wearable devices, and other outdoor applications.

physics.app-ph

Dynamics-matched Physical Reservoir Computing for Undersensed Traffic Prediction

Machine learning methods are increasingly used for traffic prediction in applications such as autonomous driving. Such predictions must be both highly accurate and immediately available, making methods with low computational costs and fast training times of interest. One such method is reservoir computing, in which the rich dynamics of a nonlinear system serves as a computational substrate and only a linear readout vector is trained. In this work we use a traffic network as the reservoir for predicting the behavior of an undersensed traffic network. This matching of the highly nonlinear dynamics allows for similar encoding between the behaviors of the reservoir and target network, enabling a more direct prediction. We show that a reservoir governed by the Improved Intelligent Driver Model (IIDM) satisfies the echo state property for a class of slowly-varying inputs. Through simulations we show that the echo state property likely holds for a larger class of inputs, and that the IIDM reservoir computer (IIDM-RC) accurately predicts an undersensed vehicle network governed by varying car-following models. We also compare with echo state networks (ESNs) and Long Short-Term Memory (LSTM) networks, finding improvements using IIDM-RC in both prediction accuracy and training time.

eess.SY

Sensitivity of p_T Fluctuations to the QCD Equation of State

We construct a novel theoretical baseline for dynamical transverse momentum correlations, $C_{pT}$, across a wide range of collision energies spanning the RHIC Beam Energy Scan (BES) program. For the first time, a unified framework is developed to describe the energy and centrality dependence of $C_{p_{\rm T}}$ from $\sqrt{s_{\text{NN}}} = 3.0$ to $200$~GeV. A central feature of this study is the implementation of Equation of State (EOS) inputs derived from Lattice QCD results at finite baryochemical potential $\mu_{\rm B}$, representing the first such application to the measured transverse momentum correlations. Despite the minimalist nature of the fluid-dynamic evolution employed, the model effectively captures the characteristic centrality scaling of the experimental data. Our results indicate that while the bulk evolution is largely governed by the EOS and system lifetime at lower energies, significant deviations in peripheral collisions at top energies highlight the onset of non-thermal correlation mechanisms. This baseline provides a necessary benchmark for interpreting transverse momentum fluctuations in heavy-ion collisions and can aid in the search for the QCD critical point.

nucl-th

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale

The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and exploration. Existing classification-based approaches operate at a coarse city-level granularity and fail to capture fine-grained details, while image retrieval methods are impractical on a global scale due to the need for extensive image galleries which are infeasible to compile. Comparatively, constructing a gallery of GPS coordinates is straightforward and inexpensive. We propose VidTAG, a dual-encoder framework that performs frame-to-GPS retrieval using both self-supervised and language-aligned features. To address temporal inconsistencies in video predictions, we introduce the TempGeo module, which aligns frame embeddings, and the GeoRefiner module, an encoder-decoder architecture that refines GPS features using the aligned frame embeddings. Evaluations on Mapillary (MSLS) and GAMa datasets demonstrate our model's ability to generate temporally consistent trajectories and outperform baselines, achieving a 20% improvement at the 1 km threshold over GeoCLIP. We also beat current State-of-the-Art by 25% on global coarse grained video geolocalization (CityGuessr68k). Our approach enables fine-grained video geolocalization and lays a strong foundation for future research. More details on the project webpage: https://parthpk.github.io/vidtag_webpage/

cs.CV

ViLL-E: Video LLM Embeddings for Retrieval

Video Large Language Models (VideoLLMs) excel at video understanding tasks where outputs are textual, such as Video Question Answering and Video Captioning. However, they underperform specialized embedding-based models in Retrieval tasks, such as Text-toVideo Retrieval and Moment Retrieval. We introduce ViLL-E (Video-LLM-Embed), a unified VideoLLM architecture endowed with a novel embedding generation mechanism that allows the model to "think longer" for complex videos and stop early for easy ones. We train this model with a three-stage training methodology combining generative and contrastive learning: initial large-scale pre-training with video-caption pairs; followed by continual training on a smaller, detailed-caption dataset; and concluding with task-specific fine-tuning on a novel multi-task dataset covering Video QA, Temporal Localization, Video Retrieval, and Video-Text Matching. Our model significantly improves temporal localization (on avg. 7% over other VideoLLMs) and video retrieval (up to 4% over dual encoder models), achieving performance comparable to state-of-the-art specialized embedding models while remaining competitive on VideoQA tasks. Furthermore, our joint contrastive-generative training unlocks new zero-shot capabilities, significantly outperforming state-of-the-art methods in composed video retrieval (+5% over SotA) and retrieval from long text (+2% over SotA).

cs.CV

Investigation of Nuclear Modification Factor from RHIC to LHC energies using Boltzmann Transport equation in conjunction with q-Weibull distribution

The study of nuclear modification factor is crucial in advancing our knowledge of the hot and dense nuclear matter created during high energy heavy-ion collision. In this direction, we have developed a theoretical model for the nuclear modification factor using the Boltzmann Transport equation in relaxation time approximation with the q-Weibull distribution as the final state distribution and studied the experimental data of nuclear modification factor of charged hadrons as well as identified particles at various energies ranging from 7.7 GeV measured at RHIC upto the maximum value of 5.44 TeV studied in LHC. We observed a good agreement between the model and the experimental data as can be quantified using the $\chi^2$/NDF values. We have also studied the mass dependence of different fit parameters that appears in the theoretical model and observe a linear mass dependence of some parameters.

hep-ph

StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales

State space models (SSMs) have emerged as a competitive alternative to transformers in various tasks. Their linear complexity and hidden-state recurrence make them particularly attractive for modeling long sequences, whereas attention becomes quadratically expensive. However, current training methods for video understanding are tailored towards transformers and fail to fully leverage the unique attributes of SSMs. For example, video models are often trained at a fixed resolution and video length to balance the quadratic scaling of attention cost against performance. Consequently, these models suffer from degraded performance when evaluated on videos with spatial and temporal resolutions unseen during training; a property we call spatio-temporal inflexibility. In the context of action recognition, this severely limits a model's ability to retain performance across both short- and long-form videos. Therefore, we propose a flexible training method that leverages and improves the inherent adaptability of SSMs. Our method samples videos at varying temporal and spatial resolutions during training and dynamically interpolates model weights to accommodate any spatio-temporal scale. This instills our SSM, which we call StretchySnake, with spatio-temporal flexibility and enables it to seamlessly handle videos ranging from short, fine-grained clips to long, complex activities. We introduce and compare five different variants of flexible training, and identify the most effective strategy for video SSMs. On short-action (UCF-101, HMDB-51) and long-action (COIN, Breakfast) benchmarks, StretchySnake outperforms transformer and SSM baselines alike by up to 28%, with strong adaptability to fine-grained actions (SSV2, Diving-48). Therefore, our method provides a simple drop-in training recipe that makes video SSMs more robust, resolution-agnostic, and efficient across diverse action recognition scenarios.

cs.CV

Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos

The recent growth in the consumption of online media by children during early childhood necessitates data-driven tools enabling educators to filter out appropriate educational content for young learners. This paper presents an approach for detecting educational content in online videos. We focus on two widely used educational content classes: literacy and math. For each class, we choose prominent codes (sub-classes) based on the Common Core Standards. For example, literacy codes include `letter names', `letter sounds', and math codes include `counting', `sorting'. We pose this as a fine-grained multilabel classification problem as videos can contain multiple types of educational content and the content classes can get visually similar (e.g., `letter names' vs `letter sounds'). We propose a novel class prototypes based supervised contrastive learning approach that can handle fine-grained samples associated with multiple labels. We learn a class prototype for each class and a loss function is employed to minimize the distances between a class prototype and the samples from the class. Similarly, distances between a class prototype and the samples from other classes are maximized. As the alignment between visual and audio cues are crucial for effective comprehension, we consider a multimodal transformer network to capture the interaction between visual and audio cues in videos while learning the embedding for videos. For evaluation, we present a dataset, APPROVE, employing educational videos from YouTube labeled with fine-grained education classes by education researchers. APPROVE consists of 193 hours of expert-annotated videos with 19 classes. The proposed approach outperforms strong baselines on APPROVE and other benchmarks such as Youtube-8M, and COIN. The dataset is available at https://github.com/rohit-gupta/MMContrast/tree/main/APPROVE

cs.CV

Cross-View Open-Vocabulary Object Detection in Aerial Imagery

Traditional object detection models are typically trained on a fixed set of classes, limiting their flexibility and making it costly to incorporate new categories. Open-vocabulary object detection addresses this limitation by enabling models to identify unseen classes without explicit training. Leveraging pretrained models contrastively trained on abundantly available ground-view image-text classification pairs provides a strong foundation for open-vocabulary object detection in aerial imagery. Domain shifts, viewpoint variations, and extreme scale differences make direct knowledge transfer across domains ineffective, requiring specialized adaptation strategies. In this paper, we propose a novel framework for adapting open-vocabulary representations from ground-view images to solve object detection in aerial imagery through structured domain alignment. The method introduces contrastive image-to-image alignment to enhance the similarity between aerial and ground-view embeddings and employs multi-instance vocabulary associations to align aerial images with text embeddings. Extensive experiments on the xView, DOTAv2, VisDrone, DIOR, and HRRSD datasets are used to validate our approach. Our open-vocabulary model achieves improvements of +6.32 mAP on DOTAv2, +4.16 mAP on VisDrone (Images), and +3.46 mAP on HRRSD in the zero-shot setting when compared to finetuned closed-vocabulary dataset-specific model performance, thus paving the way for more flexible and scalable object detection systems in aerial applications.

cs.CV

SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation

Driving scene manipulation using real-world sensor data has emerged as a promising alternative to traditional driving simulators. Despite advances in language control and neural scene representations, existing methods treat grounding, editing, and simulation as loosely connected stages, relying on heuristic object localization, manual guidance, and single-agent validation, thereby constraining semantic expressiveness and hindering scalable, reactive scenario generation. We introduce SIMSplat, a driving scene editor built on scene-graph-based 4D Gaussian Splatting augmented with language-aligned features. By embedding appearance, motion, and location semantics directly into Gaussian scene-graph nodes, SIMSplat makes reconstructed scenes queryable through free-form natural language, bridging language understanding to object-level editing and multi-agent simulation within a single framework. Building on this language-grounded scene graph, SIMSplat supports diverse edits including fine-grained pedestrian manipulation, while a multi-agent path refinement module propagates changes across all agents to ensure reactive, physically plausible simulations. The pipeline further integrates with Vision-Language Models for automated scenario mining. Experiments show that SIMSplat more than doubles baseline grounding accuracy, achieves the highest task completion rate, and produces the lowest failure rates across diverse driving scenarios.

cs.RO

An $L^\infty$ Rashevskii-Chow Theorem

Consider a finite family $\{f_1,\dots,f_\nu\}$ of $C^\infty$ vector fields on a $n$-dimensional ($n\in\mathbb{N}$), smooth manifold $\mathcal{M}$. The celebrated Rashevskii-Chow theorem states that, provided the vector fields $\{f_1,\dots,f_\nu\}$, together with their iterated Lie brackets, span the whole tangent space at some $x_*\in\mathcal{M}$, then any $x$ in a neighborhood of $x_*$ can be connected to $x_*$ by means of a finite concatenation of integral curves of $\{\pm f_1,\dots,\pm f_\nu\}$. This result finds applications in a number of areas, e.g., in control theory, in Sub-Riemannian geometry, and the theory of degenerate elliptic and parabolic partial differential equations, to mention a few. Here we extend this basic result to families of vector fields, which are considerably less regular, in particular, by allowing iterated Lie brackets to be just bounded measurable. This is technically made possible by the utilization of set-valued Lie brackets, which have already proven to be useful in extending commutativity type results, Frobenius' theorem, and also higher-order necessary conditions for optimal control problems, to the setting of non-smooth vector fields.

math.DS

The Telephone Game: Evaluating Semantic Drift in Unified Models

Unified models (UMs) combine visual understanding (I2T) and generation (T2I) in a single framework. We focus on T2I and I2T, where cross-consistency---what a model understands, it should be able to generate---is a promise of unification and a necessity when composing both capabilities. Yet, existing benchmarks evaluate them in isolation: FID/GenEval for T2I; MME/MMBench for I2T. We show this gap is consequential: models scoring competitively on these benchmarks can fail severely when understanding and generation are composed, losing entities, attributes, spatial relations, and counts, resulting in semantic drift. To quantify drift, we introduce the Semantic Drift Protocol (SDP), inspired by the Telephone Game: starting from a caption or image, we alternate I2T and T2I over multiple generations and measure semantic preservation. We propose Mean Cumulative Drift (MCD), an embedding-based measure of content retention across three representation spaces, and Multi-Generation GenEval (MGG), extending GenEval's object-level compliance scoring across generations. To stress-test models beyond COCO-style data, we create a benchmark of 400 image-text pairs sampled from NoCaps and DOCCI, emphasizing novel objects and fine-grained descriptions. Applying SDP to seven models reveals that drift varies dramatically and is not predicted by single-pass scores: BAGEL retains high semantic fidelity over multiple generations, while VILA-U and Janus variants collapse within five generations, despite comparable isolated metrics. We identify six recurring failure modes and find degradation is typically catastrophic rather than gradual: once a critical error occurs, subsequent generations compound it. SDP exposes failure modes that single-pass benchmarks miss, enabling a more faithful assessment of unified model reliability. Code and benchmark: https://github.com/mollahsabbir/telephone-game-semantic-drift

cs.CV

On Learning Closed-Loop Probabilistic Multi-Agent Simulator

The rapid iteration of autonomous vehicle (AV) deployments leads to increasing needs for building realistic and scalable multi-agent traffic simulators for efficient evaluation. Recent advances in this area focus on closed-loop simulators that enable generating diverse and interactive scenarios. This paper introduces Neural Interactive Agents (NIVA), a probabilistic framework for multi-agent simulation driven by a hierarchical Bayesian model that enables closed-loop, observation-conditioned simulation through autoregressive sampling from a latent, finite mixture of Gaussian distributions. We demonstrate how NIVA unifies preexisting sequence-to-sequence trajectory prediction models and emerging closed-loop simulation models trained on Next-token Prediction (NTP) from a Bayesian inference perspective. Experiments on the Waymo Open Motion Dataset demonstrate that NIVA attains competitive performance compared to the existing method while providing embellishing control over intentions and driving styles.

cs.RO

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior

Understanding a driver's behavior and intentions is important for potential risk assessment and early accident prevention. Safety and driver assistance systems can be tailored to individual drivers' behavior, significantly enhancing their effectiveness. However, existing datasets are limited in describing and explaining general vehicle movements based on external visual evidence. This paper introduces a benchmark, PDB-Eval, for a detailed understanding of Personalized Driver Behavior, and aligning Large Multimodal Models (MLLMs) with driving comprehension and reasoning. Our benchmark consists of two main components, PDB-X and PDB-QA. PDB-X can evaluate MLLMs' understanding of temporal driving scenes. Our dataset is designed to find valid visual evidence from the external view to explain the driver's behavior from the internal view. To align MLLMs' reasoning abilities with driving tasks, we propose PDB-QA as a visual explanation question-answering task for MLLM instruction fine-tuning. As a generic learning task for generative models like MLLMs, PDB-QA can bridge the domain gap without harming MLLMs' generalizability. Our evaluation indicates that fine-tuning MLLMs on fine-grained descriptions and explanations can effectively bridge the gap between MLLMs and the driving domain, which improves zero-shot performance on question-answering tasks by up to 73.2%. We further evaluate the MLLMs fine-tuned on PDB-X in Brain4Cars' intention prediction and AIDE's recognition tasks. We observe up to 12.5% performance improvements on the turn intention prediction task in Brain4Cars, and consistent performance improvements up to 11.0% on all tasks in AIDE.

cs.CV

Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance

While autonomous driving technologies continue to advance, current Advanced Driver Assistance Systems (ADAS) remain limited in their ability to interpret scene context or engage with drivers through natural language. These systems typically rely on predefined logic and lack support for dialogue-based interaction, making them inflexible in dynamic environments or when adapting to driver intent. This paper presents Scene-Aware Conversational ADAS (SC-ADAS), a modular framework that integrates Generative AI components including large language models, vision-to-text interpretation, and structured function calling to enable real-time, interpretable, and adaptive driver assistance. SC-ADAS supports multi-turn dialogue grounded in visual and sensor context, allowing natural language recommendations and driver-confirmed ADAS control. Implemented in the CARLA simulator with cloud-based Generative AI, the system executes confirmed user intents as structured ADAS commands without requiring model fine-tuning. We evaluate SC-ADAS across scene-aware, conversational, and revisited multi-turn interactions, highlighting trade-offs such as increased latency from vision-based context retrieval and token growth from accumulated dialogue history. These results demonstrate the feasibility of combining conversational reasoning, scene perception, and modular ADAS control to support the next generation of intelligent driver assistance.

cs.RO

VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable within individual frames or short clips. To truly understand videos as humans do, models must go beyond what is directly shown, inferring hidden relationships and contextual cues that are only implied across frames. Current benchmarks fail to capture this essential aspect of video understanding. To address this gap, we introduce VRR-QA, a benchmark for Visual Relational Reasoning Beyond Explicit Cues. We curate our benchmark from creative and cinematic videos such as movies, that deliberately employ storytelling techniques which omit direct depictions of certain events or relations, requiring viewers to infer them. VRR-QA comprises 1K meticulously expert-annotated QA pairs drawn from 1K creative video clips covering 15 genres across 7 decades of content, from both live-action and animated titles. Our extensive evaluations on 14 leading VideoQA models reveals consistent and significant performance degradation, underscoring their reliance on surface-level visual cues and highlighting the difficulty of implicit reasoning. Even the best model substantially underperforms human baselines with only 64% accuracy. Performance variations across models further illustrate the complexity and diversity of the challenges presented by VRR-QA. By releasing both dataset and data collection framework, VRR-QA establishes a rigorous, diverse, and reproducible testbed for advancing VideoQA: https://swetha5.github.io/ImplicitQA/.

cs.CV