SearcharxivSearch

arXiv subjects

Shivam Singh

Publications and source records attributed to Shivam Singh.

At least 19 recordsLinked to original sources

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often perform poorly on languages less well-represented in their training set. While it has long been known that effective language adaptation can be achieved through simple fine-tuning on monolingual data, this strategy has only been applied to a small number of languages. We massively scale up this simple approach to 102 languages covered in the FLEURS dataset, while also implementing a more complex language adaptation strategy that integrates monolingual tokenizer replacement and data augmentation using text-only fine-tuning. BuzzASR models outperform Whisper-large-v3 on 77 out of 102 languages, reducing character error rates (CER) by a factor of over 2.8 on average. Our models achieve state-of-the-art CER among open-source systems on 27 of 102 languages on the combined FLEURS and Common Voice test set. Our tokenizer replacement strategy yields an average 3.3x improvement in compression rate (characters per token) over Whisper's multilingual BPE, with gains of up to 21.7x. We release all models, code, and detailed results: https://lemn-lab.github.io/buzz-asr

cs.CL

Pulse-Burst Excitation Reveals Time-Dose Reciprocity Breakdown in Mixed-Halide Perovskites

Time-dose reciprocity, commonly associated with the Bunsen-Roscoe law, states that the response of a photosensitive system depends only on the total exposure dose, regardless of how that energy is delivered over time. Light-sensitive processes in mixed-halide perovskites, such as photoinduced halide segregation, often exhibit threshold-like behavior that may violate this principle and enable material-state control by photon timing. We test this using pulse-burst excitation, which introduces an additional temporal control dimension beyond conventional parameters such as pulse fluence, repetition rate, and average power. By redistributing the same photon dose over microsecond-to-millisecond timescales, we create distinct nonequilibrium excitation conditions and show that mixed-halide perovskites can evolve into different metastable states, revealing a breakdown of time-dose reciprocity in the combined processes of halide segregation and remixing. This additional temporal degree of freedom not only enables control of the material state but also provides a new experimental framework for disentangling the competing processes underlying photoinduced halide redistribution. Our findings establish photon timing as a control parameter for perovskite photochemistry and open additional opportunities for optical memory and neuromorphic photonic applications.

cond-mat.mtrl-sci

Unraveling the Roles of Shallow, Deep and Auger Trapping in Charge Carrier Recombination in Triple-Cation Perovskites

Understanding charge-carrier recombination in metal halide perovskites is essential for accurately identifying the factors limiting solar cell efficiency, yet it remains challenging due to the interplay of multiple competing processes. Here, we combine time-resolved photoluminescence and excitation dependent photoluminescence quantum yield measurements over a wide range of fluences and repetition rates to investigate recombination dynamics in triple-cation perovskite thin films. By jointly analyzing these multidimensional datasets, we develop a unified model that quantitatively reproduces both photoluminescence decays and absolute quantum yields across all excitation conditions. Our results reveal the coexistence of deep and shallow traps, as well as a second-order nonradiative recombination pathway attributed to Auger-assisted trapping. Importantly, this mechanism dominates under one-sun illumination, making it a critical limiting factor for photovoltaic performance. These findings provide a comprehensive framework for understanding recombination in perovskites and highlight the importance of higher-order defect-mediated processes in determining their efficiency.

cond-mat.mtrl-sci

GK-Mapper: A Stability Framework for Gustafson-Kessel Fuzzy Mapper Graphs

Topological Data Analysis uses tools from algebraic topology to study the shape and structure of data. The Mapper algorithm provides a graph-based summary of high-dimensional datasets by combining a filter function, a cover of the filter range, and clustering on the corresponding pullback sets. Several variants of Mapper have been proposed, including Conventional Mapper, F-Mapper, and Shape Fuzzy C-Means Mapper. In this article, we introduce Gustafson-Kessel Fuzzy Mapper Graphs, a geometry-adaptive extension of Shape Fuzzy C-Means Mapper. The proposed method replaces spherical fuzzy covers with ellipsoidal covers induced by the Gustafson-Kessel fuzzy clustering framework, making it more suitable for high-dimensional datasets with anisotropic and non-spherical geometry. We develop a stability framework for the graphs produced by Gustafson-Kessel Mapper and Shape Fuzzy C-Means Mapper. We prove that the membership functions depend smoothly on the fuzzifier, establish a precise condition for the existence of edges, and show that the graph is locally stable under small perturbations of the fuzzifier. We further describe the critical-event structure of graph changes in terms of threshold crossings of the membership functions and show that the graph is constant between consecutive critical events. When the threshold-crossing set is finite, this yields an eventual freezing threshold. Finally, we empirically show that Gustafson-Kessel Mapper can produce more stable graphs than Shape Fuzzy C-Means Mapper on high-dimensional and geometrically complex datasets.

math.AT

A Three Axis Evaluation Framework for Mapper Algorithms

Mapper is a well-known tool in topological data analysis, which visualizes and summarizes high-dimensional data. However, its output is sensitive to choices of lens functions, cover parameters, and clustering strategies, making evaluation challenging. Most works that have attempted to evaluate the Mapper algorithm have done so visually. In this paper, we review a roadmap for assessing Mapper algorithms along three complementary axes: stability, cluster quality, and topological shape preservation. We analyze Mapper and its variants on synthetic datasets and the UCI Digits dataset. These modes include topological explosion at high resolutions. Our findings indicate that these axes of evaluation are often in tension and that no single Mapper variant performs optimally across all three. This review provides practical guidelines for choosing Mapper variants and identifies open challenges toward a principled Mapper analysis.

math.AT

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-and-think-T, a reasoning-centric training dataset that encourages models to pause, reason over visual evidence, and produce concise, actionable responses. The dataset promotes structured reasoning prior to answer generation, guiding models toward human-like, scene-grounded assistance. We fine-tune a compact 4B-parameter model and evaluate it on our pause-and-think-B benchmark targeting contextual understanding and goal planning tasks. The model achieves 58.0% accuracy at 59x fewer parameters than Qwen3-VL-235B (58.9%), matching GPT-5.2 on scene understanding and surpassing GPT-4o. Beyond our benchmark, it also shows strong out-of-distribution performance on EgoThink and TempCompass, with substantial gains in affordance, assistance, attribution recognition, situated reasoning, and temporal order, without benchmark-specific training. Our results indicate that targeted reasoning supervision enables compact models to deliver actionable, visually grounded guidance while generalizing beyond training data, without requiring large-scale model expansion.

cs.CV

Additive-Induced Stabilization of the Energetic Landscape of PM6:Y12 Organic Solar Cells

Solvent additive engineering is a common strategy in organic photovoltaic (OPV) fabrication to improve film morphology and enhance device performance by controlling phase-separation kinetics and crystallinity. However, its effect on photostability, particularly with respect to the evolution of the energetic landscape under operational stress, remains unclear. This study investigates the impact of the additive 1-chloronaphthalene (1-CN) on the evolution of the device's energetic landscape in PM6:Y12 bulk heterojunction organic solar cells upon photoaging. Ultraviolet photoemission spectroscopy combined with argon gas cluster ion beam depth profiling is employed to probe the depth-resolved evolution of donor (PM6) and acceptor (Y12) energy levels before and after photodegradation. Our findings show that in additive-free devices, photodegradation leads to a significant 200 meV downward shift in the PM6 highest occupied molecular orbital (HOMO) level, reducing the donor-acceptor HOMO offset and impairing the driving force for hole transfer. As a consequence, the device experiences substantial efficiency loss. On the other hand, the incorporation of 1-CN effectively stabilizes the PM6 HOMO level, preserving adequate driving force for efficient exciton dissociation. Advanced X-ray diffraction characterization reveals more pronounced nanostructural degradation in blends without 1-CN than those with 1-CN upon photoaging. Collectively, these findings identify PM6 as the primary degradation pathway in PM6:Y12 blends and demonstrate that 1-CN enhances device stability by stabilizing PM6 energetics and preserving the nanostructural integrity upon photoaging.

cond-mat.mtrl-sci

Anticipate, Adapt, Act: A Hybrid Framework for Task Planning

Anticipating and adapting to failures is a key capability robots need to collaborate effectively with humans in complex domains. This continues to be a challenge despite the impressive performance of state of the art AI planning systems and Large Language Models (LLMs) because of the uncertainty associated with the tasks and their outcomes. Toward addressing this challenge, we present a hybrid framework that integrates the generic prediction capabilities of an LLM with the probabilistic sequential decision-making capability of Relational Dynamic Influence Diagram Language. For any given task, the robot reasons about the task and the capabilities of the human attempting to complete it; predicts potential failures due to lack of ability (in the human) or lack of relevant domain objects; and executes actions to prevent such failures or recover from them. Experimental evaluation in the VirtualHome 3D simulation environment demonstrates substantial improvement in performance compared with state of the art baselines.

cs.RO

Josephson Oscillation and Nonlinear Self-Trapping in Quasi-one-dimensional Quantum Liquid

In this article, we study the two-mode method to analyze the Josephson oscillation for a trapped binary Bose-Einstein condensate while taking into account the beyond mean-field and three body interactions. For this purpose, we use the archetypal model of double well potential and study the Josephson oscillation and self-trapping phases in quasi-one dimension. Additionally, our analysis provides quantitative discussion on the effect of asymmetry and dimension. We further corroborate our findings with Bogoliubov quasi-particle method and notice regions of instabilities and roton like mode.

cond-mat.quant-gas

Chimera: Compositional Image Generation using Part-based Concepting

Personalized image generative models are highly proficient at synthesizing images from text or a single image, yet they lack explicit control for composing objects from specific parts of multiple source images without user specified masks or annotations. To address this, we introduce Chimera, a personalized image generation model that generates novel objects by combining specified parts from different source images according to textual instructions. To train our model, we first construct a dataset from a taxonomy built on 464 unique (part, subject) pairs, which we term semantic atoms. From this, we generate 37k prompts and synthesize the corresponding images with a high-fidelity text-to-image model. We train a custom diffusion prior model with part-conditional guidance, which steers the image-conditioning features to enforce both semantic identity and spatial layout. We also introduce an objective metric PartEval to assess the fidelity and compositional accuracy of generation pipelines. Human evaluations and our proposed metric show that Chimera outperforms other baselines by 14% in part alignment and compositional accuracy and 21% in visual quality.

cs.CV

Molecular Cross-linking of MXenes: Tunable Interfaces and Chemiresistive Sensing

MXenes, a family of 2D transition metal compounds, have emerged as promising materials due to their unique electronic properties and tunable surface chemistry. However, the translation of these nanoscale properties into macroscopic devices is constrained by suitable cross-linking strategies that enable both processability and controlled inter-flake charge transport. Herein, we demonstrate the tunability of interfaces and the inter-layer spacing between Ti$_3$C$_2$T$_x$ MXene flakes through molecular cross-linking with homologous diamines. Oleylamine was first used to stabilize MXenes in chloroform, followed by diamine-mediated cross-linking to tune precisely the interlayer spacing. Grazing incidence X-ray scattering (GIXRD/GIWAXS) confirmed the correlation between ligand chain length and inter-layer spacing, which was further supported by Density Functional Theory (DFT) calculations. Furthermore, we investigated the charge transport properties of thin films consisting of these diamine-crosslinked Ti$_3$C$_2$T$_x$ MXenes and observed a strong dependence of the conductivity on the interlayer spacing. The dominating charge transport mechanism is variable range hopping (VRH) in accordance with the structure analysis of the films. Finally, we probed chemiresistive vapor sensing in MXene composites, observing pronounced water sensitivity and selectivity, highlighting their potential for use in humidity sensors. Insights into molecular cross-linking and its impact on charge transport open avenues for next-generation MXene-based electronic devices.

cond-mat.mtrl-sci

RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions

Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility.

cs.CV

Causal-Copilot: An Autonomous Causal Analysis Agent

Causal analysis plays a foundational role in scientific discovery and reliable decision-making, yet it remains largely inaccessible to domain experts due to its conceptual and algorithmic complexity. This disconnect between causal methodology and practical usability presents a dual challenge: domain experts are unable to leverage recent advances in causal learning, while causal researchers lack broad, real-world deployment to test and refine their methods. To address this, we introduce Causal-Copilot, an autonomous agent that operationalizes expert-level causal analysis within a large language model framework. Causal-Copilot automates the full pipeline of causal analysis for both tabular and time-series data -- including causal discovery, causal inference, algorithm selection, hyperparameter optimization, result interpretation, and generation of actionable insights. It supports interactive refinement through natural language, lowering the barrier for non-specialists while preserving methodological rigor. By integrating over 20 state-of-the-art causal analysis techniques, our system fosters a virtuous cycle -- expanding access to advanced causal methods for domain experts while generating rich, real-world applications that inform and advance causal theory. Empirical evaluations demonstrate that Causal-Copilot achieves superior performance compared to existing baselines, offering a reliable, scalable, and extensible solution that bridges the gap between theoretical sophistication and real-world applicability in causal analysis. A live interactive demo of Causal-Copilot is available at https://causalcopilot.com/.

cs.AI

AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement

An embodied agent assisting humans is often asked to complete new tasks, and there may not be sufficient time or labeled examples to train the agent to perform these new tasks. Large Language Models (LLMs) trained on considerable knowledge across many domains can be used to predict a sequence of abstract actions for completing such tasks, although the agent may not be able to execute this sequence due to task-, agent-, or domain-specific constraints. Our framework addresses these challenges by leveraging the generic predictions provided by LLM and the prior domain knowledge encoded in a Knowledge Graph (KG), enabling an agent to quickly adapt to new tasks. The robot also solicits and uses human input as needed to refine its existing knowledge. Based on experimental evaluation in the context of cooking and cleaning tasks in simulation domains, we demonstrate that the interplay between LLM, KG, and human input leads to substantial performance gains compared with just using the LLM. Project website§: https://sssshivvvv.github.io/adaptbot/

cs.RO

Anticipate & Act : Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments

Assistive agents performing household tasks such as making the bed or cooking breakfast often compute and execute actions that accomplish one task at a time. However, efficiency can be improved by anticipating upcoming tasks and computing an action sequence that jointly achieves these tasks. State-of-the-art methods for task anticipation use data-driven deep networks and Large Language Models (LLMs), but they do so at the level of high-level tasks and/or require many training examples. Our framework leverages the generic knowledge of LLMs through a small number of prompts to perform high-level task anticipation, using the anticipated tasks as goals in a classical planning system to compute a sequence of finer-granularity actions that jointly achieve these goals. We ground and evaluate our framework's abilities in realistic scenarios in the VirtualHome environment and demonstrate a 31% reduction in execution time compared with a system that does not consider upcoming tasks.

cs.RO

The Solar Ultraviolet Imaging Telescope on board Aditya-L1

The Solar Ultraviolet Imaging Telescope (SUIT) is an instrument on the Aditya-L1 mission of the Indian Space Research Organization (ISRO) launched on September 02, 2023. SUIT continuously provides, near-simultaneous full-disk and region-of-interest images of the Sun, slicing through the photosphere and chromosphere and covering a field of view up to 1.5 solar radii. For this purpose, SUIT uses 11 filters tuned at different wavelengths in the 200{--}400~nm range, including the Mg~{\sc ii} h~and~k and Ca~{\sc ii}~H spectral lines. The observations made by SUIT help us understand the magnetic coupling of the lower and middle solar atmosphere. In addition, for the first time, it allows the measurements of spatially resolved solar broad-band radiation in the near and mid ultraviolet, which will help constrain the variability of the solar ultraviolet irradiance in a wavelength range that is central for the chemistry of the Earth's atmosphere. This paper discusses the details of the instrument and data products.

astro-ph.SR

Quantum Liquid in Lower Dimensions: From the perspective of Surface Tension

We analyze the surface tension in ultra-cold atomic gases in a quasi one-dimensional and one-dimensional geometry. In recent years, experimental observations have confirmed the ``clustering of atoms" to form droplets in ultra-cold atomic gases and the emergence of this new phase is attributed to the beyond mean-field interaction. However, two decades earlier, liquid formation was predicted due to the competition of two-body and three-body interactions. Here, we review both propositions and comment on the role of beyond mean-field and three-body interaction in liquid formation by calculating the surface tension.

cond-mat.quant-gas

Anticipate & Collab: Data-driven Task Anticipation and Knowledge-driven Planning for Human-robot Collaboration

An agent assisting humans in daily living activities can collaborate more effectively by anticipating upcoming tasks. Data-driven methods represent the state of the art in task anticipation, planning, and related problems, but these methods are resource-hungry and opaque. Our prior work introduced a proof of concept framework that used an LLM to anticipate 3 high-level tasks that served as goals for a classical planning system that computed a sequence of low-level actions for the agent to achieve these goals. This paper describes DaTAPlan, our framework that significantly extends our prior work toward human-robot collaboration. Specifically, DaTAPlan planner computes actions for an agent and a human to collaboratively and jointly achieve the tasks anticipated by the LLM, and the agent automatically adapts to unexpected changes in human action outcomes and preferences. We evaluate DaTAPlan capabilities in a realistic simulation environment, demonstrating accurate task anticipation, effective human-robot collaboration, and the ability to adapt to unexpected changes. Project website: https://dataplan-hrc.github.io

cs.RO