SearcharxivSearch

arXiv subjects

Siddharth Singh

Publications and source records attributed to Siddharth Singh.

At least 19 recordsLinked to original sources

Flip-chip integrated superconducting qubits using electroplated bump bonds

Flip-chip integration offers a promising route toward scalable superconducting quantum processors and hybrid semiconductor-superconductor quantum devices. We develop a three-dimensional transmon architecture using electroplated indium in which the qubit electric field is shared nearly equally between two bump-bonded substrates while maintaining low participation at the indium-bump interface. The resulting geometry is well suited for future hybrid qubits, enabling the integration of distinct material platforms while minimizing sensitivity to bump-interface loss. Using this platform, we evaluate electroplated indium interconnects for superconducting quantum circuits. Flip-chip transmons incorporating electroplated indium bumps exhibit qubit quality factors around $10^6$. In addition, a systematic study of coplanar-waveguide resonators is used to identify losses associated with the electroplating process. In particular, we find that surface losses associated with the gold-layer, used to enable good electric contact with the indium, is likely the primary contributor to the qubit decay rate. These results demonstrate the compatibility of electroplated indium technology with high-coherence superconducting circuits and establish a promising platform for three-dimensional hybrid quantum integration.

quant-ph

Experimental Characterization and Modeling of Measurement-Induced State-Transitions in a Fluxonium Superconducting Qubit

Superconducting qubits are most often measured using dispersive readout, which, ideally, implements a projective quantum non-demolition (QND) measurement. While a larger readout drive can increase the signal and, thus, reduce discrimination errors in the readout, strong microwave drives may also cause non-QND errors by driving the qubit to a state outside the computational subspace. In this work, we experimentally characterize measurement-induced state transitions (MIST) in a fluxonium qubit over its full external flux range. We further numerically calculate the MIST errors, and find that the theory accurately predicts eleven experimentally identified regions with increased MIST. In addition to transitions to higher fluxonium levels, we also find that, at certain flux points, MIST errors are dominated by transitions that include the transmission-line-like array modes of the fluxonium's superinductor. The excellent match between theory and experiment validates that the models accurately predict the occurrence of MIST in these systems, and further highlights the influence of array modes in fluxonium readout.

quant-ph

Understanding and Improving Communication Performance in Multi-node LLM Inference

As large language models (LLMs) continue to grow in size, distributed inference has become increasingly important. Model-parallel strategies must now efficiently scale not only across multiple GPUs but also across multiple nodes. In this work, we present a detailed performance study of multi-node distributed inference using LLMs on GPU-based supercomputers. We conduct experiments with several state-of-the-art inference engines alongside YALIS, a research-oriented prototype engine designed for controlled experimentation. We analyze the strong-scaling behavior of different model-parallel schemes and identify key bottlenecks. Because all-reduce operations are a common performance bottleneck, we develop NVRAR, a hierarchical all-reduce algorithm based on recursive doubling with NVSHMEM. NVRAR achieves up to 1.9$\times$-3.6$\times$ lower latency than NCCL for message sizes between 128 KB and 2 MB on HPE Slingshot and InfiniBand interconnects. Integrated into YALIS, NVRAR achieves up to a 1.72$\times$ reduction in end-to-end batch latency for the Llama 3.1 405B model in multi-node decode-heavy workloads using tensor parallelism.

cs.DC

A Systematic Failure Analysis of Vision Foundation Models for Open Set Iris Presentation Attack Detection

Vision foundation models have demonstrated strong transferability across diverse visual recognition tasks and are increasingly considered for biometric applications. Their suitability for iris Presentation Attack Detection (PAD), particularly under realistic open-set operating conditions, remains insufficiently examined. This work presents a systematic failure analysis of general-purpose vision foundation models for open-set iris PAD using periocular imagery. Five representative foundation models are evaluated under three open-set protocols that explicitly separate different sources of distribution shift: unseen Presentation Attack Instruments (PAIs), unseen datasets captured with different sensors and cross-spectral transfer from near-infrared (NIR) to visible spectrum (VIS) imagery. Both frozen feature representations and parameter-efficient task adaptation using Low-Rank Adaptation (LoRA) are assessed within a unified experimental framework. The results indicate that foundation models can transfer across datasets with similar sensing characteristics, but fail to generalise reliably to unseen attack instruments and degrade sharply under cross-spectral evaluation. While LoRA improves performance in certain cross-dataset settings, it frequently amplifies failure under attack-level and spectral shifts. Additional validation experiments using segmented iris inputs, full backbone fine-tuning, joint cross-dataset and cross-PAI shifts, and reverse VIS to NIR transfer further confirm that these failures are not simply artefacts of periocular input, weak adaptation, or one-directional spectral evaluation. These findings show that strong closed-set or cross-dataset performance should not be treated as evidence of robust open-set security, and highlight the need for PAD representations that maintain sensitivity to presentation artefacts while remaining stable under realistic deployment variation.

cs.CV

Finite-width adiabatic shear banding and dislocation patterning in mesoscale polycrystalline aggregates

Dynamic shear banding under adiabatic conditions in a mesoscale polycrystalline aggregate is studied using a model of mesoscale dislocation mechanics and experiments. The model involves a length scale related to hardening induced by excess/polar/geometrically necessary dislocation (GND) density, and utilizes a simple classical crystal plasticity model with isotropic Voce law hardening. Simulations of statistically representative volume elements of a polycrystal determined from experimental samples are conducted. Studies in 2-d (section) and 3-d capture the experimentally observed finite-width shear bands and the formation of low-angle subgrain boundaries even in the absence of heat conduction in the model, as well as size-dependent strengthening for grain sizes from 1 to 20 $μ$m. The 2-d and large-scale 3-d simulations, the latter involving 1 million finite elements, provide access to the progressive evolution of material strength, stress state, and temperature in the course of large deformations. GND distributions accumulate at grain boundaries and form patterned structures within grain interiors, offering insight into the microstructural changes that precede failure in adiabatic shear bands. Mesh-converged, delocalized and localized plastic flow to very large deformations without softening is observed for a significant range of parameters, reflecting a competition between GND hardening and thermal softening in setting the non-softening steady state in the absence of other ductile damage mechanisms in the model.

cond-mat.mtrl-sci

Coherence limitations of a Fourier-engineered $\cos(2φ)$ transmon qubit

Intrinsically protected superconducting qubits are a promising route toward enhancing coherence times and advancing hardware towards applications in quantum computing. The $\cos(2φ)$ qubit achieves protection against qubit relaxation by allowing only the coherent tunneling of pairs of Cooper pairs, resulting in Cooper-pair parity symmetry and thereby suppressing charge-induced errors. In this work, we experimentally realize a $\cos(2φ)$ qubit by Fourier engineering the energy-phase relation in a multi-junction superconducting circuit. Using an interference-based architecture, we are able to suppress the odd harmonics of an effective qubit potential and we observe good agreement between the measured transition spectrum and the effective theoretical model. We further investigate the energy relaxation time as a function of external flux and find that the qubit lifetime at the flux symmetry point is limited by $1/f$ flux noise. This strong sensitivity arises from residual fluctuations in the first harmonic, which possesses a large prefactor despite being nominally canceled. In contrast, a fluxonium qubit with a similar energy spectrum and noise amplitude is less affected by flux noise, highlighting a key challenge for interference-based protection schemes.

quant-ph

Train-Small Deploy-Large: Leveraging Diffusion-Based Multi-Robot Planning

Learning based multi-robot path planning methods struggle to scale or generalize to changes, particularly variations in the number of robots during deployment. Most existing methods are trained on a fixed number of robots and may tolerate a reduced number during testing, but typically fail when the number increases. Additionally, training such methods for a larger number of agents can be both time consuming and computationally expensive. However, analytical methods can struggle to scale computationally or handle dynamic changes in the environment. In this work, we propose to leverage a diffusion model based planner capable of handling dynamically varying number of agents. Our approach is trained on a limited number of agents and generalizes effectively to larger numbers of agents during deployment. Results show that integrating a single shared diffusion model based planner with dedicated inter-agent attention computation and temporal convolution enables a train small deploy-large paradigm with good accuracy. We validate our method across multiple scenarios and compare the performance with existing multi-agent reinforcement learning techniques and heuristic control based methods.

cs.RO

Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training

Graph neural networks (GNNs) are widely used for learning on graph datasets derived from various real-world scenarios. Learning from extremely large graphs requires distributed training, and mini-batching with sampling is a popular approach for parallelizing GNN training. Existing distributed mini-batch approaches have significant performance bottlenecks due to expensive sampling methods and limited scaling when using data parallelism. In this work, we present ScaleGNN, a 4D parallel framework for scalable mini-batch GNN training that combines communication-free distributed sampling, 3D parallel matrix multiplication (PMM), and data parallelism. ScaleGNN introduces a uniform vertex sampling algorithm, enabling each process (GPU device) to construct its local mini-batch, i.e., subgraph partitions without any inter-process communication. 3D PMM enables scaling mini-batch training to much larger GPU counts than vanilla data parallelism with significantly lower communication overheads. We also present additional optimizations to overlap sampling with training, reduce communication overhead by sending data in lower precision, kernel fusion, and communication-computation overlap. We evaluate ScaleGNN on five graph datasets and demonstrate strong scaling up to 2048 GPUs on Perlmutter, 2048 GCDs on Frontier, and 1024 GPUs on Tuolumne. On Perlmutter, ScaleGNN achieves 3.5x end-to-end training speedup over the SOTA baseline on ogbn-products.

cs.LG

The Big Send-off: Scalable and Performant Collectives for Deep Learning

Collective communication is becoming increasingly important in data center and supercomputer workloads with an increase in distributed AI related jobs. However, existing libraries that provide collective support such as NCCL, RCCL, and Cray-MPICH exhibit several performance and scalability limitations on modern GPU supercomputers. To address these challenges, we introduce the Performant Collective Communication Library (PCCL), specifically targeted for distributed deep learning (DL) workloads. PCCL provides highly optimized implementations of key collectives used in distributed DL: all-gather, reduce-scatter, and all-reduce. PCCL uses a hierarchical design with learning-based adaptive selection of the best performing algorithms to scale efficiently to thousands of GPUs. It achieves substantial performance speedups over RCCL on 2048 GCDs of Frontier -- up to 168x for reduce-scatter, 33x for all-gather and 10x for all-reduce. More modest but still significant gains up to 5.7x over NCCL are observed on Perlmutter. These gains translate directly to performance improvement of production DL workloads: up to 4.9x speedup over RCCL in DeepSpeed ZeRO-3 training, and up to 2.4x speedup in DDP training.

cs.DC

Realistic Synthetic Household Data Generation at Scale

Advancements in foundation models have catalyzed research in Embodied AI to develop interactive agents capable of environmental reasoning and interaction. Developing such agents requires diverse, large-scale datasets. Prior frameworks generate synthetic data for long-term human-robot interactions but fail to model the bidirectional influence between human behavior and household environments. Our proposed generative framework creates household datasets at scale through loosely coupled generation of long-term human-robot interactions and environments. Human personas influence environment generation, while environment schematics and semantics shape human-robot interactions. The generated 3D data includes rich static context such as object and environment semantics, and temporal context capturing human and agent behaviors over extended periods. Our flexible tool allows users to define dataset characteristics via natural language prompts, enabling configuration of environment and human activity data through natural language specifications. The tool creates variations of user-defined configurations, enabling scalable data generation. We validate our framework through statistical evaluation using multi-modal embeddings and key metrics: cosine similarity, mutual information gain, intervention analysis, and iterative improvement validation. Statistical comparisons show good alignment with real-world datasets (HOMER) with cosine similarity (0.60), while synthetic datasets (Wang et al.) show moderate alignment (0.27). Intervention analysis across age, organization, and sleep pattern changes shows statistically significant effects (p < 0.001) with large effect sizes (Cohen's d = 0.51-1.12), confirming bidirectional coupling translates persona traits into measurable environmental and behavioral differences. These contributions enable development and testing of household smart devices at scale.

cs.RO

Physically Informed Bayesian Retrieval of SWE and Snow Depth in Forested Areas from Airborne X And Ku-Band SAR Measurements

This study presents a coupled physical statistical framework for retrieving snow water equivalent (SWE) in forested areas using dual frequency X and Ku band SAR observations. The method combines a multilayer snow hydrology model (MSHM) with microwave propagation and backscatter models, and includes a canopy parameterization based on a modified Water Cloud Model that accounts for canopy closure. The framework is applied to airborne SnowSAR measurements over Grand Mesa, Colorado, and evaluated against snow pit SWE and LiDAR snow depth from the SnowEx'17 campaign. Prior distributions of snowpack properties are generated with MSHM forced by numerical weather prediction, and vegetation and soil parameters are initialized from Ku HH observations under frozen conditions and interpolated from open to nearby forested areas using kriging. Successful SWE and snow depth retrievals in forested pixels are obtained where relative backscatter residuals are below 30% for incidence angles between 30 and 50 degrees, capturing both the mean and variance of snowpack distributions. For 90 m forested pixels, the snow depth RMSE is 0.033 m (less than 8% of maximum pit SWE), with improved spatial patterns relative to hydrology only simulations. Performance degrades in highly heterogeneous land cover such as mixed forest and wetlands and along canopy and water boundaries due to uncertainty in canopy closure, although absolute snow depth differences remain below 10% and 20% for about 62% and 82% of pixels, respectively. Retrievals at 30 m resolution for one flight further reduce spatial errors and increase the fraction of low error pixels by about 78% at a 10% absolute error threshold, demonstrating the feasibility of dual frequency Bayesian SWE retrievals in forested landscapes by combining physical modeling with SAR observations.

physics.geo-ph

Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training

Graph neural networks (GNNs) leverage the connectivity and structure of real-world graphs to learn intricate properties and relationships between nodes. Many real-world graphs exceed the memory capacity of a GPU due to their sheer size, and training GNNs on such graphs requires techniques such as mini-batch sampling to scale. The alternative approach of distributed full-graph training suffers from high communication overheads and load imbalance due to the irregular structure of graphs. We propose a three-dimensional (3D) parallel approach for full-graph training that tackles these issues and scales to billion-edge graphs. In addition, we introduce optimizations such as a double permutation scheme for load balancing, and a performance model to predict the optimal 3D configuration of our parallel implementation -- Plexus. We evaluate Plexus on six different graph datasets and show scaling results on up to 2048 GPUs of Perlmutter, and 1024 GPUs of Frontier. Plexus achieves unprecedented speedups of 2.3-12.5x over prior state of the art, and a reduction in time-to-solution by 5.2-8.7x on Perlmutter and 7.0-54.2x on Frontier.

cs.LG

Gemstones: A Model Suite for Multi-Faceted Scaling Laws

Scaling laws are typically fit using a family of models with a narrow range of frozen hyperparameter choices. In this work we study scaling laws using multiple architectural shapes and hyperparameter choices, highlighting their impact on resulting prescriptions. As a primary artifact of our research, we release the Gemstones: an open-source scaling law dataset, consisting of over 4000 checkpoints from transformers with up to 2 billion parameters and diverse architectural shapes; including ablations over learning rate and cooldown. Our checkpoints enable more complex studies of scaling, such as analyzing the relationship between width and depth. By examining our model suite, we find that the prescriptions of scaling laws can be highly sensitive to the experimental design process and the specific model checkpoints used during fitting.

cs.LG

Fast microwave-driven two-qubit gates between fluxonium qubits with a transmon coupler

Two qubit gates constitute fundamental building blocks in the realization of large-scale quantum devices. Using superconducting circuits, two-qubit gates have previously been implemented in different ways with each method aiming to maximize gate fidelity. Another important goal of a new gate scheme is to minimize the complexity of gate calibration. In this work, we demonstrate a high-fidelity two-qubit gate between two fluxonium qubits enabled by an intermediate capacitively coupled transmon. The coupling strengths between the qubits and the coupler are designed to minimize residual crosstalk while still allowing for fast gate operations. The gate is based on frequency selectively exciting the coupler using a microwave drive to complete a 2$π$ rotation, conditional on the state of the fluxonium qubits. When successful, this drive scheme implements a conditional phase gate. Using analytically derived pulse shapes, we minimize unwanted excitations of the coupler and obtain gate errors of $10^{-2}$ for gate times below 60~ns. At longer durations, our gate is limited by relaxation of the coupler. Our results show how carefully designed control pulses can speed up frequency selective entangling gates.

quant-ph

Single-Qubit Gates Beyond the Rotating-Wave Approximation for Strongly Anharmonic Low-Frequency Qubits

Single-qubit gates are in many quantum platforms applied using a linear drive resonant with the qubit transition frequency which is often theoretically described within the rotating-wave approximation (RWA). However, for fast gates on low-frequency qubits, the RWA may not hold and we need to consider the contribution from counter-rotating terms to the qubit dynamics. The inclusion of counter-rotating terms into the theoretical description gives rise to two challenges. Firstly, it becomes challenging to analytically calculate the time evolution as the Hamiltonian is no longer self-commuting. Moreover, the time evolution now depends on the carrier phase such that, in general, every operation in a sequence of gates is different. In this work, we derive and verify a correction to the drive pulses that minimizes the effect of these counter-rotating terms in a two-level system. We then derive a second correction term that arises from non-computational levels for a strongly anharmonic system. We experimentally implement these correction terms on a fluxonium superconducting qubit, which is an example of a strongly anharmonic, low-frequency qubit for which the RWA may not hold, and demonstrate how fast, high-fidelity single-qubit gates can be achieved without the need for additional hardware complexities.

quant-ph

Collaborative motion planning for multi-manipulator systems through Reinforcement Learning and Dynamic Movement Primitives

Robotic tasks often require multiple manipulators to enhance task efficiency and speed, but this increases complexity in terms of collaboration, collision avoidance, and the expanded state-action space. To address these challenges, we propose a multi-level approach combining Reinforcement Learning (RL) and Dynamic Movement Primitives (DMP) to generate adaptive, real-time trajectories for new tasks in dynamic environments using a demonstration library. This method ensures collision-free trajectory generation and efficient collaborative motion planning. We validate the approach through experiments in the PyBullet simulation environment with UR5e robotic manipulators.

cs.RO

Hybrid Robot Learning for Automatic Robot Motion Planning in Manufacturing

Industrial robots are widely used in diverse manufacturing environments. Nonetheless, how to enable robots to automatically plan trajectories for changing tasks presents a considerable challenge. Further complexities arise when robots operate within work cells alongside machines, humans, or other robots. This paper introduces a multi-level hybrid robot motion planning method combining a task space Reinforcement Learning-based Learning from Demonstration (RL-LfD) agent and a joint-space based Deep Reinforcement Learning (DRL) based agent. A higher level agent learns to switch between the two agents to enable feasible and smooth motion. The feasibility is computed by incorporating reachability, joint limits, manipulability, and collision risks of the robot in the given environment. Therefore, the derived hybrid motion planning policy generates a feasible trajectory that adheres to task constraints. The effectiveness of the method is validated through sim ulated robotic scenarios and in a real-world setup.

cs.RO

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

We study a novel language model architecture that is capable of scaling test-time computation by implicitly reasoning in latent space. Our model works by iterating a recurrent block, thereby unrolling to arbitrary depth at test-time. This stands in contrast to mainstream reasoning models that scale up compute by producing more tokens. Unlike approaches based on chain-of-thought, our approach does not require any specialized training data, can work with small context windows, and can capture types of reasoning that are not easily represented in words. We scale a proof-of-concept model to 3.5 billion parameters and 800 billion tokens. We show that the resulting model can improve its performance on reasoning benchmarks, sometimes dramatically, up to a computation load equivalent to 50 billion parameters.

cs.LG