SearcharxivSearch

arXiv subjects

Emanuele Palumbo

Publications and source records attributed to Emanuele Palumbo.

12 recordsLinked to original sources

Device-Agnostic Microwave Noise Metrology for Nonlinear Cryogenic Quantum Devices

Microwave devices capable of near-quantum-limited signal processing are essential components in the toolbox of solid-state quantum technologies. The manipulation and readout of single-photon microwave signals through amplifiers, mixers, isolators, etc. must fulfill strict requirements in terms of signal integrity to ensure reliable operation. These active microwave quantum devices operate in complex cryo-electronic setups. This poses challenges to their characterization, since all relevant figures of merit must be expressed at the reference planes of their ports. Even though cryogenic S-parameter calibration is non-trivial, metrological approaches are converging toward rigorous methods. Furthermore, preserving signal integrity must be quantified via absolute noise levels at the ports of the Device Under Test (DUT), requiring an absolute power reference. In this work, we present an in situ noise metrology protocol based on substituting a controllable noise source for the DUT. We motivate this choice by showing that placing the noise source at the DUT input impacts the separability of the calibration from the DUT characteristics. Our proposed architecture combines Planck spectroscopy using a Variable Temperature Stage with Short-Open-Load-Reciprocal scattering-parameter calibration, so that noise and scattering quantities are referred to the same cryogenic reference planes. In this configuration, the readout-chain calibration is separated from the internal dynamics of the DUT. As a demanding use case, we apply the protocol to a Josephson Traveling Wave Parametric Amplifier and extract its gain and input-referred added noise under pump conditions activating multimode nonlinear behavior. This illustrates how our device-agnostic protocol supports portable noise characterization of nonlinear cryogenic microwave devices.

quant-ph

Traceable In Situ Microwave Power Measurement at the Cryogenic Device Plane in a Dilution Refrigerator

Accurate knowledge of the microwave power delivered to a cryogenic device under test (DUT) is essential for the characterization and operation of superconducting quantum circuits. However, this information is difficult to obtain inside dilution refrigerators because of distributed attenuation, impedance mismatch, switch-path repeatability, and temperature-dependent microwave components. This paper presents an in situ measurement method for RF power at the cryogenic device plane. The method uses a custom variable temperature stage (VTS) as a cryogenic thermal-transfer element. The TVS is alternately heated by a four-wire DC heater and by microwave power dissipated in a 20 dB pass-through attenuator. By fitting the thermal transients and comparing the corresponding steady-state temperatures, the absorbed microwave power is inferred from a directly measured DC electrical power through an AC/DC substitution procedure. The finite reflection and transmission of the attenuator are then accounted for by cryogenic two-port scattering-parameter measurements based on a switch-assisted Short--Open--Load--Reciprocal calibration, so that the result is referred to the DUT reference plane. The system is demonstrated in a dilution refrigerator with powers between -43 and -58 dBm at the DUT input plane. The demonstrated relative standard uncertainty ranges from about 2% at -43.9 dBm to about 40% at -57.6 dBm. The proposed approach combines thermal RF power transfer, cryogenic S-parameter correction, and uncertainty evaluation in a measurement architecture compatible with quantum-device experiments, providing a practical route toward traceable microwave-power calibration at millikelvin stages.

physics.ins-det

TOAST: Transformer Optimization using Adaptive and Simple Transformations

Foundation models achieve state-of-the-art performance across different tasks, but their size and computational demands raise concerns about accessibility and sustainability. Existing efficiency methods often require additional retraining or finetuning, limiting their practicality. Recent findings suggest that deep neural networks exhibit internal representation similarities. While such similarities across different models have been exploited for enabling techniques such as model stitching and merging, intra-network redundancy remains underexplored as a source for efficiency gains. In this paper, we introduce Transformer Optimization using Adaptive and Simple Transformations (TOAST), a framework that exploits these redundancies to approximate entire transformer blocks with lightweight closed-form mappings, such as linear transformations or even the identity function, without any additional training. Across state-of-the-art pretrained vision models (e.g., ViT, DINOv2, DeiT) and datasets ranging from MNIST to ImageNet-1k, TOAST reduces parameters and computation while preserving, and in some cases improving, downstream performance. These results show that large portions of transformer depth can be replaced by trivial functions, opening a new perspective on efficient foundation models.

cs.LG

JCO: Optimization Framework for Nonlinear Superconducting Circuits Using a Lumped-Element Approach and Harmonic Balance

In this contribution we present JosephsonCircuitsOptimizer.jl (JCO), a simulation and optimization framework based on the JosephsonCircuits.jl library for Julia. It models superconducting circuits that include Josephson junctions (JJs) and other nonlinear elements within a lumped-element approach, leveraging harmonic balance, a frequency-domain technique that provides a computationally efficient alternative to traditional time-domain simulations. JCO automates the evaluation of optimal circuit parameters by implementing Bayesian optimization with Gaussian processes through a device-specific metric and identifying the optimal working point to achieve a defined performance function. This makes it well suited for circuits with strong nonlinearity and a high-dimensional set of coupled design parameters. To demonstrate its capabilities, we focus on optimizing a Josephson Traveling-Wave Parametric Amplifier (JTWPA) based on Superconducting Nonlinear Asymmetric Inductive eLements (SNAILs), operating in the three-wave mixing regime. The device consists of an array of unit cells, each containing a loop with multiple JJs, that amplifies weak quantum signals near the quantum noise limit. By integrating efficient simulation and optimization strategies, the framework supports the systematic development of superconducting circuits for a broad range of applications.

quant-ph

Post-hoc Stochastic Concept Bottleneck Models

Concept Bottleneck Models (CBMs) are interpretable models that predict the target variable through high-level human-understandable concepts, allowing users to intervene on mispredicted concepts to adjust the final output. While recent work has shown that modeling dependencies between concepts can improve CBM performance, especially under interventions, such approaches typically require retraining the entire model, which may be infeasible when access to the original data or compute is limited. In this paper, we introduce Post-hoc Stochastic Concept Bottleneck Models (PSCBMs), a lightweight method that augments any pre-trained CBM with a multivariate normal distribution over concepts by adding only a small covariance-prediction module, without retraining the backbone model. We propose two training strategies and show on real-world data that PSCBMs consistently match or improve both concept and target accuracy over standard CBMs at test time. Furthermore, we show that due to the modeling of concept dependencies, PSCBMs perform much better than CBMs under interventions, while remaining far more efficient than retraining a similar stochastic model from scratch.

cs.LG

Steering Generative Models for Accessibility: EasyRead Image Generation

EasyRead pictograms are simple, visually clear images that represent specific concepts and support comprehension for people with intellectual disabilities, low literacy, or language barriers. The large-scale production of EasyRead content has traditionally been constrained by the cost and expertise required to manually design pictograms. In contrast, automatic generation of such images could significantly reduce production time and cost, enabling broader accessibility across digital and printed materials. However, modern diffusion-based image generation models tend to produce outputs that exhibit excessive visual detail and lack stylistic stability across random seeds, limiting their suitability for clear and consistent pictogram generation. This challenge highlights the need for methods specifically tailored to accessibility-oriented visual content. In this work, we present a unified pipeline for generating EasyRead pictograms by fine-tuning a Stable Diffusion model using LoRA adapters on a curated corpus that combines augmented samples from multiple pictogram datasets. Since EasyRead pictograms lack a unified formal definition, we introduce an EasyRead score to benchmark pictogram quality and consistency. Our results demonstrate that diffusion models can be effectively steered toward producing coherent EasyRead-style images, indicating that generative models can serve as practical tools for scalable and accessible pictogram production.

cs.HC

From Logits to Hierarchies: Hierarchical Clustering made Simple

The hierarchical structure inherent in many real-world datasets makes the modeling of such hierarchies a crucial objective in both unsupervised and supervised machine learning. While recent advancements have introduced deep architectures specifically designed for hierarchical clustering, we adopt a critical perspective on this line of research. Our findings reveal that these methods face significant limitations in scalability and performance when applied to realistic datasets. Given these findings, we present an alternative approach and introduce a lightweight method that builds on pre-trained non-hierarchical clustering models. Remarkably, our approach outperforms specialized deep models for hierarchical clustering, and it is broadly applicable to any pre-trained clustering model that outputs logits, without requiring any fine-tuning. To highlight the generality of our approach, we extend its application to a supervised setting, demonstrating its ability to recover meaningful hierarchies from a pre-trained ImageNet classifier. Our results establish a practical and effective alternative to existing deep hierarchical clustering methods, with significant advantages in efficiency, scalability and performance.

cs.LG

Hybrid Modeling of Photoplethysmography for Non-invasive Monitoring of Cardiovascular Parameters

Continuous cardiovascular monitoring can play a key role in precision health. However, some fundamental cardiac biomarkers of interest, including stroke volume and cardiac output, require invasive measurements, e.g., arterial pressure waveforms (APW). As a non-invasive alternative, photoplethysmography (PPG) measurements are routinely collected in hospital settings. Unfortunately, the prediction of key cardiac biomarkers from PPG instead of APW remains an open challenge, further complicated by the scarcity of annotated PPG measurements. As a solution, we propose a hybrid approach that uses hemodynamic simulations and unlabeled clinical data to estimate cardiovascular biomarkers directly from PPG signals. Our hybrid model combines a conditional variational autoencoder trained on paired PPG-APW data with a conditional density estimator of cardiac biomarkers trained on labeled simulated APW segments. As a key result, our experiments demonstrate that the proposed approach can detect fluctuations of cardiac output and stroke volume and outperform a supervised baseline in monitoring temporal changes in these biomarkers.

cs.LG

Towards Quantifying Two-Mode Correlation Linewidths in Quantum Circuits

This paper aims to quantify the linewidth of two-mode correlations in Traveling Wave Parametric Amplifiers (TWPAs). Artifacts induced by data acquisition and processing, such as windowing effects and acquisition time, are examined to understand their influence on the linewidth estimation of these correlations. The findings underscore the significance of acquisition parameters in optimizing two-mode correlation measurements, enhancing device characterization for quantum applications.

quant-ph

Characterization of a Transmon Qubit in a 3D Cavity for Quantum Machine Learning and Photon Counting

In this paper we report the use of superconducting transmon qubit in a 3D cavity for quantum machine learning and photon counting applications. We first describe the realization and characterization of a transmon qubit coupled to a 3D resonator, providing a detailed description of the simulation framework and of the experimental measurement of important parameters, like the dispersive shift and the qubit anharmonicity. We then report on a Quantum Machine Learning application implemented on the single-qubit device to fit the u-quark parton distribution function of the proton. In the final section of the manuscript we present a new microwave photon detection scheme based on two qubits coupled to the same 3D resonator. This could in principle decrease the dark count rate, favouring applications like axion dark matter searches.

quant-ph

Identifiability Results for Multimodal Contrastive Learning

Contrastive learning is a cornerstone underlying recent progress in multi-view and multimodal learning, e.g., in representation learning with image/caption pairs. While its effectiveness is not yet fully understood, a line of recent work reveals that contrastive learning can invert the data generating process and recover ground truth latent factors shared between views. In this work, we present new identifiability results for multimodal contrastive learning, showing that it is possible to recover shared factors in a more general setup than the multi-view setting studied previously. Specifically, we distinguish between the multi-view setting with one generative mechanism (e.g., multiple cameras of the same type) and the multimodal setting that is characterized by distinct mechanisms (e.g., cameras and microphones). Our work generalizes previous identifiability results by redefining the generative process in terms of distinct mechanisms with modality-specific latent variables. We prove that contrastive learning can block-identify latent factors shared between modalities, even when there are nontrivial dependencies between factors. We empirically verify our identifiability results with numerical simulations and corroborate our findings on a complex multimodal dataset of image/text pairs. Zooming out, our work provides a theoretical basis for multimodal representation learning and explains in which settings multimodal contrastive learning can be effective in practice.

cs.LG

On the Limitations of Multimodal VAEs

Multimodal variational autoencoders (VAEs) have shown promise as efficient generative models for weakly-supervised data. Yet, despite their advantage of weak supervision, they exhibit a gap in generative quality compared to unimodal VAEs, which are completely unsupervised. In an attempt to explain this gap, we uncover a fundamental limitation that applies to a large family of mixture-based multimodal VAEs. We prove that the sub-sampling of modalities enforces an undesirable upper bound on the multimodal ELBO and thereby limits the generative quality of the respective models. Empirically, we showcase the generative quality gap on both synthetic and real data and present the tradeoffs between different variants of multimodal VAEs. We find that none of the existing approaches fulfills all desired criteria of an effective multimodal generative model when applied on more complex datasets than those used in previous benchmarks. In summary, we identify, formalize, and validate fundamental limitations of VAE-based approaches for modeling weakly-supervised data and discuss implications for real-world applications.

cs.LG