SearcharxivSearch

arXiv subjects

Thomas Hummel

Publications and source records attributed to Thomas Hummel.

17 recordsLinked to original sources

Woosh: A Sound Effects Foundation Model

The audio research community depends on open generative models as foundational tools for building novel approaches and establishing baselines. In this report, we present Woosh, Sony AI's publicly released sound effect foundation model, detailing its architecture, training process, and an evaluation against other popular open models. Being optimized for sound effects, we provide (1) a high-quality audio encoder/decoder model and (2) a text-audio alignment model for conditioning, together with (3) text-to-audio and (4) video-to-audio generative models. Distilled text-to-audio and video-to-audio models are also included in the release, allowing for low-resource operation and fast inference. Our evaluation on both public and private data shows competitive or better performance for each module when compared to existing open alternatives like StableAudio-Open and TangoFlux. Inference code and model weights are available at https://github.com/SonyResearch/Woosh. Demo samples can be found at https://sonyresearch.github.io/Woosh/.

cs.SD

Interfacing superconducting nanowire single photon detectors with cryogenic opto-electronics for quantum photonic applications

Interfacing single-photon detectors with active photonic components is a cornerstone photonic quantum technology. We describe how the output signal of commercial superconducting nanowire single-photon detectors can be used in situ to drive photonic components such as lasers and electro-optic modulators, co-located in the cryostat. This is enabled by developing custom circuitry using cryogenic-compatible discrete components in the SiGe-BiCMOS platform. We have demonstrated this with a number of experiments, in particular optical readout of an SNSPD and low-latency feed-forward modulation based on single-photon measurement events, all at or below 4 K. This manuscript is an abridged version of the Master thesis of the primary author N. Lamberty.

quant-ph

Cryogenic Feedforward of a Photonic Quantum State

Modulation conditioned on measurements on entangled photonic quantum states is a cornerstone technology of optical quantum information processing. Performing this task with low latency requires combining single-photon-level detectors with both electronic logic processing and optical modulation in close proximity. In the technologically relevant telecom wavelength band, detection of photonic quantum states is best performed with high-efficiency, low-noise, and high-speed detectors based on the photon-induced breakdown of superconductivity. Therefore, using these devices for feedforward requires mutual compatibility of all components under cryogenic conditions. Here, we demonstrate low-latency feedforward using a quasi-photon-number-resolved measurement on a quantum light source. Specifically, we use a multipixel superconducting nanowire single-photon detector, amplifier, logic, and an integrated electro-optic modulator in situ below 4K. We modulate the signal mode of a spontaneous parametric down-conversion source, conditional on a photon-number measurement of the idler mode, with a total latency of (23+/-3)ns. The photon-number discrimination actively manipulates the signal mode photon statistics, which is itself a central component in photonic quantum computing reliant on heralded single-photon sources. This represents an important benchmark for the fastest quantum photonic feedforward experiments comprising measurement, amplification, logic and modulation. This has direct applications in quantum computing, communication, and simulation protocols.

quant-ph

EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this challenging task, the first step is to acquire large-scale training datasets and collect high-quality benchmarks for evaluation. In this work, we introduce EgoCVR, a new evaluation benchmark for fine-grained Composed Video Retrieval using large-scale egocentric video datasets. EgoCVR consists of 2,295 queries that specifically focus on high-quality temporal video understanding. We find that existing Composed Video Retrieval frameworks do not achieve the necessary high-quality temporal video understanding for this task. To address this shortcoming, we adapt a simple training-free method, propose a generic re-ranking framework for Composed Video Retrieval, and demonstrate that this achieves strong results on EgoCVR. Our code and benchmark are freely available at https://github.com/ExplainableML/EgoCVR.

cs.CV

Optical Bias and Cryogenic Laser Readout of a Multipixel Superconducting Nanowire Single Photon Detector

Cryogenic opto-electronic interconnects are gaining increasing interest as a means to control and read out cryogenic electronic components. The challenge is to achieve sufficient signal integrity with low heat load processing. In this context, we demonstrate the opto-electronic bias and readout of a commercial four-pixel superconducting nanowire single-photon detector array using a cryogenic photodiode and laser. We show that this approach has a similar system detection efficiency to a conventional bias. Furthermore, multi-pixel detection events are faithfully converted between the optical and electrical domain, which allows reliable extraction of amplitude multiplexed photon statistics. Our device has a passive heat dissipation of 2.6mW, maintains the signal rise time of 3ns, and operates in free-running (self-resetting) mode at a repetition rate of 600kHz. This demonstrates the potential of high-bandwidth, low noise, and low heat load opto-electronic interconnects for scalable cryogenic signal processing and transmission.

physics.optics

How well can superconducting nanowire single-photon detectors resolve photon number?

We apply principal component analysis (PCA) to a set of electrical output signals from a commercially available superconducting nanowire single-photon detector (SNSPD) to investigate their photon-number-resolving capability. We find that the rising edge as well as the amplitude of the electrical signal have the most dependence on photon number. Accurately measuring the rising edge while simultaneously measuring the voltage of the pulse amplitude maximizes the photon-number resolution of SNSPDs. Using an optimal basis of principle components, we show unambiguous discrimination between one- and two-photon events, as well as partial resolution up to five photons. This expands the use-case of SNSPDs to photon-counting experiments, without the need of detector multiplexing architectures.

quant-ph

Video-adverb retrieval with compositional adverb-action embeddings

Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice versa) that aligns video embeddings with their matching compositional adverb-action text embedding in a joint embedding space. The compositional adverb-action text embedding is learned using a residual gating mechanism, along with a novel training objective consisting of triplet losses and a regression target. Our method achieves state-of-the-art performance on five recent benchmarks for video-adverb retrieval. Furthermore, we introduce dataset splits to benchmark video-adverb retrieval for unseen adverb-action compositions on subsets of the MSR-VTT Adverbs and ActivityNet Adverbs datasets. Our proposed framework outperforms all prior works for the generalisation task of retrieving adverbs from videos for unseen adverb-action compositions. Code and dataset splits are available at https://hummelth.github.io/ReGaDa/.

cs.CV

Text-to-feature diffusion for audio-visual few-shot learning

Training deep learning models for video classification from audio-visual data commonly requires immense amounts of labeled training data collected via a costly process. A challenging and underexplored, yet much cheaper, setup is few-shot learning from video data. In particular, the inherently multi-modal nature of video data with sound and visual information has not been leveraged extensively for the few-shot video classification task. Therefore, we introduce a unified audio-visual few-shot video classification benchmark on three datasets, i.e. the VGGSound-FSL, UCF-FSL, ActivityNet-FSL datasets, where we adapt and compare ten methods. In addition, we propose AV-DIFF, a text-to-feature diffusion framework, which first fuses the temporal and audio-visual features via cross-modal attention and then generates multi-modal features for the novel classes. We show that AV-DIFF obtains state-of-the-art performance on our proposed benchmark for audio-visual (generalised) few-shot learning. Our benchmark paves the way for effective audio-visual classification when only limited labeled data is available. Code and data are available at https://github.com/ExplainableML/AVDIFF-GFSL.

cs.CV

Pyroelectric Influence on Lithium Niobate During the Thermal Transition for Cryogenic Integrated Photonics

Lithium niobate has emerged as a promising platform for integrated quantum optics, enabling efficient generation, manipulation, and detection of quantum states of light. However, integrating single-photon detectors requires cryogenic operating temperatures, since the best performing detectors are based on narrow superconducting wires. While previous studies have demonstrated the operation of quantum light sources and electro-optic modulators in LiNbO3 at cryogenic temperatures, the thermal transition between room temperature and cryogenic conditions introduces additional effects that can significantly influence device performance. In this paper, we investigate the generation of pyroelectric charges and their impact on the optical properties of lithium niobate waveguides when changing from room temperature to 25K, and vice versa. We measure the generated pyroelectric charge flow and correlate this with fast changes in the birefringence acquired through the Senarmont method. Both electrical and optical influence of the pyroelectric effect occurs predominantly at temperatures above 100K.

physics.optics

All optical operation of a superconducting photonic interface

Advanced electro-optic processing combines electrical control with optical modulation and detection. For quantum photonic applications these processes must be carried out at the single photon level with high efficiency and low noise. Integrated quantum photonics has made great strides achieving single photon manipulation by combining key components on integrated chips which are operated by external driving electronics. Nevertheless, electrical interconnects between driving electronics and the electro-optic components, some of which require cryogenic operating conditions, can introduce parasitic effects. Here we show an all-optical interface which simultaneously delivers the operation power to, and extracts the measurement signal from, an advanced photonic circuit, namely, bias and readout of a superconducting nanowire single photon detector (SNSPD) on a single stage in a 1K cryostat. To do so, we supply all power for the single photon detector, output signal conditioning, and electro-optic readout using optical interconnects alone, thereby fully decoupling the cryogenic circuitry from the external environment. This removes the need to heatsink electrical connections, and potentially offers low-loss, high-bandwidth signal processing. This method opens the possibility to operate other advanced electrically decoupled photonic circuits such as optical control and readout of superconducting circuits, and feedforward for photonic quantum computing.

quant-ph

Nanosecond gating of superconducting nanowire single-photon detectors using cryogenic bias circuitry

Superconducting nanowire single-photon detectors (SNSPDs) show near unity efficiency, low dark count rate, and short recovery time. Combining these characteristics with temporal control of SNSPDs broadens their applications as in active de-latching for higher dynamic range counting or temporal filtering for pump-probe spectroscopy or LiDAR. To that end, we demonstrate active gating of an SNSPD with a minimum off-to-on rise time of 2.4 ns and a total gate length of 5.0 ns. We show how the rise time depends on the inductance of the detector in combination with the control electronics. The gate window is demonstrated to be fully and freely, electrically tunable up to 500 ns at a repetition rate of 1.0 MHz, as well as ungated, free-running operation. Control electronics to generate the gating are mounted on the 2.3 K stage of a closed-cycle sorption cryostat, while the detector is operated on the cold stage at 0.8 K. We show that the efficiency and timing jitter of the detector is not altered during the on-time of the gating window. We exploit gated operation to demonstrate a method to increase in the photon counting dynamic range by a factor 11.2, as well as temporal filtering of a strong pump in an emulated pump-probe experiment.

physics.ins-det

Semantic Image Synthesis with Semantically Coupled VQ-Model

Semantic image synthesis enables control over unconditional image generation by allowing guidance on what is being generated. We conditionally synthesize the latent space from a vector quantized model (VQ-model) pre-trained to autoencode images. Instead of training an autoregressive Transformer on separately learned conditioning latents and image latents, we find that jointly learning the conditioning and image latents significantly improves the modeling capabilities of the Transformer model. While our jointly trained VQ-model achieves a similar reconstruction performance to a vanilla VQ-model for both semantic and image latents, tying the two modalities at the autoencoding stage proves to be an important ingredient to improve autoregressive modeling performance. We show that our model improves semantic image synthesis using autoregressive models on popular semantic image datasets ADE20k, Cityscapes and COCO-Stuff.

cs.CV

Temporal and cross-modal attention for audio-visual zero-shot learning

Audio-visual generalised zero-shot learning for video classification requires understanding the relations between the audio and visual information in order to be able to recognise samples from novel, previously unseen classes at test time. The natural semantic and temporal alignment between audio and visual data in video data can be exploited to learn powerful representations that generalise to unseen classes at test time. We propose a multi-modal and Temporal Cross-attention Framework (\modelName) for audio-visual generalised zero-shot learning. Its inputs are temporally aligned audio and visual features that are obtained from pre-trained networks. Encouraging the framework to focus on cross-modal correspondence across time instead of self-attention within the modalities boosts the performance significantly. We show that our proposed framework that ingests temporal features yields state-of-the-art performance on the \ucf, \vgg, and \activity benchmarks for (generalised) zero-shot learning. Code for reproducing all results is available at \url{https://github.com/ExplainableML/TCAF-GZSL}.

cs.CV

Cryogenic electro-optic modulation in titanium in-diffused lithium niobate waveguides

Lithium niobate is a promising platform for integrated quantum optics. In this platform we aim to efficiently manipulate and detect quantum states by combining superconducting single photon detectors and modulators. The cryogenic operation of a superconducting single photon detector dictates the optimisation of the electro-optic modulators under the same operating conditions. To that end, we characterise a phase modulator, directional coupler, and polarisation converter at both ambient and cryogenic temperatures. The operation voltage $V_{π/2}$ of these modulators increases due to the decrease of the electro-optic effect by 74% for the phase modulator, 84% for the directional coupler and 35% for the polarisation converter below 8.5$\,\mathrm{K}$. The phase modulator preserves its broadband nature and modulates light in the characterised wavelength range. The unbiased bar state of the directional coupler changed by a wavelength shift of 85$\,\mathrm{nm}$ while cooling the device down to 5$\,\mathrm{K}$. The polarisation converter uses periodic poling to phasematch the two orthogonal polarisations. The phasematched wavelength of the used poling changes by 112$\,\mathrm{nm}$ when cooling to 5$\,\mathrm{K}$

physics.optics

Where and When: Space-Time Attention for Audio-Visual Explanations

Explaining the decision of a multi-modal decision-maker requires to determine the evidence from both modalities. Recent advances in XAI provide explanations for models trained on still images. However, when it comes to modeling multiple sensory modalities in a dynamic world, it remains underexplored how to demystify the mysterious dynamics of a complex multi-modal model. In this work, we take a crucial step forward and explore learnable explanations for audio-visual recognition. Specifically, we propose a novel space-time attention network that uncovers the synergistic dynamics of audio and visual data over both space and time. Our model is capable of predicting the audio-visual video events, while justifying its decision by localizing where the relevant visual cues appear, and when the predicted sounds occur in videos. We benchmark our model on three audio-visual video event datasets, comparing extensively to multiple recent multi-modal representation learners and intrinsic explanation models. Experimental results demonstrate the clear superior performance of our model over the existing methods on audio-visual video event recognition. Moreover, we conduct an in-depth study to analyze the explainability of our model based on robustness analysis via perturbation tests and pointing games using human annotations.

cs.CV

Crossmodal Language Grounding in an Embodied Neurocognitive Model

Human infants are able to acquire natural language seemingly easily at an early age. Their language learning seems to occur simultaneously with learning other cognitive functions as well as with playful interactions with the environment and caregivers. From a neuroscientific perspective, natural language is embodied, grounded in most, if not all, sensory and sensorimotor modalities, and acquired by means of crossmodal integration. However, characterising the underlying mechanisms in the brain is difficult and explaining the grounding of language in crossmodal perception and action remains challenging. In this paper, we present a neurocognitive model for language grounding which reflects bio-inspired mechanisms such as an implicit adaptation of timescales as well as end-to-end multimodal abstraction. It addresses developmental robotic interaction and extends its learning capabilities using larger-scale knowledge-based data. In our scenario, we utilise the humanoid robot NICO in obtaining the EMIL data collection, in which the cognitive robot interacts with objects in a children's playground environment while receiving linguistic labels from a caregiver. The model analysis shows that crossmodally integrated representations are sufficient for acquiring language merely from sensory input through interaction with objects in an environment. The representations self-organise hierarchically and embed temporal and spatial information through composition and decomposition. This model can also provide the basis for further crossmodal integration of perceptually grounded cognitive representations.

cs.NE

Efficient demultiplexed single-photon source with a quantum dot coupled to a nanophotonic waveguide

Planar nanostructures allow near-ideal extraction of emission from a quantum emitter embedded within, thereby realizing deterministic single-photon sources. Such a source can be transformed into M single-photon sources by implementing active temporal-to-spatial mode demultiplexing. We report on the realization of such a demultiplexed source based on a quantum dot embedded in a nanophotonic waveguide. Efficient outcoupling (>60%) from the waveguide into a single mode optical fiber is obtained with high-efficiency grating couplers. As a proof-of-concept, active demultiplexing into M=4 spatial channels is demonstrated by the use of electro-optic modulators with an end-to-end efficiency of >81% into single-mode fibers. Overall we demonstrate four-photon coincidence rates of >1 Hz even under non-resonant excitation of the quantum dot. The main limitation of the current source is the residual population of other exciton transitions that corresponds to a finite preparation efficiency of the desired transition. We quantitatively extract a preparation efficiency of 15% using the second-order correlation function measurements. The experiment highlights the applicability of planar nanostructures as efficient multiphoton sources through temporal-to-spatial demultiplexing and lays out a clear path way of how to scale up towards demonstrating quantum advantages with the quantum dot sources.

quant-ph