SearcharxivSearch

arXiv subjects

William Ho

Publications and source records attributed to William Ho.

6 recordsLinked to original sources

When Vocal Tone and Literal Meaning Diverge: An Acoustic-Semantic Incongruity Study for Large Audio-Language Models

Affective cues across modalities may be incongruous (e.g., sarcasm or mocking praise), potentially leading to misinterpretation when relying on a single modality. Large Audio-Language Models (LALMs) have recently gained popularity and been applied to multimodal emotion recognition, but their ability to disentangle acoustic and semantic cues, especially in incongruent cases, remains underexplored. To address this gap, we introduce CREMA-ASIS, a dataset specifically created to investigate incongruence between acoustic emotion and semantic sentiment cues. It pairs acoustic emotion labels with semantic sentiment polarities. Using this dataset, we evaluate LALM biases within a multitask framework and conduct a layer-wise analysis to identify modality dominance across layers. Our findings reveal that LALMs struggle with semantic-acoustic incongruent cases, rarely predicting incongruity, and that LALMs are predominantly influenced by semantic information. However, supervised fine-tuning significantly improves LALM performance on our CREMA-ASIS test set while preserving transcription accuracy and joint emotion recognition. Results demonstrate potential for enhancing both acoustic and semantic understanding on out-of-domain data.

eess.AS

Dark matter searches with a 13 meV threshold superconducting sensor array

Many well-motivated dark matter models predict meV-scale energy deposits in interactions with terrestrial experiments, but this regime is challenging to probe due to a lack of mature single-quantum detectors. Here we report results from QUALIPHIDE (QUAntum LImited PHotons In the Dark Experiment), a cryogenic dark matter search using a $41$-pixel array of energy-resolving microwave kinetic inductance detectors with a $13$ meV threshold, simultaneously used to look for both conversion photons from THz wavelength hidden photon dark matter and phonons from particle-like light dark matter interactions. The experimental design, with on- and off-focus pixels for the hidden photon search, allows for a data-driven background model, giving the experiment discovery potential. A blind analysis of $22$ hours of data shows no significant excess, setting the strongest constraints on the hidden photon kinetic mixing parameter $\chi$ over the mass range of $13$-$90$ meV/$c^2$, reaching $1.5\times10^{-12}$ at $50$ meV/$c^2$. These data also yield among the first terrestrial limits on dark matter scattering off nuclei and electrons, down to $5$ MeV/$c^2$ and $20$ keV/$c^2$, respectively. The low threshold also enables future study of the low-energy excess limiting cryogenic detectors and, as we project, will allow for a terahertz-scale QCD axion search with a magnetic field.

hep-ex

Visualizing Mathieu-Type Dynamics in a Tabletop Magnetic Trap: A Coil-Driven Parametric Oscillator

We present a tabletop demonstration of dynamic stabilization and ponderomotive-like trapping using a pair of sinusoidally-driven anti-Helmholtz coils and a suspended permanent magnet. The oscillating field produces a rapid micromotion superimposed on a slower secular oscillation, with micromotion amplitude increasing with displacement and peaking near the turning points. This behavior reveals a ponderomotive-like mechanism: a spatial gradient of micromotion amplitude that drives slow secular motion. The time-averaged effect provides a time-averaged harmonic (ponderomotive) restoring force that confines the magnet between the coils. Driving at 12-18 Hz places the system in a small-q regime where the two time scales are clearly separated and directly visible to the eye. Video tracking (included with this article) quantifies the motion and reveals a stability edge as the drive frequency is lowered (near 6-7 Hz in our apparatus). From trajectories in the 12-18 Hz range, we extract an effective Mathieu parameter q ~ 0.16 from the measured timescale separation of the secular versus drive frequencies. The apparatus uses inexpensive, readily available parts, and we provide a concise materials list, analysis code, field-gradient calibration data, and demonstration videos.

physics.atom-ph

Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction

The growing demand for home healthcare calls for tools that can support care delivery. In this study, we explore automatic health assessment from voice using real-world home care visit data, leveraging the diverse patient information it contains. First, we utilize Large Language Models (LLMs) to integrate Subjective, Objective, Assessment, and Plan (SOAP) notes derived from unstructured audio transcripts and structured vital signs into a holistic illness score that reflects a patient's overall health. This compact representation facilitates cross-visit health status comparisons and downstream analysis. Next, we design a multi-stage preprocessing pipeline to extract short speech segments from target speakers in home care recordings for acoustic analysis. We then employ an Audio Language Model (ALM) to produce plain-language descriptions of vocal biomarkers and examine their association with individuals' health status. Our experimental results benchmark both commercial and open-source LLMs in estimating illness scores, demonstrating their alignment with actual clinical outcomes, and revealing that SOAP notes are substantially more informative than vital signs. Building on the illness scores, we provide the first evidence that ALMs can identify health-related acoustic patterns from home care recordings and present them in a human-readable form. Together, these findings highlight the potential of LLMs and ALMs to harness heterogeneous in-home visit data for better patient monitoring and care.

eess.AS

Development Status of the KIPM Detector Consortium

A Kinetic Inductance Phonon-Mediated Detector is a calorimeter that uses kinetic inductance detectors to read out phonon signals from the device substrate. We have established a consortium comprising university and national lab groups dedicated to advancing the state of the art in these detectors, with the ultimate goal of designing a detector sub-eV threshold on energy deposited in the substrate, enabling searches for both light dark matter and low-energy neutrino interactions. This consortium brings together experts in kinetic inductance detector design, phonon and quasiparticle dynamics, and noise modeling, along with specialized fabrication facilities, test platforms, and unique calibration capabilities. Recently, our consortium has demonstrated a resolution on energy absorbed by the sensor of 2.1 eV, the current record for such devices. The current focus of the consortium is modeling and improving the phonon collection efficiency and implementing low-$\boldsymbol{T_c}$ superconductors, both of which serve to improve the overall energy resolution and threshold of the detectors.

physics.ins-det

From Who Said What to Who They Are: Modular Training-free Identity-Aware LLM Refinement of Speaker Diarization

Speaker diarization (SD) remains challenging in real-world scenarios due to dynamic environments and unknown speaker numbers. SD is rarely used alone and is typically paired with automatic speech recognition (ASR). However, existing non-modular SD+ASR frameworks lack flexibility and do not provide true speaker identities. We propose a training-free modular pipeline combining off-the-shelf SD, ASR, and a large language model (LLM) to determine who spoke, what was said, and who they are. Using structured LLM prompting on reconciled SD and ASR outputs, our method leverages semantic continuity in conversational context to refine low-confidence speaker labels and assigns role identities while correcting split speakers. On a real-world patient-clinician dataset, our approach achieves a 29.7% relative error reduction over baseline reconciled SD and ASR. It enhances diarization performance without additional training and delivers a complete pipeline for SD, ASR, and speaker identity detection in practical applications.

eess.AS