SearcharxivSearch

arXiv subjects

Minu Kim

Publications and source records attributed to Minu Kim.

At least 19 recordsLinked to original sources

Language Orthogonalization of Self-Supervised Speech Representations for Cross-lingual Parkinson's Detection

Self-supervised speech models (S3Ms) provide powerful representations for Parkinson's disease (PD) detection, making cross-lingual transfer attractive for languages lacking labeled patient speech. However, these representations also encode language identity, which can confound this transfer: without target-language PD speech, classifiers may separate languages rather than pathology, yielding high specificity but low sensitivity on target patients. We propose \emph{language orthogonalization}, a closed-form ridge residualization of S3M features against external VoxLingua107 language embeddings, fitted using only healthy-control (HC) speech. By removing language-predictable components while retaining pathology-related variation, it produces a less language-dependent geometry in which HC representations concentrate while PD representations disperse. Across five S3M backbones, three speech tasks, and three target languages, our method consistently improves cross-lingual PD-detection performance while correcting the high-specificity/low-sensitivity failure.

eess.AS

A Hierarchical Feature Engineering Framework for Automated Classification of Phonotraumatic and Non-Phonotraumatic Vocal Hyperfunction

Ambulatory neck-surface acceleration enables non-invasive monitoring of vocal hyperfunction, yet robust biomarkers for its subtypes remain limited. This study investigates the NeckVibe Challenge dataset to distinguish phonotraumatic (PVH) and non-phonotraumatic (NPVH) from healthy controls. We propose a hierarchical feature engineering framework comprising: (i) static, (ii) dynamic, (iii) ratio-based, (iv) coupling features capturing source filter interactions. While univariate statistical analysis shows strong separability for PVH but limited significance for NPVH, our machine learning pipeline, tailored for high-dimensional feature integration, identifies that coupling features are crucial for both tasks. We achieve an AUC of 0.891 for PVH and 0.728 for NPVH, suggesting that while PVH is near-linearly separable, NPVH discrimination benefits from modeling non-linear feature interactions.

cs.SD

EXAONE 4.5 Technical Report

This technical report introduces EXAONE 4.5, the first open-weight vision language model released by LG AI Research. EXAONE 4.5 is architected by integrating a dedicated visual encoder into the existing EXAONE 4.0 framework, enabling native multimodal pretraining over both visual and textual modalities. The model is trained on large-scale data with careful curation, particularly emphasizing document-centric corpora that align with LG's strategic application domains. This targeted data design enables substantial performance gains in document understanding and related tasks, while also delivering broad improvements across general language capabilities. EXAONE 4.5 extends context length up to 256K tokens, facilitating long-context reasoning and enterprise-scale use cases. Comparative evaluations demonstrate that EXAONE 4.5 achieves competitive performance in general benchmarks while outperforming state-of-the-art models of similar scale in document understanding and Korean contextual reasoning. As part of LG's ongoing effort toward practical industrial deployment, EXAONE 4.5 is designed to be continuously extended with additional domains and application scenarios to advance AI for a better life.

cs.CL

Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster

Similarities between language representations derived from Self-Supervised Speech Models (S3Ms) have been observed to primarily reflect geographic proximity or surface typological similarities driven by recent expansion or contact, potentially missing deeper genealogical signals. We investigate how scaling an S3M-based language identification system from 126 to 4,017 languages reshapes this topology, and find a non-linear effect: phylogenetic recovery stays flat up to the 1K scale, but the 4K model undergoes a qualitative shift, resolving both clear lineages and long-term linguistic contact. Most strikingly, a robust Pacific macro-cluster emerges, grouping genealogically unrelated Papuan, Oceanic, and Australian languages, and we trace its driver to a concentrated encoding that captures shared acoustic signatures such as global energy dynamics. These results suggest that massive S3Ms internalize multiple layers of language history, offering a promising perspective for computational phylogenetics and the study of language contact.

cs.CL

K-EXAONE Technical Report

This technical report presents K-EXAONE, a large-scale multilingual language model developed by LG AI Research. K-EXAONE is built on a Mixture-of-Experts architecture with 236B total parameters, activating 23B parameters during inference. It supports a 256K-token context window and covers six languages: Korean, English, Spanish, German, Japanese, and Vietnamese. We evaluate K-EXAONE on a comprehensive benchmark suite spanning reasoning, agentic, general, Korean, and multilingual abilities. Across these evaluations, K-EXAONE demonstrates performance comparable to open-weight models of similar size. K-EXAONE, designed to advance AI for a better life, is positioned as a powerful proprietary AI foundation model for a wide range of industrial and research applications.

cs.CL

How Far Do SSL Speech Models Listen for Tone? Temporal Focus of Tone Representation under Low-resource Transfer

Lexical tone is central to many languages but remains underexplored in self-supervised learning (SSL) speech models, especially beyond Mandarin. We study four languages with complex and diverse tone systems (Burmese, Thai, Lao, and Vietnamese) to ask how far such models "listen" for tone and how transfer operates in low-resource conditions. As a baseline reference, we estimate the temporal span of tone cues: approximately 100ms (Burmese/Thai) and 180ms (Lao/Vietnamese). Probes and gradient analysis on fine-tuned SSL models reveal that tone transfer varies by downstream task: automatic speech recognition fine-tuning aligns spans with language-specific tone cues, while prosody- and voice-related tasks bias toward overly long spans. These findings indicate that tone transfer is shaped by downstream task, highlighting task effects on temporal focus in tone modeling.

eess.AS

ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction

Noise-robust speaker verification leverages joint learning of speech enhancement (SE) and speaker verification (SV) to improve robustness. However, prevailing approaches rely on implicit noise suppression, which struggles to separate noise from speaker characteristics as they do not explicitly distinguish noise from speech during training. Although integrating SE and SV helps, it remains limited in handling noise effectively. Meanwhile, recent SE studies suggest that explicitly modeling noise, rather than merely suppressing it, enhances noise resilience. Reflecting this, we propose ParaNoise-SV, with dual U-Nets combining a noise extraction (NE) network and a speech enhancement (SE) network. The NE U-Net explicitly models noise, while the SE U-Net refines speech with guidance from NE through parallel connections, preserving speaker-relevant features. Experimental results show that ParaNoise-SV achieves a relatively 8.4% lower equal error rate (EER) than previous joint SE-SV models.

eess.AS

Guiding Reasoning in Small Language Models with LLM Assistance

The limited reasoning capabilities of small language models (SLMs) cast doubt on their suitability for tasks demanding deep, multi-step logical deduction. This paper introduces a framework called Small Reasons, Large Hints (SMART), which selectively augments SLM reasoning with targeted guidance from large language models (LLMs). Inspired by the concept of cognitive scaffolding, SMART employs a score-based evaluation to identify uncertain reasoning steps and injects corrective LLM-generated reasoning only when necessary. By framing structured reasoning as an optimal policy search, our approach steers the reasoning trajectory toward correct solutions without exhaustive sampling. Our experiments on mathematical reasoning datasets demonstrate that targeted external scaffolding significantly improves performance, paving the way for collaborative use of both SLM and LLM to tackle complex reasoning tasks that are currently unsolvable by SLMs alone.

cs.CL

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language selection. Previous cross-lingual research has used various source languages to enhance performance for the target low-resource language without thorough consideration of selection. Our study stands out by providing an in-depth analysis of language selection, supported by a practical approach to assess phonetic proximity among multiple language families. We investigate how within-family similarity impacts performance in multilingual training, which aids in understanding language dynamics. We also evaluate the effect of using phonologically similar languages, regardless of family. For the phoneme recognition task, utilizing phonologically similar languages consistently achieves a relative improvement of 55.6% over monolingual training, even surpassing the performance of a large-scale self-supervised learning model. Multilingual training within the same language family demonstrates that higher phonological similarity enhances performance, while lower similarity results in degraded performance compared to monolingual training.

eess.AS

Compositional Phoneme Approximation for L1-Grounded L2 Pronunciation Training

Learners of a second language (L2) often map non-native phonemes to similar native-language (L1) phonemes, making conventional L2-focused training slow and effortful. To address this, we propose an L1-grounded pronunciation training method based on compositional phoneme approximation (CPA), a feature-based representation technique that approximates L2 sounds with sequences of L1 phonemes. Evaluations with 20 Korean non-native English speakers show that CPA-based training achieves a 76% in-box formant rate in acoustic analysis, 17.6% relative improvement in phoneme recognition accuracy, and over 80% of speech being rated as more native-like, with minimal training. Project page: https://gsanpark.github.io/CPA-Pronunciation.

cs.CL

Impact of synthesis method on the structure and function of high entropy oxides

The term sample dependence describes the troublesome tendency of nominally equivalent samples to exhibit different physical properties. High entropy oxides (HEOs) are a class of materials where sample dependence has the potential to be particularly profound due to their inherent chemical complexity. In this work, we prepare a spinel HEO of identical nominal composition by five distinct methods, spanning a range of thermodynamic and kinetic conditions: solid state, high pressure, hydrothermal, molten salt, and combustion syntheses. By structurally characterizing these five samples across all length scales with a variety of x-ray methods, we find that while the average structure is unaltered, the samples vary significantly in their local structures and their microstructures. The most profound differences are observed at intermediate length scales, both in terms of crystallite morphology and cation homogeneity. As revealed by x-ray fluorescence microscopy ideal cation homogeneity is achieved only in the case of combustion synthesis. These structural differences in turn significantly alter the observed functional properties, which we demonstrate via characterization of their magnetic response. While ferrimagnetic order is retained across all five samples, the sharpness of the transition, the size of the saturated moment, and the coercivity all show marked variations with synthesis method. We conclude that the chemical flexibility inherent to HEOs is complemented by strong synthesis method dependence, providing another axis along which to optimize these materials for a wide range of applications.

cond-mat.mtrl-sci

Effect of high pressure synthesis conditions on the formation of high entropy oxides

High entropy materials are often entropy stabilized, meaning that the configurational entropy from multiple elements sharing a single lattice site stabilizes the structure. In this work, we study how high-pressure synthesis conditions can stabilize or destabilize a high entropy oxide (HEO). We study the high-pressure and high-temperature phase equilibria of two well-known families of HEOs: the rock-salt structured compound (Mg,Co,Ni,Cu,Zn)O including some cation substitutions and the spinel structured (Cr,Mn,Fe,Co,Ni)$_3$O$_4$. Syntheses were performed at various temperatures, pressures, and oxygen activity levels resulting in dramatically different synthesis outcomes. In particular, in the rock salt HEO we observe the competing tenorite and wurtzite phases and the possible formation of a layered rock salt phase, while the spinel HEO is highly susceptible to decomposition into a mixture of rock-salt and corundum phases. At the highest tested pressures, 15 GPa, we discover the transformation of the spinel HEO into a metastable modified ludwigite-type structure with nominal formula (Cr,Mn,Fe,Co,Ni)$_4$O$_5$. The relationship between the synthesis conditions and the final reaction product is not straightforward. Nonetheless, we conclude that high-pressure conditions provide an important opportunity to synthesize high entropy phases that cannot be formed any other way.

cond-mat.mtrl-sci

Preference Alignment with Flow Matching

We present Preference Flow Matching (PFM), a new framework for preference-based reinforcement learning (PbRL) that streamlines the integration of preferences into an arbitrary class of pre-trained models. Existing PbRL methods require fine-tuning pre-trained models, which presents challenges such as scalability, inefficiency, and the need for model modifications, especially with black-box APIs like GPT-4. In contrast, PFM utilizes flow matching techniques to directly learn from preference data, thereby reducing the dependency on extensive fine-tuning of pre-trained models. By leveraging flow-based models, PFM transforms less preferred data into preferred outcomes, and effectively aligns model outputs with human preferences without relying on explicit or implicit reward function estimation, thus avoiding common issues like overfitting in reward models. We provide theoretical insights that support our method's alignment with standard PbRL objectives. Experimental results indicate the practical effectiveness of our method, offering a new direction in aligning a pre-trained model to preference. Our code is available at https://github.com/jadehaus/preference-flow-matching.

cs.LG

Crystallization of heavy fermions via epitaxial strain in spinel LiV$_{2}$O$_{4}$ thin film

The mixed-valent spinel LiV$_{2}$O$_{4}$ is known as the first oxide heavy-fermion system. There is a general consensus that a subtle interplay of charge, spin, and orbital degrees of freedom of correlated electrons plays a crucial role in the enhancement of quasi-particle mass, but the specific mechanism has remained yet elusive. A charge-ordering (CO) instability of V$^{3+}$ and V$^{4+}$ ions that is geometrically frustrated by the V pyrochlore sublattice from forming a long-range CO down to $T$ = 0 K has been proposed as a prime candidate for the mechanism. To uncover the hidden CO instability, we applied epitaxial strain from a substrate on single-crystalline thin films of LiV$_{2}$O$_{4}$. Here we show a strain-induced crystallization of heavy fermions in a LiV$_{2}$O$_{4}$ film on MgO, where a charge-ordered insulator comprising of a stack of V$^{3+}$ and V$^{4+}$ layers along [001], the historical Verwey-type ordering, is stabilized by the in-plane tensile and out-of-plane compressive strains from the substrate. Our discovery of the [001] Verwey-type CO, together with previous realizations of a distinct [111] CO, evidence the close proximity of the heavy-fermion state to degenerate CO states mirroring the geometrical frustration of the pyrochlore lattice, which supports the CO instability scenario for the mechanism behind the heavy-fermion formation.

cond-mat.str-el

Discovery of Superconductivity in (Ba,K)SbO$_{3}$

Superconducting bismuthates (Ba,K)BiO$_{3}$ (BKBO) constitute an interesting class of superconductors in that superconductivity with a remarkably high $T_\mathrm{c}$ of 30 K arises in proximity to charge density wave (CDW) order. Prior understanding on the driving mechanism of the CDW and superconductivity emphasizes the role of either bismuth (negative $U$ model) or oxygen ions (ligand hole model). While holes in BKBO presumably reside on oxygen owing to their negative charge transfer energy, so far there has been no other comparative material studied. Here, we introduce (Ba,K)SbO$_{3}$ (BKSO) in which the Sb 5$s$ orbital energy is higher than that of the Bi 6$s$ orbitals enabling tuning of the charge transfer energy from negative to slightly positive. The parent compound BaSbO$_{3-\delta}$ shows a larger CDW gap compared to the undoped bismuthate BaBiO$_{3}$. As the CDW order is suppressed via potassium substitution up to 65 %, superconductivity emerges, rising up to $T_\mathrm{c}$ = 15 K. This value is lower than the maximum $T_\mathrm{c}$ of BKBO, but higher by more than a factor of two at comparable potassium concentrations. The discovery of an enhanced CDW gap and superconductivity in BKSO indicates that the sign of the charge transfer energy may not be crucial, but instead strong metal-oxygen covalency plays the essential role in constituting a CDW and high-$T_\mathrm{c}$ superconductivity in the main-group perovskite oxides.

cond-mat.supr-con

Electronic Structure of the Bond Disproportionated Bismuthate Ag$_2$BiO$_3$

We present a comprehensive study on the silver bismuthate Ag$_2$BiO$_3$, synthesized under high-pressure high-temperature conditions, which has been the subject of recent theoretical work on topologically complex electronic states. We present X-ray photoelectron spectroscopy results showing two different bismuth states, and X-ray absorption spectroscopy results on the oxygen $K$-edge showing holes in the oxygen bands. These results support a bond disproportionated state with holes on the oxygen atoms for Ag$_2$BiO$_3$. We estimate a band gap of $\sim$1.25~eV for Ag$_2$BiO$_3$ from optical conductivity measurements, which matches the band gap in density functional calculations of the electronic band structure in the non-symmorphic space group $Pnn2$, which supports two inequivalent Bi sites. In our band structure calculations the disproportionated Ag$_2$BiO$_3$ is expected to host Weyl nodal chains, one of which is located $\sim$0.5~eV below the Fermi level. Furthermore, we highlight similarities between Ag$_2$BiO$_3$ and the well-known disproportionated bismuthate BaBiO$_3$, including breathing phonon modes with similar energy. In both compounds hybridization of Bi-$6s$ and O-$2p$ atomic orbitals is important in shaping the band structure, but in contrast to the Ba-$5p$ in BaBiO$_3$, the Ag-$4d$ bands in Ag$_2$BiO$_3$ extend up to the Fermi level.

cond-mat.mtrl-sci

Universal bound to the amplitude of the vortex Nernst signal in superconductors

A liquid of superconducting vortices generates a transverse thermoelectric response. This Nernst signal has a tail deep in the normal state due to superconducting fluctuations. Here, we present a study of the Nernst effect in two-dimensional hetero-structures of Nb-doped strontium titanate (STO) and in amorphous MoGe. The Nernst signal generated by ephemeral Cooper pairs above the critical temperature has the magnitude expected by theory in STO. On the other hand, the peak amplitude of the vortex Nernst signal below $T_c$ is comparable in both and in numerous other superconductors despite the large distribution of the critical temperature and the critical magnetic fields. In four superconductors belonging to different families, the maximum Nernst signal corresponds to an entropy per vortex per layer of $\approx$ k$_Bln2$.

cond-mat.supr-con

Observation of the signatures of sub-resolution defects in two-dimensional superconductors with scanning SQUID

The diamagnetic susceptibility of a superconductor is directly related to its superfluid density. Mutual inductance is a highly sensitive method for characterizing thin films; however, in traditional mutual inductance measurements, the measured response is a non-trivial average over the area of the mutual inductance coils, which are typically of millimeter size. Here we image localized, isolated features in the diamagnetic susceptibility of δ-doped SrTiO3, the 2-DES at the interface between LaAlO3 and SrTiO3, and Nb superconducting thin film systems using scanning superconducting quantum interference device susceptometry, with spatial resolution as fine as 0.7 μm. We show that these features can be modeled as locally suppressed superfluid density, with a single parameter that characterizes the strength of each feature. This method provides a systematic means of finding and quantifying submicron defects in two-dimensional superconductors.

cond-mat.supr-con