SearcharxivSearch

arXiv subjects

Taewoo Kang

Publications and source records attributed to Taewoo Kang.

4 recordsLinked to original sources

Hilbert entropy for measuring the complexity of high-dimensional systems

Measuring the complexity of high-dimensional data in physical systems becomes a critical factor in determining the information and quality of the systems. However, traditional metrics, such as Lyapunov exponent, fractal dimension, and information entropy, are limited in measuring contextual higher-dimensional data in that they do not elucidate the intrinsic nature of physical systems. Herein, we introduce a novel methodology for quantifying the complexity of high-dimensional data through dimension reduction yet retaining context using a space-filling curve such as the Hilbert curve along with generalized entropy measures. We validate this methodology in measuring critical phenomena, including phase transitions in spin and percolation models. Our findings demonstrate a high degree of concordance between the Hilbert entropy and theoretical phase transition points. Moreover, we further proceed to an exploration of the hidden relationship between the Hilbert entropy and the fractal dimension, such as a linear relationship between scaling exponent and the Euclidean dimension of scale-invariant 2D/3D geometries. The present methodology offers a promising new framework for understanding and analyzing complex systems in higher dimensions, with potential applications across various fields of physics.

physics.soc-ph

M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models

We present M3-SLU, a new multimodal large language model (MLLM) benchmark for evaluating multi-speaker, multi-turn spoken language understanding. While recent models show strong performance in speech and text comprehension, they still struggle with speaker-attributed reasoning, the ability to understand who said what and when in natural conversations. M3-SLU is built from four open corpora (CHiME-6, MELD, MultiDialog, and AMI) and comprises over 12,000 validated instances with paired audio, transcripts, and metadata. It includes two tasks: (1) Speaker-Attributed Question Answering and (2) Speaker Attribution via Utterance Matching. We provide baseline results for both cascaded pipelines and end-to-end MLLMs, evaluated using an LLM-as-Judge and accuracy metrics. Results show that while models can capture what was said, they often fail to identify who said it, revealing a key gap in speaker-aware dialogue understanding. M3-SLU offers as a challenging benchmark to advance research in speaker-aware multimodal understanding.

cs.CL

Embracing Dialectic Intersubjectivity: Coordination of Different Perspectives in Content Analysis with LLM Persona Simulation

This study attempts to advancing content analysis methodology from consensus-oriented to coordination-oriented practices, thereby embracing diverse coding outputs and exploring the dynamics among differential perspectives. As an exploratory investigation of this approach, we evaluate six GPT-4o configurations to analyze sentiment in Fox News and MSNBC transcripts on Biden and Trump during the 2020 U.S. presidential campaign, examining patterns across these models. By assessing each model's alignment with ideological perspectives, we explore how partisan selective processing could be identified in LLM-Assisted Content Analysis (LACA). Findings reveal that partisan persona LLMs exhibit stronger ideological biases when processing politically congruent content. Additionally, intercoder reliability is higher among same-partisan personas compared to cross-partisan pairs. This approach enhances the nuanced understanding of LLM outputs and advances the integrity of AI-driven social science research, enabling simulations of real-world implications.

cs.CL

STAR: Spatio-Temporal Prediction of Air Quality Using A Multimodal Approach

With the increase of global economic activities and high energy demand, many countries have raised concerns about air pollution. However, air quality prediction is a challenging issue due to the complex interaction of many factors. In this paper, we propose a multimodal approach for spatio-temporal air quality prediction. Our model learns the multimodal fusion of critical factors to predict future air quality levels. Based on the analyses of data, we also assessed the impacts of critical factors on air quality prediction. We conducted experiments on two real-world air pollution datasets. For Seoul dataset, our method achieved 11% and 8.2% improvement of the mean absolute error in long-term predictions of PM2.5 and PM10, respectively, compared to baselines. Our method also reduced the mean absolute error of PM2.5 predictions by 20% compared to the previous state-of-the-art results on China 1-year dataset.

eess.SP