SearcharxivSearch

arXiv subjects

Kawon Lee

Publications and source records attributed to Kawon Lee.

5 recordsLinked to original sources

Attractive statistical forces and Pauli crystal formation in trapped Fermi gases

Exchange statistics endows identical particles with an effective "statistical potential", whose familiar exact form is two-body and purely repulsive for fermions. Here we construct an exact collective many-body form: the thermodynamics of $N$ trapped ideal fermions maps onto classical distinguishable particles governed by a single potential -- exactly for harmonic confinement at all temperatures, and to leading semiclassical order for arbitrary potentials. The associated force separates canonically into pairwise contributions, which for $N\geq 3$ can turn attractive, governed by a simple geometric criterion. Classical minimization reproduces observed few-body Pauli-crystal symmetries and agrees with the $N=55$ ground-state probability maximum at sub-percent shell accuracy. Heating drives discrete structural transitions accompanied by a crossover of the strongest force from attractive to repulsive. Both the potential and its forces are directly computable from existing single-shot imaging data, turning quantum exchange into measurable classical mechanics.

cond-mat.stat-mech

Universal Box Operator: $\mathbf{O}(D,D)$-Symmetry and $\alpha^{\prime}$-Corrections

We construct a fully covariant,$\mathbf{O}(D,D)$-symmetric d'Alembertian -- or box operator -- that acts on tensor fields of arbitrary rank and provides a universal kinetic term for all bosonic closed-string states. In its Riemannian parametrization, the operator packages the Riemann curvature, $H$-flux, and dilaton gradient into a single duality-covariant object. This yields $\mathbf{O}(D,D)$-symmetric gravitational-wave equations for the massless sector, governs the tachyon and all massive modes, and clarifies how higher excitations contribute to $\alpha^{\prime}$-corrections. The box operator thus supplies a unified description of closed-string dynamics across the entire spectrum. Our analysis shows that any apparent breaking of $\mathbf{O}(D,D)$ symmetry arises only after integrating out massive modes in a Wilsonian sense, where loop momentum integrals obscure half of the doubled momenta. We stand on the view that $\mathbf{O}(D,D)$ symmetry and doubled diffeomorphisms remain exact and undeformed at the fundamental level of string theory.

hep-th

Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations

Imitation learning (IL) enables agents to mimic expert behavior without reward signals but faces challenges in cross-domain scenarios with high-dimensional, noisy, and incomplete visual observations. To address this, we propose Domain-Invariant Per-Frame Feature Extraction for Imitation Learning (DIFF-IL), a novel IL method that extracts domain-invariant features from individual frames and adapts them into sequences to isolate and replicate expert behaviors. We also introduce a frame-wise time labeling technique to segment expert behaviors by timesteps and assign rewards aligned with temporal contexts, enhancing task performance. Experiments across diverse visual environments demonstrate the effectiveness of DIFF-IL in addressing complex visual tasks.

cs.CV

Improving Perceptual Quality, Intelligibility, and Acoustics on VoIP Platforms

In this paper, we present a method for fine-tuning models trained on the Deep Noise Suppression (DNS) 2020 Challenge to improve their performance on Voice over Internet Protocol (VoIP) applications. Our approach involves adapting the DNS 2020 models to the specific acoustic characteristics of VoIP communications, which includes distortion and artifacts caused by compression, transmission, and platform-specific processing. To this end, we propose a multi-task learning framework for VoIP-DNS that jointly optimizes noise suppression and VoIP-specific acoustics for speech enhancement. We evaluate our approach on a diverse VoIP scenarios and show that it outperforms both industry performance and state-of-the-art methods for speech enhancement on VoIP applications. Our results demonstrate the potential of models trained on DNS-2020 to be improved and tailored to different VoIP platforms using VoIP-DNS, whose findings have important applications in areas such as speech recognition, voice assistants, and telecommunication.

cs.SD

Speech Enhancement for Virtual Meetings on Cellular Networks

We study speech enhancement using deep learning (DL) for virtual meetings on cellular devices, where transmitted speech has background noise and transmission loss that affects speech quality. Since the Deep Noise Suppression (DNS) Challenge dataset does not contain practical disturbance, we collect a transmitted DNS (t-DNS) dataset using Zoom Meetings over T-Mobile network. We select two baseline models: Demucs and FullSubNet. The Demucs is an end-to-end model that takes time-domain inputs and outputs time-domain denoised speech, and the FullSubNet takes time-frequency-domain inputs and outputs the energy ratio of the target speech in the inputs. The goal of this project is to enhance the speech transmitted over the cellular networks using deep learning models.

cs.SD