SearcharxivSearch

arXiv subjects

Kyungmin Lee

Publications and source records attributed to Kyungmin Lee.

At least 19 recordsLinked to original sources

A rigorous hybridization of variational quantum eigensolver with classical neural network

Combining variational quantum process with classical neural learning offers a flexible route to improve ground-state estimation. We establish an integrated framework that connects neural transformations, measurement statistics, and guarantees physical stability through three requirements: self-contained training, polynomial resource scaling, and variational consistency. Its constructive realization, termed \emph{unitary variational quantum-neural hybrid eigensolver}~(U-VQNHE), couples a variational quantum circuit to a neural phase function through norm-preserving post-processing. The learned transformation is evaluated from measurement records, preserves the exact variational bound, and admits range-independent concentration guarantees for independent finite-shot evaluations. Complementing this construction, we characterize the statistical and representational constraints of amplitude reweighting: sampled-support mismatch can destabilize empirical normalization, while exact distribution matching can require exponentially large dynamic range for Haar-random targets and structured ansatz--target pairs under specified near-tensorizability and mismatch conditions. Finite-shot simulations on Ising and disordered XYZ spin models demonstrate improved energy accuracy over the underlying variational quantum eigensolver and greater robustness than the amplitude-reweighting baselines. Together, these results provide a principled foundation for quantum--neural eigensolvers in which physical consistency, expressive capacity, and measurement cost are treated as a single design problem.

quant-ph

The $ν$EYE Neutrino Telescope: Conceptual Design Report

The $\bfνEYE$ neutrino project leverages the existing large pit at Yemilab located in South Korea, to reveal the existence of sterile neutrino, the up-turn of the neutrinos from the Sun, and the first minimum of the neutrino oscillation over distances on the order of tens of kilometers for the first time. This initiative is expected to facilitate a wide range of significant scientific and technological advancements within both South Korean and international communities engaged in neutrino science and technology. The $\bfνEYE$ aims to investigate the largely unexplored sector of almost-massless lepton in the elementary particle physics in detail. The emphasis will be placed on the study of real time nuclear processes and reactions involving possible sterile neutrinos on timescales down to nanoseconds in ultra-high intense or radioactive neutrino beams for the first time in the world; the $\bfνEYE$ looks at to-be universal oscillation (``up-turn'' in the electron neutrino survival probability) of neutrinos predicted by the three neutrino oscillation paradigm. This will confirm or deny our current understanding on the particle interactions of the lepton sector; and measurement of the first oscillation minimum between the first and second neutrinos in mass.

hep-ex

Energy Transport Velocity in Photonic Time Crystals

Steep or near-vertical Floquet dispersion in photonic time crystals (PTCs) is often read as fast, even apparently superluminal, transport. Here, we demonstrate that this anomaly arises from modulation-driven geometric drift, not energy flow. By deriving a Maxwell-flux Hellmann-Feynman relation, we prove that the cycle-averaged energy velocity remains strictly bounded. We further establish a universal velocity-product law conserved throughout the passband, $ v_E v_g=\langle v_{\rm ph}^2\rangle_T $, fixing transport solely by the temporal average of the inverse permittivity. The divergent group velocity is then traced to a mismatch between electric and magnetic geometric phase connections, revealing apparent superluminality as a geometric effect of temporal modulation.

physics.optics

K-EXAONE 2.0 Technical Report

This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundation models. Rather than training from scratch, we upcycle K-EXAONE and expand its architecture, yielding a Mixture-of-Experts (MoE) model with 750B total parameters and approximately 37B activated per token---more than three times the capacity of its predecessor. K-EXAONE 2.0 supports context lengths of up to 256K tokens and expands multilingual coverage from six to ten languages. Its training pipeline combines continual pre-training, difficulty-focused mid-training, and post-training to strengthen reasoning, agentic coding, multilingual capability, and safety grounded in Korean sociocultural contexts. Across nine evaluation categories selected to reflect the conditions of practical use, K-EXAONE 2.0 improves over K-EXAONE and remains competitive with open-weight models, showing its largest gains in agentic coding and long-context understanding and its clearest strengths in long-context retrieval and safety. Released under the Apache 2.0 license, K-EXAONE 2.0 enables the wider AI ecosystem to evaluate, deploy, adapt, and build upon it, while marking the beginning---rather than the endpoint---of our challenge toward the global frontier.

cs.CL

Classical Petermann Factor as a Measure of Quantum Squeezing in Photonic Time Crystals

Photonic time crystals realize a continuum of momentum-resolved SU(1,1) parametric amplifiers. We show that a classical quantity, the Petermann factor of the effective Floquet Bogoliubov de Gennes (BdG) dynamical matrix, sets the scale of their quantum noise. In stable bands it fixes the Bogoliubov mixing and hence the mean bare-photon occupation of the Floquet vacuum, while in momentum gaps it sets the photon-number prefactor and enhances the squeezing dynamics, with the Floquet growth rate setting the time scale. This converts classical measurements of mode nonorthogonality into quantitative predictions for squeezing and photon generation, and offers a compact design parameter for engineering quantum resources in two-mode BdG platforms.

physics.optics

Multimode Phonon-Number Measurement and Single-shot Superparity Measurement using Dispersive Shifts in a Trapped Ion

Dispersive shifts are a widely used tool for bosonic readout and control in circuit quantum electrodynamics, yet they remain relatively unexplored in trapped-ion motional systems. Here we introduce a unified framework for multimode phonon-number measurement and nondestructive single-shot superparity measurement, i.e., phonon-number measurement modulo 2^k, using dispersive shifts in the far-detuned multimode Jaynes-Cummings interaction of a trapped ion system. We implement a Ramsey sequence that realizes a multimode spin-dependent rotation (SDR) together with a selective decoupling scheme that cancels the phase induced by the carrier AC-Stark shift while preserving the phonon-number-dependent phase induced by the dispersive shift. Within this framework, we infer single-mode and two-mode Fock-state distributions from spin-population dynamics, use SDR-based conditional parity operators with postselection to generate cat states and entangled coherent states, and realize nondestructive single-shot measurements of phonon number modulo 2, 4, and 8 in the single-mode setting. These results open a new avenue for the use of multimode parity operators in trapped-ion bosonic systems.

quant-ph

SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creating an urgent need for reliable deepfake detection. Yet most existing approaches rely on massive real--fake datasets, which are increasingly difficult to maintain as new generators continue to emerge. In this work, we investigate how much information about image authenticity is already encoded in modern multimodal vision representations. We find that frozen multimodal encoders naturally separate real and synthetic images in their embedding space, enabling a simple linear classifier to achieve strong performance without task-specific fine-tuning. Motivated by this observation, we develop a representation-aware data curation strategy that selects a compact set of representative generators for training. The resulting training set contains only 10K images, compared to 288K in AIGIBench and 4M in OpenFake, while improving robustness to unseen generators and distribution shifts. We additionally introduce RealWorldBench, a benchmark consisting of modern camera photographs, contemporary stock images, and outputs from recent commercial generators. Experiments across multiple benchmarks show that combining frozen multimodal representations with carefully curated training data provides a simple and effective approach to AI-generated image detection.

cs.CV

PHUMA: Physically Reliable Humanoid Locomotion Dataset

Motion imitation is a promising approach for humanoid locomotion, enabling agents to acquire humanlike behaviors. Existing methods typically rely on high-quality motion capture datasets such as AMASS, but these are scarce and expensive, limiting scalability and diversity. Recent studies attempt to scale data collection by converting large-scale internet videos, exemplified by Humanoid-X. However, they often suffer from physical artifacts such as floating, penetration, and foot skating, which hinder stable imitation. To address this, we introduce PHUMA, a Physically Reliable HUMAnoid locomotion dataset produced by a two-stage pipeline combining physics-aware curation and physics-constrained retargeting, aggregating both motion capture and internet video into a physically reliable, 73-hour corpus. On motion tracking benchmarks, PHUMA-trained policies achieve higher success rates than those trained on AMASS and Humanoid-X, and successfully transfer zero-shot to a real Unitree G1. The code is available at https://davian-robotics.github.io/PHUMA.

cs.RO

Quantum Photonic Time Crystals: From Temporal Boundaries to Floquet Light-Matter Interactions

Photonic time crystals (PTCs) are temporally periodic media whose Floquet spectra can exhibit momentum gaps, parametric amplification, and effective non-Hermitian descriptions, making them an idealized setting for vacuum amplification and nonequilibrium light-matter dynamics. Their classical electrodynamics is now well developed; the quantum side is less so, and this focused review is an attempt to organize what exists. We trace that account from temporal boundaries to homogeneous Floquet media and light-matter dynamics. A single temporal boundary induces Bogoliubov mode mixing and photon-pair creation; in homogeneous bulk media, momentum conservation isolates counter-propagating $(k,-k)$ sectors and yields a two-mode $SU(1,1)$ squeezing structure. Temporal periodicity promotes this to a Floquet problem with band and momentum-gap regimes, compactly described in a fixed Nambu basis. We then relate PTCs to the dynamical Casimir effect and parametric amplification, which share the same pair-creation mechanism but organize it through discrete resonances rather than a momentum-resolved bulk spectrum. We close with light-matter settings: spontaneous-emission decay and modulation-assisted excitation, atom-PTC dynamics, LDOS-based observables and their limits, and finite, dispersive, and experimentally accessible platforms.

physics.optics

Electrolyte Bonding Engineering for Highly Uniform GeTe-based CBRAM and Parallel Hebbian Learning in Selector-free Hopfield Networks

Hopfield networks offer a hardware-friendly framework for energy-efficient associative memory, yet their practical realization in memristor crossbar arrays is critically hindered by device-to-device (D2D) variability, which prevents reliable parallel programming. Here, we address this bottleneck through systematic composition engineering of the Ge-Te solid electrolyte in conductive bridge random access memory (CBRAM) devices. By varying the Ge:Te ratio, we identify Ge3.5Te1 as an optimal electrolyte composition that suppresses stochastic resistance variation by approximately three orders of magnitude compared to GeSe-based devices. Raman spectroscopy reveals that this dramatic improvement originates from a bonding network dominated by asymmetric-stretching GeTe4 tetrahedral units, which form interconnected free-volume channels that confine and stabilize Cu+ ion migration pathways. Leveraging this enhanced uniformity, we fabricate a selector-less 16x16 Cu/Ge3.5Te1 CBRAM crossbar array and demonstrate a 4x4 Hopfield associative network capable of learning and recalling binary pattern pairs via fully parallel programming using a half-selection scheme. Successful pattern recall is achieved for up to two stored associations despite the absence of selector elements, establishing a proof-of-concept for selector-free hardware implementations of associative memory. These results highlight the critical role of electrolyte bonding structure in determining memristor uniformity and provide a materials-driven pathway toward scalable, parallel neuromorphic computing systems.

physics.app-ph

Contrastive Representation Regularization for Vision-Language-Action Models

Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained Vision-Language Models (VLMs). However, their representations arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and proprioceptive information. To address the issue, we introduce Robot State-aware Contrastive Loss (RS-CL), a simple and effective representation regularization for VLA models, designed to bridge the gap between VLM representations and robotic signals. In particular, RS-CL aligns the representations more closely with the robot's proprioceptive states by using relative distances between the states as soft supervision. Complementing the original action prediction objective, RS-CL enhances control-relevant representation learning, while being lightweight and fully compatible with standard VLA training pipelines. Our empirical results demonstrate that RS-CL substantially improves the performance of state-of-the-art VLA models; it pushes the prior art to 69.7% achieving the state-of-the-art performance on the RoboCasa-Kitchen benchmark, and boosts success rates from 45.0% to 58.3% on challenging real-robot manipulation tasks.

cs.RO

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring a multimodal diffusion transformer that maintains separate modality streams while enabling cross-modal knowledge sharing. In addition, DUST utilizes independent noise perturbations and a decoupled flow matching loss to learn cross-modal causal relationships. We further introduce an asynchronous sampling method for action and vision tokens that enhances performance through inference-time scaling. Experimental results on simulated benchmarks like RoboCasa and GR-1 show that DUST achieves up to 6% gains over state-of-the-art VLA and world-modeling baselines, with inference-time scaling providing an additional 2-5% improvement. In real-world tasks using the Franka Research 3, DUST outperforms baselines by 10% in success rate. Finally, we demonstrate that DUST enables effective transfer learning through both pretraining on action-free videos and joint-training with heterogeneous robot and human datasets.

cs.CV

Development of a system for testing full-size CMS LGAD sensors

Low-Gain Avalanche Diode (LGAD) sensors, offering timing resolutions of the order of tens of picoseconds, are being widely adopted in particle physics experiments and related applications. As these applications scale to large numbers of sensors with varying pixel geometries, conventional manual characterization techniques become inadequate for large-scale quality control. We present a modular probe card system for automated electrical characterization of pixelated LGAD sensors, consisting of a probe card, a switching board, precision measurement instruments, and control software. The system supports flexible pixel selection and measurement. Its performance is demonstrated through current-voltage (I-V) and capacitance-voltage (C-V) measurements of a $16 \times 16$ LGAD array. A rapid row-wise I-V scan of the full array is completed in approximately 20 minutes, while a pixel-by-pixel I-V scan from 0 to 300 V with a 1 V step requires about 340 minutes. The switching matrix introduces less than 1 nA of leakage current even in a conservative worst-case configuration, remaining small compared with the leakage current of a normal LGAD pixel. The modular architecture and automation capability make the system a practical and scalable solution for large-scale LGAD sensor quality control and distributed testing environments.

physics.ins-det

Trust Region Q Adjoint Matching

Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-learning with Adjoint Matching (QAM) addressed this issue by reformulating into a memoryless stochastic optimal control (SOC) problem with a learned critic. However, QAM inherits a fundamental fragility of critic-guided improvement: small critic errors are amplified when critics are ill-conditioned, often leading to model collapse. This paper introduces Trust Region Q-Adjoint Matching (TRQAM), a stable off-policy fine-tuning algorithm that adaptively controls the path-space KL with pretrained flow policies through projected dual descent. Specifically, we optimize the trust-region parameter $λ$ in SOC dynamics, and theoretically show that the path-space KL can be represented by a closed-form function of $λ$. As a result, our method can precisely control the exact deviation from pretrained flow policies, achieving stable off-policy RL. Through experiments on 50 OGBench tasks, TRQAM consistently outperforms prior arts in both offline RL and offline-to-online RL. In particular, TRQAM achieves an overall success rate of 68% in offline RL, substantially improves the strongest baseline at 46%.

cs.LG

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure step-by-step reasoning alone, research-level problems use such reasoning to advance the frontier of mathematical knowledge itself, emerging as a compelling alternative. Yet research-level math benchmarks remain scarce because such problems are difficult to source (e.g., Riemann Bench and FrontierMath-Tier 4 contain 25 and 50 problems, respectively). To support reliable evaluation of next-generation frontier models, we introduce Soohak, a 439-problem benchmark newly authored from scratch by 64 mathematicians. Soohak comprises two subsets. On the Challenge subset, frontier models including Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach 30.4%, 26.4%, and 10.4% respectively, leaving substantial headroom, while leading open-weight models such as Qwen3-235B, GPT-OSS-120B, and Kimi-2.5 remain below 15%. Notably, beyond standard problem solving, Soohak introduces a refusal subset that probes a capability intrinsic to research mathematics: recognizing ill-posed problems and pausing rather than producing confident but unjustified answers. On this subset, no model exceeds 50%, identifying refusal as a new optimization target that current models do not directly address. To prevent contamination, the dataset will be publicly released in late 2026, with model evaluations available upon request in the interim.

cs.CL

Self-organized photonic time quasicrystal from a single imposed clock

A photonic time crystal usually writes a clock into a medium. Here one clock does more than program the medium: it seeds a quasiperiodic temporal order that the nonlinear medium selects for itself. In a guided-wave lattice of nonlinear dipoles, a single-tone pump modulates the polarization sector, while Maxwell--polarization back-action selects two response frequencies whose only resolved low-order relation is the pump-locked sum condition. Their sum phase locks to the pump and the complementary phase winds, producing a photonic discrete time quasicrystal with torus-like phase dynamics and a discrete combination spectrum. Site-resolved measurements show locked-phase coherence across the measured lattice sites over a finite control-parameter window. These results establish a route from externally programmed time-varying media to self-organized temporal order in nonlinear photonic systems.

physics.optics

A Hardware-aware Hopfield Network with a Nonlinear Memristor Array for Robust Associative Memory with Superlinear Capacity

Associative memory retrieves complete patterns from partial or corrupted inputs and constitutes a primitive form of generative inference. Classical Hopfield networks (CHN) provide a canonical framework for associative memory but suffer from limited memory capacity. Recently, modern Hopfield networks (MHN) were introduced to achieve higher capacity by using explicit pattern-wise storage and neurons with the softmax activation function, which makes the MHN vulnerable to noise and the hardware implementation complicated due to its network size varying with the number of stored patterns. Here, we introduce a hardware-aware Hopfield network (HHN), in which the intrinsic nonlinear current-voltage characteristics of a charge-trap memristor are leveraged to engineer the energy landscape of the HN, increasing the memory capacity. Using a 25 x 25 nonlinear memristor array, we demonstrate reliable reconstruction of corrupted patterns with memory capacity far exceeding the classical limit (K ~ 0.14N, where N is the number of neurons). The HHN preserves Hopfield-type energy-minimization dynamics and remains robust to synaptic conductance noise. Large-scale simulations on high-dimensional image data reveal an empirical memory capacity scaling of K ~ 0.3 x N^1.2 under a fixed synaptic budget. These results establish HHN as a scalable hardware-native architecture for low-power associative memory and generative inference.

cond-mat.dis-nn

RLDX-1 Technical Report

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs (e.g. $π_{0.5}$ and GR00T N1.6) across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8% while $π_{0.5}$ and GR00T N1.6 achieve around 40%, highlighting the ability of RLDX-1 to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

cs.RO