SearcharxivSearch

arXiv subjects

Jinho Park

Publications and source records attributed to Jinho Park.

At least 19 recordsLinked to original sources

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-temporal reasoning benchmark datasets primarily rely on static image sets or passively curated video data, which limits the evaluation of fine-grained reasoning capabilities. In this paper, we introduce VGenST-Bench, a video benchmark that employs generative models to actively synthesize highly controlled and diverse evaluation scenarios. To construct VGenST-Bench, we propose a multi-agent pipeline incorporating a human quality control stage, ensuring the quality of all generated videos and QA pairs. We establish a comprehensive 3x2x2 video taxonomy, encompassing Spatial Scale, Perspective, and Scene Dynamics to span diverse scenarios. Furthermore, we design a hierarchical task suite that decouples low-level visual perception from high-level spatio-temporal reasoning. By shifting the paradigm from passive curation to active synthesis, VGenST-Bench enables fine-grained diagnosis of spatio-temporal understanding in MLLMs.

cs.CV

EXAONE 4.5 Technical Report

This technical report introduces EXAONE 4.5, the first open-weight vision language model released by LG AI Research. EXAONE 4.5 is architected by integrating a dedicated visual encoder into the existing EXAONE 4.0 framework, enabling native multimodal pretraining over both visual and textual modalities. The model is trained on large-scale data with careful curation, particularly emphasizing document-centric corpora that align with LG's strategic application domains. This targeted data design enables substantial performance gains in document understanding and related tasks, while also delivering broad improvements across general language capabilities. EXAONE 4.5 extends context length up to 256K tokens, facilitating long-context reasoning and enterprise-scale use cases. Comparative evaluations demonstrate that EXAONE 4.5 achieves competitive performance in general benchmarks while outperforming state-of-the-art models of similar scale in document understanding and Korean contextual reasoning. As part of LG's ongoing effort toward practical industrial deployment, EXAONE 4.5 is designed to be continuously extended with additional domains and application scenarios to advance AI for a better life.

cs.CL

Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection

Open-vocabulary 3D object detection aims to localize and recognize objects beyond a fixed training taxonomy. In multi-view RGB settings, recent approaches often decouple geometry-based instance construction from semantic labeling, generating class-agnostic fragments and assigning open-vocabulary categories post hoc. While flexible, such decoupling leaves instance construction governed primarily by geometric consistency, without semantic constraints during merging. When geometric evidence is view-dependent and incomplete, this geometry-only merging can lead to irreversible association errors, including over-merging of distinct objects or fragmentation of a single instance. We propose Group3D, a multi-view open-vocabulary 3D detection framework that integrates semantic constraints directly into the instance construction process. Group3D maintains a scene-adaptive vocabulary derived from a multimodal large language model (MLLM) and organizes it into semantic compatibility groups that encode plausible cross-view category equivalence. These groups act as merge-time constraints: 3D fragments are associated only when they satisfy both semantic compatibility and geometric consistency. This semantically gated merging mitigates geometry-driven over-merging while absorbing multi-view category variability. Group3D supports both pose-known and pose-free settings, relying only on RGB observations. Experiments on ScanNet and ARKitScenes demonstrate that Group3D achieves state-of-the-art performance in multi-view open-vocabulary 3D detection, while exhibiting strong generalization in zero-shot scenarios. The project page is available at https://ubin108.github.io/Group3D/.

cs.CV

AdaRadar: Rate Adaptive Spectral Compression for Radar-based Perception

Radar is a critical perception modality in autonomous driving systems due to its all-weather characteristics and ability to measure range and Doppler velocity. However, the sheer volume of high-dimensional raw radar data saturates the communication link to the computing engine (e.g., an NPU), which is often a low-bandwidth interface with data rate provisioned only for a few low-resolution range-Doppler frames. A generalized codec for utilizing high-dimensional radar data is notably absent, while existing image-domain approaches are unsuitable, as they typically operate at fixed compression ratios and fail to adapt to varying or adversarial conditions. In light of this, we propose radar data compression with adaptive feedback. It dynamically adjusts the compression ratio by performing gradient descent from the proxy gradient of detection confidence with respect to the compression rate. We employ a zeroth-order gradient approximation as it enables gradient computation even with non-differentiable core operations--pruning and quantization. This also avoids transmitting the gradient tensors over the band-limited link, which, if estimated, would be as large as the original radar data. In addition, we have found that radar feature maps are heavily concentrated on a few frequency components. Thus, we apply the discrete cosine transform to the radar data cubes and selectively prune out the coefficients effectively. We preserve the dynamic range of each radar patch through scaled quantization. Combining those techniques, our proposed online adaptive compression scheme achieves over 100x feature size reduction at minimal performance drop (~1%p). We validate our results on the RADIal, CARRADA, and Radatron datasets.

cs.CV

K-EXAONE Technical Report

This technical report presents K-EXAONE, a large-scale multilingual language model developed by LG AI Research. K-EXAONE is built on a Mixture-of-Experts architecture with 236B total parameters, activating 23B parameters during inference. It supports a 256K-token context window and covers six languages: Korean, English, Spanish, German, Japanese, and Vietnamese. We evaluate K-EXAONE on a comprehensive benchmark suite spanning reasoning, agentic, general, Korean, and multilingual abilities. Across these evaluations, K-EXAONE demonstrates performance comparable to open-weight models of similar size. K-EXAONE, designed to advance AI for a better life, is positioned as a powerful proprietary AI foundation model for a wide range of industrial and research applications.

cs.CL

Coherent and compact van der Waals transmon qubits

State-of-the-art superconducting qubits rely on a limited set of thin-film materials. Expanding their materials palette can improve performance, extend operating regimes, and introduce new functionalities, but conventional thin-film fabrication hinders systematic exploration of new material combinations. Van der Waals (vdW) materials offer a highly modular crystalline platform that facilitates such exploration while enabling gate-tunability, higher-temperature operation, and compact qubit geometries. Yet it remains unknown whether a fully vdW superconducting qubit can support quantum coherence and what mechanisms dominate loss at both low and elevated temperatures in such a device. Here we demonstrate quantum-coherent merged-element transmons made entirely from vdW Josephson junctions. These first-generation, fully crystalline qubits achieve microsecond lifetimes in an ultra-compact footprint without external shunt capacitors. Energy relaxation measurements, together with microwave characterization of vdW capacitors, point to dielectric loss as the dominant relaxation channel up to hundreds of millikelvin. These results establish vdW materials as a viable platform for compact superconducting quantum devices.

quant-ph

Measuring Reactive-Load Impedance with Transmission-Line Resonators Beyond the Perturbative Limit

We develop an analytic framework to extract circuit parameters and loss tangent from superconducting transmission-line resonators terminated by reactive loads, extending analysis beyond the perturbative regime. The formulation yields closed-form relations between resonant frequency, participation ratio, and internal quality factor, removing the need for full-wave simulations. We validate the framework through circuit simulations, finite-element modeling, and experimental measurements of van der Waals parallel-plate capacitors, using it to extract the dielectric constant and loss tangent of hexagonal boron nitride. Statistical analysis across multiple reference resonators, together with multimode self-calibration, demonstrates consistent and reproducible extraction of both capacitance and loss tangent in close agreement with literature values. In addition to parameter extraction, the analytic relations provide practical design guidelines for maximizing energy participation ratio in the load and improving the precision of resonator-based material metrology.

quant-ph

Crystalline superconductor-semiconductor Josephson junctions for compact superconducting qubits

The narrow bandgap of semiconductors allows for thick, uniform Josephson junction barriers, potentially enabling reproducible, stable, and compact superconducting qubits. We study vertically stacked van der Waals Josephson junctions with semiconducting weak links, whose crystalline structures and clean interfaces offer a promising platform for quantum devices. We observe robust Josephson coupling across 2--12 nm (3--18 atomic layers) of semiconducting WSe$_2$ and, notably, a crossover from proximity- to tunneling-type behavior with increasing weak link thickness. Building on these results, we fabricate a prototype all-crystalline merged-element transmon qubit with transmon frequency and anharmonicity closely matching design parameters. We demonstrate dispersive coupling between this transmon and a microwave resonator, highlighting the potential of crystalline superconductor-semiconductor structures for compact, tailored superconducting quantum devices.

cond-mat.mes-hall

Characterizing Pattern Matching and Its Limits on Compositional Task Structures

Despite impressive capabilities, LLMs' successes often rely on pattern-matching behaviors, yet these are also linked to OOD generalization failures in compositional tasks. However, behavioral studies commonly employ task setups that allow multiple generalization sources (e.g., algebraic invariances, structural repetition), obscuring a precise and testable account of how well LLMs perform generalization through pattern matching and their limitations. To address this ambiguity, we first formalize pattern matching as functional equivalence, i.e., identifying pairs of subsequences of inputs that consistently lead to identical results when the rest of the input is held constant. Then, we systematically study how decoder-only Transformer and Mamba behave in controlled tasks with compositional structures that isolate this mechanism. Our formalism yields predictive and quantitative insights: (1) Instance-wise success of pattern matching is well predicted by the number of contexts witnessing the relevant functional equivalence. (2) We prove a tight sample complexity bound of learning a two-hop structure by identifying the exponent of the data scaling law for perfect in-domain generalization. Our empirical results align with the theoretical prediction, under 20x parameter scaling and across architectures. (3) Path ambiguity is a structural barrier: when a variable influences the output via multiple paths, models fail to form unified intermediate state representations, impairing accuracy and interpretability. (4) Chain-of-Thought reduces data requirements yet does not resolve path ambiguity. Hence, we provide a predictive, falsifiable boundary for pattern matching and a foundational diagnostic for disentangling mixed generalization mechanisms.

cs.LG

The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think

Long chain-of-thought (CoT) is an essential ingredient in effective usage of modern large language models, but our understanding of the reasoning strategies underlying these capabilities remains limited. While some prior works have attempted to categorize CoTs using predefined strategy types, such approaches are constrained by human intuition and fail to capture the full diversity of model behaviors. In this work, we introduce the CoT Encyclopedia, a bottom-up framework for analyzing and steering model reasoning. Our method automatically extracts diverse reasoning criteria from model-generated CoTs, embeds them into a semantic space, clusters them into representative categories, and derives contrastive rubrics to interpret reasoning behavior. Human evaluations show that this framework produces more interpretable and comprehensive analyses than existing methods. Moreover, we demonstrate that this understanding enables performance gains: we can predict which strategy a model is likely to use and guide it toward more effective alternatives. Finally, we provide practical insights, such as that training data format (e.g., free-form vs. multiple-choice) has a far greater impact on reasoning behavior than data domain, underscoring the importance of format-aware model design.

cs.CL

Engineering Andreev Bound States for Thermal Sensing in Proximity Josephson Junctions

The thermal response of proximity Josephson junctions (JJs) is governed by the temperature ($T$)-dependent occupation of Andreev bound states (ABS), making them promising candidates for sensitive thermal detection. In this study, we systematically engineer ABS to enhance the thermal sensitivity of the critical current ($I_c$) of proximity JJs, quantified as $|\,dI_c/dT\,|$ for the threshold readout scheme and $|\,dI_c/dT \cdot I_c^{-1}\,|$ for the inductive readout scheme. Using a gate-tunable graphene-based JJ platform, we explore the impact of key parameters -- including channel length, transparency, carrier density, and superconducting material -- on the thermal response. Our results reveal that the proximity-induced superconducting gap plays a crucial role in optimizing thermal sensitivity. Notably, we see a maximum $|\,dI_c/dT \cdot I_c^{-1}\,|$ value of $0.6\,\mathrm{K}^{-1}$ at low temperatures with titanium-based graphene JJs. By demonstrating a systematic approach to engineering ABS in proximity JJs, this work establishes a versatile framework for optimizing thermal sensors and advancing the study of ABS-mediated transport.

cond-mat.supr-con

Color Universal Design Neural Network for the Color Vision Deficiencies

Information regarding images should be visually understood by anyone, including those with color deficiency. However, such information is not recognizable if the color that seems to be distorted to the color deficiencies meets an adjacent object. The aim of this paper is to propose a color universal design network, called CUD-Net, that generates images that are visually understandable by individuals with color deficiency. CUD-Net is a convolutional deep neural network that can preserve color and distinguish colors for input images by regressing the node point of a piecewise linear function and using a specific filter for each image. To generate CUD images for color deficiencies, we follow a four-step process. First, we refine the CUD dataset based on specific criteria by color experts. Second, we expand the input image information through pre-processing that is specialized for color deficiency vision. Third, we employ a multi-modality fusion architecture to combine features and process the expanded images. Finally, we propose a conjugate loss function based on the composition of the predicted image through the model to address one-to-many problems that arise from the dataset. Our approach is able to produce high-quality CUD images that maintain color and contrast stability. The code for CUD-Net is available on the GitHub repository

eess.IV

Unveiling Topological Hinge States in the Higher-Order Topological Insulator WTe$_2$ Based on the Fractional Josephson Effect

Higher-order topological insulators (HOTIs) represent a novel class of topological materials, characterised by the emergence of topological boundary modes at dimensions two or more lower than those of bulk materials. Recent experimental studies have identified conducting channels at the hinges of HOTIs, although their topological nature remains unexplored. In this study, we investigated Shapiro steps in Al-WTe$_2$-Al proximity Josephson junctions (JJs) under microwave irradiation and examined the topological properties of the hinge states in WTe$_2$. Specifically, we analysed the microwave frequency dependence of the absence of the first Shapiro step in hinge-dominated JJs, attributing this phenomenon to the 4$\pi$-periodic current-phase relationship characteristic of topological JJs. These findings may encourage further research into topological superconductivity with topological hinge states in superconducting hybrid devices based on HOTIs. Such advances could lead to the realisation of Majorana zero modes for topological quantum physics and pave the way for applications in spintronic devices.

cond-mat.mes-hall

How language models extrapolate outside the training data: A case study in Textualized Gridworld

Language models' ability to extrapolate learned behaviors to novel, more complex environments beyond their training scope is highly unknown. This study introduces a path planning task in a textualized Gridworld to probe language models' extrapolation capabilities. We show that conventional approaches, including next token prediction and Chain of Thought (CoT) finetuning, fail to extrapolate in larger, unseen environments. Inspired by human cognition and dual process theory, we propose cognitive maps for path planning, a novel CoT framework that simulates humanlike mental representations. Our experiments show that cognitive maps not only enhance extrapolation to unseen environments but also exhibit humanlike characteristics through structured mental simulation and rapid adaptation. Our finding that these cognitive maps require specialized training schemes and cannot be induced through simple prompting opens up important questions about developing general-purpose cognitive maps in language models. Our comparison with exploration-based methods further illuminates the complementary strengths of offline planning and online exploration.

cs.CL

How Do Large Language Models Acquire Factual Knowledge During Pretraining?

Despite the recent observation that large language models (LLMs) can store substantial factual knowledge, there is a limited understanding of the mechanisms of how they acquire factual knowledge through pretraining. This work addresses this gap by studying how LLMs acquire factual knowledge during pretraining. The findings reveal several important insights into the dynamics of factual knowledge acquisition during pretraining. First, counterintuitively, we observe that pretraining on more data shows no significant improvement in the model's capability to acquire and maintain factual knowledge. Next, there is a power-law relationship between training steps and forgetting of memorization and generalization of factual knowledge, and LLMs trained with duplicated training data exhibit faster forgetting. Third, training LLMs with larger batch sizes can enhance the models' robustness to forgetting. Overall, our observations suggest that factual knowledge acquisition in LLM pretraining occurs by progressively increasing the probability of factual knowledge presented in the pretraining data at each step. However, this increase is diluted by subsequent forgetting. Based on this interpretation, we demonstrate that we can provide plausible explanations for recently observed behaviors of LLMs, such as the poor performance of LLMs on long-tail knowledge and the benefits of deduplicating the pretraining corpus.

cs.CL

UL-VIO: Ultra-lightweight Visual-Inertial Odometry with Noise Robust Test-time Adaptation

Data-driven visual-inertial odometry (VIO) has received highlights for its performance since VIOs are a crucial compartment in autonomous robots. However, their deployment on resource-constrained devices is non-trivial since large network parameters should be accommodated in the device memory. Furthermore, these networks may risk failure post-deployment due to environmental distribution shifts at test time. In light of this, we propose UL-VIO -- an ultra-lightweight (<1M) VIO network capable of test-time adaptation (TTA) based on visual-inertial consistency. Specifically, we perform model compression to the network while preserving the low-level encoder part, including all BatchNorm parameters for resource-efficient test-time adaptation. It achieves 36X smaller network size than state-of-the-art with a minute increase in error -- 1% on the KITTI dataset. For test-time adaptation, we propose to use the inertia-referred network outputs as pseudo labels and update the BatchNorm parameter for lightweight yet effective adaptation. To the best of our knowledge, this is the first work to perform noise-robust TTA on VIO. Experimental results on the KITTI, EuRoC, and Marulan datasets demonstrate the effectiveness of our resource-efficient adaptation method under diverse TTA scenarios with dynamic domain shifts.

cs.CV

Twisted van der Waals Josephson junction based on high-Tc superconductor

Stacking two-dimensional van der Waals (vdW) materials rotated with respect to each other show versatility for the study of exotic quantum phenomena. Especially, anisotropic layered materials have great potential for such twistronics applications, providing high tunability. Here, we report anisotropic superconducting order parameters in twisted Bi2Sr2CaCu2O8+x (Bi-2212) vdW junctions with an atomically clean vdW interface, achieved using the microcleave-and-stack technique. The vdW Josephson junctions with twist angles of 0° and 90° showed the maximum Josephson coupling, which was comparable to that of intrinsic Josephson (IJ) junctions in the bulk crystal. As the twist angle approaches 45°, Josephson coupling is suppressed, and eventually disappears at 45°. The observed twist angle dependence of the Josephson coupling can be explained quantitatively by theoretical calculation with the d-wave superconducting order parameter of Bi-2212 and finite tunneling incoherence of the junction. Our results revealed the anisotropic nature of Bi-2212 and provided a novel fabrication technique for vdW-based twistronics platforms compatible with air-sensitive vdW materials.

cond-mat.supr-con

Steady Floquet-Andreev States Probed by Tunnelling Spectroscopy

Engineering quantum states through light-matter interaction has created a new paradigm in condensed matter physics. A representative example is the Floquet-Bloch state, which is generated by time-periodically driving the Bloch wavefunctions in crystals. Previous attempts to realise such states in condensed matter systems have been limited by the transient nature of the Floquet states produced by optical pulses, which masks the universal properties of non-equilibrium physics. Here, we report the generation of steady Floquet Andreev (F-A) states in graphene Josephson junctions by continuous microwave application and direct measurement of their spectra by superconducting tunnelling spectroscopy. We present quantitative analysis of the spectral characteristics of the F-A states while varying the phase difference of superconductors, temperature, microwave frequency and power. The oscillations of the F-A state spectrum with phase difference agreed with our theoretical calculations. Moreover, we confirmed the steady nature of the F-A states by establishing a sum rule of tunnelling conductance, and analysed the spectral density of Floquet states depending on Floquet interaction strength. This study provides a basis for understanding and engineering non-equilibrium quantum states in nano-devices.

cond-mat.mes-hall