SearcharxivSearch

arXiv subjects

Jimin Lee

Publications and source records attributed to Jimin Lee.

At least 19 recordsLinked to original sources

Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services

Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments such as call center conversations. When privacy restrictions preclude audio access, error correction must rely on text-based post-editing. Existing text-only approaches face significant challenges in low-resource languages, mainly due to a critical scarcity of annotated corpora and tailored correction methodologies. For Korean, this resource gap is particularly pronounced, as existing resources are predominantly designed for ASR training rather than text-based error correction. To address this, we introduce DasanCallDial, the first large-scale Korean benchmark dataset specifically curated for dialogue-level ASR error correction. Derived from genuine call center interactions, it comprises 1,974 dialogues with 115,460 utterances. Leveraging this resource, we propose Detector-Gated Contextual Span Correction (DCSC), a text-only post-editing framework for error-sparse Korean speech recognition transcripts. DCSC combines an encoder-based detector that first performs token-level error detection, followed by a language model-based corrector trained to rectify fine-grained span-level errors. Additionally, we employ dialogue-level context augmentation to enable the model to leverage discourse history for disambiguation. By employing multi-level granularity, our method achieves state-of-the-art performance, effectively overcoming the limitations of general LLMs in low-resource settings.

cs.CL

RLDX-1 Technical Report

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs (e.g. $\pi_{0.5}$ and GR00T N1.6) across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8% while $\pi_{0.5}$ and GR00T N1.6 achieve around 40%, highlighting the ability of RLDX-1 to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

cs.RO

Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models

Humans understand and interact with the real world by relying on diverse physical feedback beyond visual perception. Motivated by this, recent approaches attempt to incorporate physical sensory signals into Vision-Language-Action models (VLAs). However, they typically focus on a single type of physical signal, failing to capture the heterogeneous and complementary nature of real-world interactions. In this paper, we propose MoSS, a modular sensory stream framework that adapts VLAs to leverage multiple sensory signals for action prediction. Specifically, we introduce decoupled modality streams that integrate heterogeneous physical signals into the action stream via joint cross-modal self-attention. To enable stable incorporation of new modalities, we adopt a two-stage training scheme that freezes pretrained VLA parameters in the early stage. Furthermore, to better capture contact interaction dynamics, we incorporate an auxiliary task that predicts future physical signals. Through extensive real-world experiments, we demonstrate that MoSS successfully augments VLAs to leverage diverse physical signals (i.e., tactile and torque), integrating multiple signals to achieve synergistic performance gains.

cs.RO

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

Dense Video Captioning (DVC) is a challenging multimodal task that involves temporally localizing multiple events within a video and describing them with natural language. While query-based frameworks enable the simultaneous, end-to-end processing of localization and captioning, their reliance on shared queries often leads to significant multi-task interference between the two tasks, as well as temporal redundancy in localization. In this paper, we propose utilizing role-specific queries that separate localization and captioning into independent components, allowing each to exclusively learn its role. We then employ contrastive alignment to enforce semantic consistency between the corresponding outputs, ensuring coherent behavior across the separated queries. Furthermore, we design a novel suppression mechanism in which mutual temporal overlaps across queries are penalized to tackle temporal redundancy, supervising the model to learn distinct, non-overlapping event regions for more precise localization. Additionally, we introduce a lightweight module that captures core event concepts to further enhance semantic richness in captions through concept-level representations. We demonstrate the effectiveness of our method through extensive experiments on major DVC benchmarks YouCook2 and ActivityNet Captions.

cs.CV

MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages

Social media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automatic processing. Since language use is more informal, spontaneous, and adheres to many different sociolects, the performance of NLP models often deteriorates. One solution to this problem is to transform data to a standard variant before processing it, which is also called lexical normalization. There has been a wide variety of benchmarks and models proposed for this task. The MultiLexNorm benchmark proposed to unify these efforts, but it consists almost solely of languages from the Indo-European language family in the Latin script. Hence, we propose an extension to MultiLexNorm, which covers 5 Asian languages from different language families in 4 different scripts. We show that the previous state-of-the-art model performs worse on the new languages and propose a new architecture based on Large Language Models (LLMs), which shows more robust performance. Finally, we analyze remaining errors, revealing future directions for this task.

cs.CL

AI-Enhanced High-Density NIRS Patch for Real-Time Brain Layer Oxygenation Monitoring in Neurological Emergencies

Photon scattering has traditionally limited the ability of near-infrared spectroscopy (NIRS) to extract accurate, layer-specific information from the brain. This limitation restricts its clinical utility for precise neurological monitoring. To address this, we introduce an AI-driven, high-density NIRS system optimized to provide real-time, layer-specific oxygenation data from the brain cortex, specifically targeting acute neuro-emergencies. Our system integrates high-density NIRS reflectance data with a neural network trained on MRI-based synthetic datasets. This approach achieves robust cortical oxygenation accuracy across diverse anatomical variations. In simulations, our AI-assisted NIRS demonstrated a strong correlation (R2=0.913) with actual cortical oxygenation, markedly outperforming conventional methods (R2=0.469). Furthermore, biomimetic phantom experiments confirmed its superior anatomical reliability (R2=0.986) compared to standard commercial devices (R2=0.823). In clinical validation with healthy subjects and ischemic stroke patients, the system distinguished between the two groups with an AUC of 0.943. This highlights its potential as an accessible, high-accuracy diagnostic tool for emergency and point-of-care settings. These results underscore the system's capability to advance neuro-monitoring precision through AI, enabling timely, data-driven decisions in critical care environments.

q-bio.NC

Reduced Variability in Threshold Switches Using Heterostructures of SiO${_x}$ and Vertically Aligned MoS${_2}$

Layered two-dimensional (2D) materials provide unique structural features, such as physical gaps between their layers that are only connected through van der Waals (vdW) forces. These vdW gaps can guide the migration of intercalated ions and thus regulate filament growth in resistive switching (RS) devices. Vertically aligned 2D materials and their heterostructures provide vdW gap-mediated ion transport in memristor crossbars, providing great potential for high-density integration and reliable RS performance. Nevertheless, the fundamental switching mechanisms and their contributions to the RS remain inadequately understood. In this work, we investigate silver (Ag) filament-based threshold switching (TS) in heterostructures comprising vertically aligned 2D molybdenum disulfide (VAMoS${_2}$) grown via sulfurization and silicon oxide (SiO${_x}$). Compared to SiO${_x}$-only devices, the SiO${_x}$/VAMoS${_2}$ devices exhibit TS with higher on-threshold and hold voltages, each approximately 0.4 V, faster switching times down to 356 ns under a 4 V pulse, and a lower cycle-to-cycle on-current variability of 3.0%. A physics-based, variability-aware model reveals that confined Ag ion migration within the vdW gaps in VAMoS${_2}$ forms ultrathin seed filaments, which guide filament growth in the SiO${_x}$ layer. These findings establish SiO${_x}$/VAMoS${_2}$ heterostructures as a promising concept for reliable TS in vertical device architectures for emerging memories and neuromorphic computing.

physics.app-ph

Intermediate Resistive State in Wafer-Scale MoS${_2}$ Memristors through Lateral Silver Filament Growth for Artificial Synapse Applications

Memristors based on two-dimensional materials (2DMs) have garnered significant attention due to their fast resistive switching (RS) behavior and atomic-level thickness, which enables low power consumption, making them promising candidates for neuromorphic computing. Among these, memristors based on molybdenum disulfide (MoS${_2}$) have been extensively studied. Their RS has been attributed to the formation and rupture of conductive filaments (CFs). However, the underlying mechanism of filament formation remains underexplored, and the inherently stochastic nature of RS leads to high variability and limited reproducibility. Additionally, the lack of scalable fabrication techniques for 2DM-based memristors restricts their integration into standard semiconductor technology. Here, we demonstrate memristors based on metal-organic chemical vapor-deposited MoS${_2}$ on the wafer-scale. Our devices exhibit volatile and nonvolatile RS behavior, tunable by modulating the current compliance. Notably, we observe stable RS characteristics in an intermediate resistive state (IRS), featuring set and reset voltages within $\pm$1 V, an endurance exceeding 2500 cycles in direct current mode, and a state retention over 10${^6}$ s. The experimental data, complemented with simulations, suggest that the IRS originates from the lateral growth of the CF within the MoS${_2}$ layer. Furthermore, the devices successfully emulate synaptic plasticity with current responses on the microsecond timescale, highlighting their potential for large-scale integration in neuromorphic computing architectures.

physics.app-ph

Contrastive Representation Regularization for Vision-Language-Action Models

Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained Vision-Language Models (VLMs). However, their representations arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and proprioceptive information. To address the issue, we introduce Robot State-aware Contrastive Loss (RS-CL), a simple and effective representation regularization for VLA models, designed to bridge the gap between VLM representations and robotic signals. In particular, RS-CL aligns the representations more closely with the robot's proprioceptive states by using relative distances between the states as soft supervision. Complementing the original action prediction objective, RS-CL enhances control-relevant representation learning, while being lightweight and fully compatible with standard VLA training pipelines. Our empirical results demonstrate that RS-CL substantially improves the performance of state-of-the-art VLA models; it pushes the prior art to 69.7% achieving the state-of-the-art performance on the RoboCasa-Kitchen benchmark, and boosts success rates from 45.0% to 58.3% on challenging real-robot manipulation tasks.

cs.RO

Think Clearly: Improving Reasoning via Redundant Token Pruning

Recent large language models have shown promising capabilities in long-form reasoning, following structured chains of thought before arriving at a final answer. However, we observe that these reasoning paths tend to include substantial redundancy; analyzing attention patterns reveals that attention scores are widely scattered, particularly incorrect answers exhibit greater attention sparsity. In this paper, we demonstrate that deliberately removing this redundancy in the reasoning process significantly improves performance through clear thinking, i.e., removing distraction. Specifically, we systematically identify reasoning redundancy by measuring token-level attention scores to a special end-of-thinking token, which is appended to an explicit instruction inserted to conclude each intermediate reasoning step. Furthermore, we propose structure-aware pruning that prioritizes removing tokens in low-contributing reasoning chunks over individual tokens. After evicting redundant tokens, we remove the injected end-of-thinking instruction, then resume the reasoning generation. We demonstrate that our method significantly improves overall accuracy across reasoning-intensive benchmarks without any training involved. In particular, our method shows strong performance on challenging mathematical competition benchmarks such as AIME and AMC, where reasoning redundancy is more prevalent.

cs.AI

Threshold Switching in Vertically Aligned MoS${_2}$/SiO${_x}$ Heterostructures based on Silver Ion Migration

Threshold switching (TS) is a phenomenon where non-permanent changes in electrical resistance of a two-terminal device can be controlled by modulating the voltage bias. TS based on silver (Ag) conductive filaments has been observed in many materials, including layered two-dimensional (2D) transition metal dichalcogenides (TMDs). 2D TMDs are particularly promising for metal ion movement due to their van der Waals (vdW) gaps between their sheets, facilitating ion migration and filament formation without disturbing covalent chemical bonds. In this work, we demonstrate the heterostructure growth of vertically aligned molybdenum disulfide (VAMoS${_2}$) with an amorphous silicon oxide (SiO${_x}$) layer on top after sulfurization. We show that Ag ions migrate through this material stack, enabling TS. Our Ag/SiO${_x}$/VAMoS${_2}$/gold (Au) devices exhibit TS at low voltages of ~0.63 V, with high on-state currents over 200 ${\mu}$A and stable switching exceeding 10${^4}$ cycles. Moreover, we identify two rate-limiting steps for filament formation through a physics-based dynamical model and simulate the switching kinetics. Our devices show a fast on-switching time of 311 ns and spontaneous relaxation in 233 ns. These findings deepen the understanding of SiOx/MoS${_2}$-based RS devices and demonstrate the promise for applications in emerging memories and neuromorphic computing systems.

cond-mat.mtrl-sci

Volatile and Nonvolatile Resistive Switching in Lateral 2D Molybdenum Disulfide-Based Memristive Devices

Developing electronic devices capable of emulating biological functions is essential for advancing brain-inspired computation paradigms such as neuromorphic computing. In recent years, two-dimensional materials have emerged as promising candidates for neuromorphic electronic devices. This work addresses the coexistence of volatile and nonvolatile resistive switching in lateral memristors based on molybdenum disulfide with silver as the active electrode. The fabricated devices exhibited switching voltages of ~0.16 V and ~0.52 V for volatile and nonvolatile operation, respectively, under direct-current measurements. They also displayed the essential synaptic functions of paired-pulse facilitation and short- and long-term plasticity under pulse stimulation. The operation mechanism was investigated by in-situ transmission electron microscopy, which showed lateral migration of silver ions along the molybdenum disulfide between electrodes. Based on the experimental data, a macroscopic semi-classical electron transport model was used to reproduce the current-voltage characteristics and support the proposed underlying switching mechanisms.

cond-mat.mtrl-sci

SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL

Text-to-SQL aims to convert natural language questions into executable SQL queries. While previous approaches, such as skeleton-masked selection, have demonstrated strong performance by retrieving similar training examples to guide large language models (LLMs), they struggle in real-world scenarios where such examples are unavailable. To overcome this limitation, we propose Self-Augmentation in-context learning with Fine-grained Example selection for Text-to-SQL (SAFE-SQL), a novel framework that improves SQL generation by generating and filtering self-augmented examples. SAFE-SQL first prompts an LLM to generate multiple Text-to-SQL examples relevant to the test input. Then SAFE-SQL filters these examples through three relevance assessments, constructing high-quality in-context learning examples. Using self-generated examples, SAFE-SQL surpasses the previous zero-shot, and few-shot Text-to-SQL frameworks, achieving higher execution accuracy. Notably, our approach provides additional performance gains in extra hard and unseen scenarios, where conventional methods often fail.

cs.CL

Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving

Humans have long relied on visual aids like sketches and diagrams to support reasoning and problem-solving. Visual tools, like auxiliary lines in geometry or graphs in calculus, are essential for understanding complex ideas. However, many tutoring systems remain text-based, providing feedback only through natural language. Leveraging recent advances in Large Multimodal Models (LMMs), this paper introduces Interactive Sketchpad, a tutoring system that combines language-based explanations with interactive visualizations to enhance learning. Built on a pre-trained LMM, Interactive Sketchpad is fine-tuned to provide step-by-step guidance in both text and visuals, enabling natural multimodal interaction with the student. Accurate and robust diagrams are generated by incorporating code execution into the reasoning process. User studies conducted on math problems such as geometry, calculus, and trigonometry demonstrate that Interactive Sketchpad leads to improved task comprehension, problem-solving accuracy, and engagement levels, highlighting its potential for transforming educational technologies. All code is available at: https://stevenshinechen.github.io/interactivesketchpad/.

cs.HC

Influence of Humidity on the Resistive Switching of Hexagonal Boron Nitride-Based Memristors

Two-dimensional material-based memristors have recently gained attention as components of future neuromorphic computing concepts. However, their surrounding atmosphere can influence their behavior. In this work, we investigate the resistive switching behavior of hexagonal boron nitride-based memristors with active nickel electrodes under vacuum conditions. Our cells exhibit repeatable, bipolar, nonvolatile switching under voltage stress after initial forming, with a switching window > 10${^3}$ under ambient conditions. However, in a vacuum, the forming is suppressed, and hence, no switching is observed. Compact model simulations can reproduce the set kinetics of our cells under ambient conditions and predict highly suppressed resistive switching in a water-deficient environment, supporting the experimental results. Our findings have important implications for the application of h-BN-based memristors with electrochemically active electrodes since semiconductor chips are typically processed under high vacuum conditions and encapsulated to protect them from atmospheric influences.

cond-mat.mtrl-sci

Probing-RAG: Self-Probing to Guide Language Models in Selective Document Retrieval

Retrieval-Augmented Generation (RAG) enhances language models by retrieving and incorporating relevant external knowledge. However, traditional retrieve-and-generate processes may not be optimized for real-world scenarios, where queries might require multiple retrieval steps or none at all. In this paper, we propose a Probing-RAG, which utilizes the hidden state representations from the intermediate layers of language models to adaptively determine the necessity of additional retrievals for a given query. By employing a pre-trained prober, Probing-RAG effectively captures the model's internal cognition, enabling reliable decision-making about retrieving external documents. Experimental results across five open-domain QA datasets demonstrate that Probing-RAG outperforms previous methods while reducing the number of redundant retrieval steps.

cs.CL

SoccerNet 2024 Challenges Results

The SoccerNet 2024 challenges represent the fourth annual video understanding challenges organized by the SoccerNet team. These challenges aim to advance research across multiple themes in football, including broadcast video understanding, field understanding, and player understanding. This year, the challenges encompass four vision-based tasks. (1) Ball Action Spotting, focusing on precisely localizing when and which soccer actions related to the ball occur, (2) Dense Video Captioning, focusing on describing the broadcast with natural language and anchored timestamps, (3) Multi-View Foul Recognition, a novel task focusing on analyzing multiple viewpoints of a potential foul incident to classify whether a foul occurred and assess its severity, (4) Game State Reconstruction, another novel task focusing on reconstructing the game state from broadcast videos onto a 2D top-view map of the field. Detailed information about the tasks, challenges, and leaderboards can be found at https://www.soccer-net.org, with baselines and development kits available at https://github.com/SoccerNet.

cs.CV

Volatile MoS${_2}$ Memristors with Lateral Silver Ion Migration for Artificial Neuron Applications

Layered two-dimensional (2D) semiconductors have shown enhanced ion migration capabilities along their van der Waals (vdW) gaps and on their surfaces. This effect can be employed for resistive switching (RS) in devices for emerging memories, selectors, and neuromorphic computing. To date, all lateral molybdenum disulfide (MoS${_2}$)-based volatile RS devices with silver (Ag) ion migration have been demonstrated using exfoliated, single-crystal MoS${_2}$ flakes requiring a forming step to enable RS. Here, we present volatile RS with multilayer MoS${_2}$ grown by metal-organic chemical vapor deposition (MOCVD) with repeatable forming-free operation. The devices show highly reproducible volatile RS with low operating voltages of approximately 2 V and fast switching times down to 130 ns considering their micrometer scale dimensions. We investigate the switching mechanism based on Ag ion surface migration through transmission electron microscopy, electronic transport modeling, and density functional theory. Finally, we develop a physics-based compact model and explore the implementation of our volatile memristors as artificial neurons in neuromorphic systems.

physics.app-ph