SearcharxivSearch

arXiv subjects

Hyunseok Lee

Publications and source records attributed to Hyunseok Lee.

At least 19 recordsLinked to original sources

Development and Initial Performance of an Upgraded NaI(Tl) Crystal Encapsulation for COSINE-100U

The COSINE-100 experiment was designed to test the DAMA/LIBRA annual-modulation claim using low-background NaI(Tl) detectors. For the COSINE-100U upgrade, we developed a new crystal-encapsulation system to increase light-collection efficiency while preserving long-term detector stability, thereby improving sensitivity to low-mass dark matter. The upgraded design eliminates the quartz optical windows used in COSINE-100 and directly couples the photomultiplier tubes (PMTs) to the crystal end faces through 2-mm-thick silicone optical pads, thereby reducing the number of optical interfaces. For the larger crystals, the crystal edges were beveled to guide scintillation light more efficiently onto 3-inch high-quantum-efficiency PMTs. The performance study uses 2462~h (102.6~days) of room-temperature COSINE-100U data and, for direct background comparisons, reference COSINE-100 data acquired near the end of operation. 698~h (29.1~days) of COSINE-100 data acquired near the end of operation in March 2023. All eight crystals showed higher light yields than in COSINE-100, with values ranging from 15.8 to 27.7~p.e./keV; six crystals exceeded 20~p.e./keV. The measured bulk-$α$ rates were lower than the COSINE-100 values and consistent with the expected time evolution of internal $^{210}$Pb, while the 1--2-MeV surface-$α$ rates were substantially reduced. The upgrade also restored two crystals that had previously been excluded from the COSINE-100 physics analysis because of poor optical performance. Independent validation tests demonstrated that the encapsulation remains mechanically robust and optically stable during long-term immersion in liquid scintillator at low temperature. This paper presents the encapsulation design, the room-temperature detector performance, and the reduction in surface-related backgrounds achieved at the Yemilab facility.

physics.ins-det

Topological flowscape reveals state transitions in nonreciprocal living matter

Nonreciprocal interactions -- where forces between entities are asymmetric -- govern a wide range of nonequilibrium phenomena, yet their role in structural transitions in living and active systems remains elusive. Here, we demonstrate a transition between nonreciprocal states using starfish embryos at different stages of development, where interactions are inherently asymmetric and tunable. Experiments, interaction inference, and topological analysis yield a nonreciprocal state diagram spanning crystalline, flocking, and fragmented states, revealing that weak nonreciprocity promotes structural order while stronger asymmetry disrupts it. To capture these transitions, we introduce topological landscapes, mapping the distribution of structural motifs across state space. We further develop topological flowscapes, a dynamic framework that quantifies transitions between collective states and detects an informational rate shift from the experimental state transition. Together, these results establish a general approach for decoding nonequilibrium transitions and uncover how asymmetric interactions sculpt the dynamical and structural architecture of active and living matter.

cond-mat.soft

Naive Visual Memory is Not Enough: A Failure-Mode Study of GUI Agents

Graphical User Interface (GUI) agents are increasingly used to automate complex computer tasks across applications, websites, and operating systems. To improve their reliability, recent work has introduced experiential memory, where agents retrieve prior trajectories to guide decision-making in similar states. More recent approaches further extend this idea to visual memory by storing and retrieving screenshots from past interactions, providing agents with richer contextual information than text-only memories. However, the effect of visual memory in GUI agents remains insufficiently understood: it is unclear which failures visual memory mitigates, or which failures it exacerbates. To systematically analyze the effect of visual memory, we introduce a taxonomy of four GUI agent failures (i.e., cognitive failure, visual state misunderstanding, hidden operation blindness, and grounding error) that map to distinct stages of the perception-reasoning-action pipeline. We find that prepending full-image memory has a divergent effect on the failure distribution: it reduces state-level failures but worsens action-level ones, and increases hidden operation blindness and grounding error. Motivated by this finding, we propose Action-Grounded Visual Memory (AGMem), an action-grounded memory framework for GUI agents. The core idea of AGMem is to store image crops that capture the local GUI region closely related to a successful action or a recovery, rather than storing full screenshots. Experiments on OSWorld show that AGMem improves task success rates by 33.3 % over full-image memory. These results demonstrate that AGMem is an effective representation for visual memory in GUI agents.

cs.MA

Benchmarking Visual State Tracking in Multimodal Video Understanding

Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacity for visual state tracking is fundamental to video understanding, yet remains underexplored in current evaluations of Multimodal Large Language Models (MLLMs). We introduce Visual STAte Tracking benchmark (VSTAT), a video-based benchmark designed to diagnose visual state tracking in MLLMs. VSTAT consists of 834 clips drawn from both synthetic and real-world videos, paired with 1,500 questions that cannot be answered from any single frame or short segment, requiring continuous perception and integration of events across the entire video stream. Despite their strong performance on existing video benchmarks, we find that state-of-the-art MLLMs perform far below humans and only modestly above answer-prior baselines. To analyze this gap, we compare MLLMs' thinking traces with the underlying video stream to understand why and when MLLMs fail on VSTAT. We find that MLLMs reason and track correctly in text, but fail at visually perceiving the events they need to track. Finally, our preliminary evaluation suggests that recent agentic approaches, including MLLM-based video agents and coding agents, do not readily resolve these failures, still falling short on VSTAT.

cs.CV

Vision-aligned Latent Reasoning for Multi-modal Large Language Model

Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems which require extensive multi-step reasoning. This is primarily due to the progressive dilution of visual information during long-context generation, which hinders their ability to fully exploit test-time scaling. To address this issue, we introduce Vision-aligned Latent Reasoning (VaLR), a simple, yet effective reasoning framework that dynamically generates vision-aligned latent tokens before each Chain of Thought reasoning step, guiding the model to reason based on perceptual cues in the latent space. Specifically, VaLR is trained to preserve visual knowledge during reasoning by aligning intermediate embeddings of MLLM with those from vision encoders. Empirical results demonstrate that VaLR consistently outperforms existing approaches across a wide range of benchmarks requiring long-context understanding or precise visual perception, while exhibiting test-time scaling behavior not observed in prior MLLMs. In particular, VaLR improves the performance significantly from 33.0% to 52.9% on VSI-Bench, achieving a 19.9%p gain over Qwen2.5-VL.

cs.CV

Cog3DMap: Multi-View Vision-Language Reasoning with 3D Cognitive Maps

Precise spatial understanding from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs), as their visual representations are predominantly semantic and lack explicit geometric grounding. While existing approaches augment visual tokens with geometric cues from visual geometry models, their MLLM is still required to implicitly infer the underlying 3D structure of the scene from these augmented tokens, limiting its spatial reasoning capability. To address this issue, we introduce Cog3DMap, a framework that recurrently constructs an explicit 3D memory from multi-view images, where each token is grounded in 3D space and possesses both semantic and geometric information. By feeding these tokens into the MLLM, our framework enables direct reasoning over a spatially structured 3D map, achieving state-of-the-art performance on various spatial reasoning benchmarks. Code will be made publicly available.

cs.CV

New sector morphologies emerge from anisotropic colony growth

Competition during range expansions is of great interest from both practical and theoretical view points. Experimentally, range expansions are often studied in homogeneous Petri dishes, which lack spatial anisotropy that might be present in realistic populations. Here, we analyze a model of anisotropic growth, based on coupled Kardar-Parisi-Zhang and Fisher-Kolmogorov-Petrovsky-Piskunov equations that describe surface growth and lateral competition. Compared to a previous study of isotropic growth, anisotropy relaxes a constraint between parameters of the model. We completely characterize spatial patterns and invasion velocities in this generalized model. In particular, we find that strong anisotropy results in a distinct morphology of spatial invasion with a kink in the displaced strain ahead of the boundary between the strains. This morphology of the out-competed strain is similar to a shock wave and serves as a signature of anisotropic growth.

nlin.PS

Beyond Correctness: Learning Robust Reasoning via Transfer

Reinforcement Learning with Verifiable Rewards (RLVR) has recently strengthened LLM reasoning, but its focus on final answer correctness leaves a critical gap: it does not ensure the robustness of the reasoning process itself. We adopt a simple philosophical view, robust reasoning should remain useful beyond the mind that produced it, and treat reasoning as a form of meaning transfer that must survive truncation, reinterpretation, and continuation. Building on this principle, we introduce Reinforcement Learning with Transferable Reward (RLTR), which operationalizes robustness via transfer reward that tests whether a partial reasoning prefix from one model can guide a separate model to the correct answer. This encourages LLMs to produce reasoning that is stable, interpretable, and genuinely generalizable. Our approach improves sampling consistency while improving final answer accuracy, and it reaches comparable performance in substantially fewer training steps. For example, on MATH500, RLTR achieves a +3.6%p gain in Maj@64 compared to RLVR and matches RLVR's average accuracy with roughly 2.5x fewer training steps, providing both more reliable reasoning and significantly more sample efficient.

cs.LG

Degenerate Euler- Seidel Method for degenerate Bernoulli, Euler, and Genocchi polynomials

This paper introduces a degenerate version of the Euler-Seidel method by incorporating a parameter lambda into the classical recurrence relation. We define a degenerate Euler-Seidel matrix associated with an initial sequence and establish corresponding lambda-generalized binomial identities and generating function relations. By applying this method to the degenerate Bernoulli, Euler, and Genocchi polynomials, we derive several new combinatorial identities. This work extends the classical Euler-Seidel method to the domain of degenerate special polynomials and numbers, providing a new framework for studying their properties.

math.NT

ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification

Self-awareness, i.e., the ability to assess and correct one's own generation, is a fundamental aspect of human intelligence, making its replication in large language models (LLMs) an important yet challenging task. Previous works tackle this by employing extensive reinforcement learning or rather relying on large external verifiers. In this work, we propose Refine via Intrinsic Self-Verification (ReVISE), an efficient and effective framework that enables LLMs to self-correct their outputs through self-verification. The core idea of ReVISE is to enable LLMs to verify their reasoning processes and continually rethink reasoning trajectories based on its verification. We introduce a structured curriculum based upon online preference learning to implement this efficiently. Specifically, as ReVISE involves two challenging tasks (i.e., self-verification and reasoning correction), we tackle each task sequentially using curriculum learning, collecting both failed and successful reasoning paths to construct preference pairs for efficient training. During inference, our approach enjoys natural test-time scaling by integrating self-verification and correction capabilities, further enhanced by our proposed confidence-aware decoding mechanism. Our experiments on various reasoning tasks demonstrate that ReVISE achieves efficient self-correction and significantly improves reasoning performance.

cs.LG

ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search

Recent advances in Multimodal Large Language Models (MLLMs) have enabled autonomous agents to interact with computers via Graphical User Interfaces (GUIs), where accurately localizing the coordinates of interface elements (e.g., buttons) is often required for fine-grained actions. However, this remains significantly challenging, leading prior works to rely on large-scale web datasets to improve the grounding accuracy. In this work, we propose Reasoning Graphical User Interface Grounding for Data Efficiency (ReGUIDE), a novel and effective framework for web grounding that enables MLLMs to learn data efficiently through self-generated reasoning and spatial-aware criticism. More specifically, ReGUIDE learns to (i) self-generate a language reasoning process for the localization via online reinforcement learning, and (ii) criticize the prediction using spatial priors that enforce equivariance under input transformations. At inference time, ReGUIDE further boosts performance through a test-time scaling strategy, which combines spatial search with coordinate aggregation. Our experiments demonstrate that ReGUIDE significantly advances web grounding performance across multiple benchmarks, outperforming baselines with substantially fewer training data points (e.g., only 0.2% samples compared to the best open-sourced baselines).

cs.LG

New Constraints on Axion-Like Particles with the NEON Detector at a Nuclear Reactor

We report new constraints on axion-like particles (ALPs) using data from the NEON experiment, which features a 16.7 kg of NaI(Tl) target located 23.7 meters from a 2.8 GW thermal power nuclear reactor. Analyzing a total exposure of 3063 kg$\cdot$days, with 1596 kg$\cdot$days during reactor-on and 1467 kg$\cdot$days during reactor-off periods, we compared energy spectra to search for ALP-induced signals. No significant signal was observed, enabling us to set exclusion limits at the 95\% confidence level. These limits probe previously unexplored regions of the ALP parameter space, particularly for axion mass ($m_a$) near $1$ MeV/c$^2$. For ALP-photon coupling (${g_{aγ}}$), limits reach as low as 6.24$\times$ 10$^{-6}$ GeV$^{-1}$ at $m_a$ = 3.0 MeV/c$^2$, while for ALP-electron coupling (${g_{ae}}$), limits reach 4.95$\times$ 10$^{-8}$ at $m_a$ = 1.02 MeV/c$^2$. This work demonstrates the potential for future reactor experiments to probe unexplored ALP parameter space.

hep-ex

ReMoDetect: Reward Models Recognize Aligned LLM's Generations

The remarkable capabilities and easy accessibility of large language models (LLMs) have significantly increased societal risks (e.g., fake news generation), necessitating the development of LLM-generated text (LGT) detection methods for safe usage. However, detecting LGTs is challenging due to the vast number of LLMs, making it impractical to account for each LLM individually; hence, it is crucial to identify the common characteristics shared by these models. In this paper, we draw attention to a common feature of recent powerful LLMs, namely the alignment training, i.e., training LLMs to generate human-preferable texts. Our key finding is that as these aligned LLMs are trained to maximize the human preferences, they generate texts with higher estimated preferences even than human-written texts; thus, such texts are easily detected by using the reward model (i.e., an LLM trained to model human preference distribution). Based on this finding, we propose two training schemes to further improve the detection ability of the reward model, namely (i) continual preference fine-tuning to make the reward model prefer aligned LGTs even further and (ii) reward modeling of Human/LLM mixed texts (a rephrased texts from human-written texts using aligned LLMs), which serves as a median preference text corpus between LGTs and human-written texts to learn the decision boundary better. We provide an extensive evaluation by considering six text domains across twelve aligned LLMs, where our method demonstrates state-of-the-art results. Code is available at https://github.com/hyunseoklee-ai/ReMoDetect.

cs.LG

Selective excitation of work-generating cycles in nonreciprocal living solids

Emergent nonreciprocity in active matter drives the formation of self-organized states that transcend the behaviors of equilibrium systems. Integrating experiments, theory and simulations, we demonstrate that active solids composed of living starfish embryos spontaneously transition between stable fluctuating and oscillatory steady states. The nonequilibrium steady states arise from two distinct chiral symmetry breaking mechanisms at the microscopic scale: the spinning of individual embryos resulting in a macroscopic odd elastic response, and the precession of their rotation axis, leading to active gyroelasticity. In the oscillatory state, we observe long-wavelength optical vibrational modes that can be excited through mechanical perturbations. Strikingly, these excitable nonreciprocal solids exhibit nonequilibrium work generation without cycling protocols, due to coupled vibrational modes. Our work introduces a novel class of tunable nonequilibrium processes, offering a framework for designing and controlling soft robotic swarms and adaptive active materials, while opening new possibilities for harnessing nonreciprocal interactions in engineered systems.

cond-mat.soft

Deep Learning Techniques for Automatic Lateral X-ray Cephalometric Landmark Detection: Is the Problem Solved?

Localization of the craniofacial landmarks from lateral cephalograms is a fundamental task in cephalometric analysis. The automation of the corresponding tasks has thus been the subject of intense research over the past decades. In this paper, we introduce the "Cephalometric Landmark Detection (CL-Detection)" dataset, which is the largest publicly available and comprehensive dataset for cephalometric landmark detection. This multi-center and multi-vendor dataset includes 600 lateral X-ray images with 38 landmarks acquired with different equipment from three medical centers. The overarching objective of this paper is to measure how far state-of-the-art deep learning methods can go for cephalometric landmark detection. Following the 2023 MICCAI CL-Detection Challenge, we report the results of the top ten research groups using deep learning methods. Results show that the best methods closely approximate the expert analysis, achieving a mean detection rate of 75.719% and a mean radial error of 1.518 mm. While there is room for improvement, these findings undeniably open the door to highly accurate and fully automatic location of craniofacial landmarks. We also identify scenarios for which deep learning methods are still failing. Both the dataset and detailed results are publicly available online, while the platform will remain open for the community to benchmark future algorithm developments at https://cl-detection2023.grand-challenge.org/.

cs.CV

Competition on the edge of an expanding population

In growing populations, the fate of mutations depends on their competitive ability against the ancestor and their ability to colonize new territory. Here we present a theory that integrates both aspects of mutant fitness by coupling the classic description of one-dimensional competition (Fisher equation) to the minimal model of front shape (KPZ equation). We solved these equations and found three regimes, which are controlled solely by the expansion rates, solely by the competitive abilities, or by both. Collectively, our results provide a simple framework to study spatial competition.

q-bio.PE

Laser pulse compression by a density gradient plasma for exawatt to zettawatt lasers

We propose a new method for laser pulse compression that uses the spatially varying dispersion of a plasma plume with a density gradient. This novel scheme can be used to compress ultrahigh power lasers. A long, negatively frequency-chirped, laser pulse reflects off the plasma ramp of an over-dense plasma. As the density increases longitudinally the high frequency photons at the leading part of the laser pulse propagates more deeply than low frequency photons, the pulse is compressed in a similar way to compression off a chirped mirror. Proof-of-principle simulations, using a one-dimensional (1-D) particle-in-cell (PIC) simulation code demonstrates the compression of 2.35 ps laser pulse to 10.3 fs, with a compression ratio of 225. As plasmas is robust and resistant to high intensities, unlike gratings in a chirped-pulse amplification (CPA) technique [1], the method could be used as a compressor to reach exawatt or zettawatt peak power lasers.

physics.plasm-ph

Assessing generalisability of deep learning-based polyp detection and segmentation methods through a computer vision challenge

Polyps are well-known cancer precursors identified by colonoscopy. However, variability in their size, location, and surface largely affect identification, localisation, and characterisation. Moreover, colonoscopic surveillance and removal of polyps (referred to as polypectomy ) are highly operator-dependent procedures. There exist a high missed detection rate and incomplete removal of colonic polyps due to their variable nature, the difficulties to delineate the abnormality, the high recurrence rates, and the anatomical topography of the colon. There have been several developments in realising automated methods for both detection and segmentation of these polyps using machine learning. However, the major drawback in most of these methods is their ability to generalise to out-of-sample unseen datasets that come from different centres, modalities and acquisition systems. To test this hypothesis rigorously we curated a multi-centre and multi-population dataset acquired from multiple colonoscopy systems and challenged teams comprising machine learning experts to develop robust automated detection and segmentation methods as part of our crowd-sourcing Endoscopic computer vision challenge (EndoCV) 2021. In this paper, we analyse the detection results of the four top (among seven) teams and the segmentation results of the five top teams (among 16). Our analyses demonstrate that the top-ranking teams concentrated on accuracy (i.e., accuracy > 80% on overall Dice score on different validation sets) over real-time performance required for clinical applicability. We further dissect the methods and provide an experiment-based hypothesis that reveals the need for improved generalisability to tackle diversity present in multi-centre datasets.

cs.CV