SearcharxivSearch

arXiv subjects

Jinwon Lee

Publications and source records attributed to Jinwon Lee.

At least 19 recordsLinked to original sources

Direct Observation of Channelised Supercurrents in a Kagome Superconductor

Superconductors are many-body quantum states in which current flows without dissipation. Theory predicts that supercurrents follow a relatively simple spatial pattern in both conventional and unconventional superconductors. Recent studies into the AV3Sb5 (A = Cs, K, Rb) family of Kagome superconductors indicate that CsV3Sb5 has unconventional transport properties that cannot be accounted for with these simple theories, including reports of intrinsic Josephson junctions, higher order Cooper pairing and the zero field diode effect. Attempts to interpret these findings have focused on the interplay of superconductivity with the unconventional charge density wave (CDW) order in these materials, with which superconductivity competes. A current roadblock to understanding how these kagome superconductors give rise to their intriguing properties is the lack of spatially resolved information about transport. Here we show, using a recently developed superconducting quantum interference device (SQUID) microscope, that flakes of CsV3Sb5-xSnx host a network of narrow supercurrent channels. These supercurrent channels emerge at the critical temperature and remain stable for all temperatures and currents. Their non-linear behaviour is consistent with a network of Josephson junctions linked by narrow supercurrent filaments, which naturally leads to the observed transport anomalies. Intriguingly, these observations are much weaker in undoped samples, which suggests links to the physics of charge density waves, disorder, and electronic correlations, all of which are greatly influenced by the doping strength. These results establish new frontiers for the local investigation of charge transport and competing orders in strongly correlated electron systems, and shine a new light on the anomalous transport properties of the AV3Sb5 kagome superconductors.

cond-mat.supr-con

Nuclear magnetic resonance on a single atom with a local probe

The nuclear spin is a prime candidate for quantum information applications due to its weak coupling to the environment and inherently long coherence times. However, this weak coupling also challenges the addressability of the nuclear spin. Here we demonstrate nuclear magnetic resonance (NMR) on a single on-surface atom using a local scanning probe. We employ an electron-nuclear double resonance measurement scheme and resolve nuclear spin transitions of a single 47Ti isotope with a nuclear spin of I = 5/2. The quadrupole interaction enables to resolve multiple NMR transitions, which are consistent with our eigenenergy calculations. Our experimental results indicate that the nuclear spin can be driven efficiently irrespective of its hybridization with the electron spin, which is required for direct control of the nuclear spin in the long-lifetime regime. This investigation of NMR on a single atom in a platform with atomic-scale control is a valuable development for other platforms deploying nuclear spins for characterization techniques or quantum information technology.

quant-ph

The Path Not Taken: RLVR Provably Learns Off the Principals

Reinforcement Learning with Verifiable Rewards (RLVR) reliably improves the reasoning performance of large language models, yet it appears to modify only a small fraction of parameters. We revisit this paradox and show that sparsity is a surface artifact of a model-conditioned optimization bias: for a fixed pretrained model, updates consistently localize to preferred parameter regions, highly consistent across runs and largely invariant to datasets and RL recipes. We mechanistically explain these dynamics with a Three-Gate Theory: Gate I (KL Anchor) imposes a KL-constrained update; Gate II (Model Geometry) steers the step off principal directions into low-curvature, spectrum-preserving subspaces; and Gate III (Precision) hides micro-updates in non-preferred regions, making the off-principal bias appear as sparsity. We then validate this theory and, for the first time, provide a parameter-level characterization of RLVR's learning dynamics: RLVR learns off principal directions in weight space, achieving gains via minimal spectral drift, reduced principal-subspace rotation, and off-principal update alignment. In contrast, SFT targets principal weights, distorts the spectrum, and even lags RLVR. Together, these results provide the first parameter-space account of RLVR's training dynamics, revealing clear regularities in how parameters evolve. Crucially, we show that RL operates in a distinct optimization regime from SFT, so directly adapting SFT-era parameter-efficient fine-tuning (PEFT) methods can be flawed, as evidenced by our case studies on advanced sparse fine-tuning and LoRA variants. We hope this work charts a path toward a white-box understanding of RLVR and the design of geometry-aware, RLVR-native learning algorithms, rather than repurposed SFT-era heuristics.

cs.LG

High-Resolution Casimir Force Sensing Across a Superconducting Transition

The Casimir effect and superconductivity are foundational quantum phenomena whose interplay is an open question in physics, with significant implications for electron physics, quantum gravity, and high-temperature superconductivity. Determining how Casimir forces behave across a superconducting transition remains elusive due to the difficulty of realizing precise alignment, cryogenic operation, and isolating small force changes from competing effects. Recent theories predict milli-Pascal jumps in Casimir pressure across the transition, motivating experiments capable of reaching well below this regime. Here, we demonstrate an on-chip superconducting nanomechanical platform that overcomes these long-standing challenges, achieving the most parallel Casimir configurations to date. Our microchip-based parallel plates reach unprecedented area-to-separation ratios, exceeding past experiments across superconducting transitions by three orders of magnitude and yielding the strongest Casimir forces generated between compliant surfaces. Scanning tunneling microscopy (STM) directly detects the resonant motion of a suspended nanoscale plate with subatomic precision in lateral positioning and displacement, enabling suppression of van der Waals, electrostatic, and thermal effects. With verified micro-Pascal pressure resolution, our platform provides a credible entry point into a new field of quantum experiments, enabling exploration of Casimir-superconductivity interactions with the stability, parallelism, and sensitivity required to access this regime of physics.

quant-ph

APOLLO: SGD-like Memory, AdamW-level Performance

Large language models (LLMs) are notoriously memory-intensive during training, particularly with the popular AdamW optimizer. This memory burden necessitates using more or higher-end GPUs or reducing batch sizes, limiting training scalability and throughput. To address this, various memory-efficient optimizers have been proposed to reduce optimizer memory usage. However, they face critical challenges: (i) reliance on costly SVD operations; (ii) significant performance trade-offs compared to AdamW; and (iii) still substantial optimizer memory overhead to maintain competitive performance. In this work, we identify that AdamW's learning rate adaptation rule can be effectively coarsened as a structured learning rate update. Based on this insight, we propose Approximated Gradient Scaling for Memory-Efficient LLM Optimization (APOLLO), which approximates learning rate scaling using an auxiliary low-rank optimizer state based on pure random projection. This structured learning rate update rule makes APOLLO highly tolerant to further memory reductions while delivering comparable pre-training performance. Even its rank-1 variant, APOLLO-Mini, achieves superior pre-training performance compared to AdamW with SGD-level memory costs. Extensive experiments demonstrate that the APOLLO series performs on-par with or better than AdamW, while achieving greater memory savings by nearly eliminating the optimization states of AdamW. These savings provide significant system-level benefits: (1) Enhanced Throughput: 3x throughput on an 8xA100-80GB setup compared to AdamW by supporting 4x larger batch sizes. (2) Improved Model Scalability: Pre-training LLaMA-13B with naive DDP on A100-80GB GPUs without system-level optimizations. (3) Low-End GPU Friendly Pre-training: Pre-training LLaMA-7B on a single GPU using less than 12 GB of memory with weight quantization.

cs.LG

Single-shot readout of the nuclear spin of an on-surface atom

Nuclear spins owe their long-lived magnetic states to their excellent isolation from the environment. At the same time, a finite degree of interaction with their surroundings is necessary for reading and writing the spin state. Therefore, detailed knowledge of and control over the atomic environment of a nuclear spin is key to optimizing conditions for quantum information applications. While various platforms enabled single-shot readout of nuclear spins, their direct environments were either unknown or impossible to controllably modify on the atomic scale. Scanning tunnelling microscopy (STM), combined with electron spin resonance (ESR), provides atomic-scale information of individual nuclear spins via the hyperfine interaction. Here, we demonstrate single-shot readout of an individual $^{\text{49}}$Ti nuclear spin with an STM. Employing a pulsed measurement scheme, we find its lifetime to be in the order of seconds. Furthermore, we shed light on the pumping and relaxation mechanisms of the nuclear spin by investigating its response to both ESR driving and tunnelling current, which is supported by model calculations. These findings give an atomic-scale insight into the nature of nuclear spin relaxation and are relevant for the development of atomically assembled qubit platforms.

cond-mat.mes-hall

Signatures of Amorphous Shiba State in FeTe$_{0.55}$Se$_{0.45}$

The iron-based superconductor FeTe$_{0.55}$Se$_{0.45}$ is a peculiar material: it hosts a surface state with a Dirac dispersion, is a putative topological superconductor hosting Majorana modes in vortices, and has an unusually low Fermi energy. The superconducting state is generally thought to be characterized by three gaps in different bands, with the usual homogenous, spatially extended Bogoliubov excitations -- in this work, we uncover evidence that it is instead of a very different nature. Our scanning tunneling spectroscopy data shows several peaks in the density of states above a full gap, and by analyzing the spatial and junction-resistance dependence of the peaks, we conclude that the peaks above the first one are not coherence peaks from different bands. Instead, comparisons with our simulations indicate that they originate from generalized Shiba states that are spatially overlapping. This can lead to an amorphous state of Bogoliubov quasiparticles, reminiscent of impurity bands in semiconductors. We discuss the origin and implications of this new state.

cond-mat.supr-con

LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference

The explosive arrival of OpenAI's ChatGPT has fueled the globalization of large language model (LLM), which consists of billions of pretrained parameters that embodies the aspects of syntax and semantics. HyperAccel introduces latency processing unit (LPU), a latency-optimized and highly scalable processor architecture for the acceleration of LLM inference. LPU perfectly balances the memory bandwidth and compute logic with streamlined dataflow to maximize performance and efficiency. LPU is equipped with expandable synchronization link (ESL) that hides data synchronization latency between multiple LPUs. HyperDex complements LPU as an intuitive software framework to run LLM applications. LPU achieves 1.25 ms/token and 20.9 ms/token for 1.3B and 66B model, respectively, which is 2.09x and 1.37x faster than the GPU. LPU, synthesized using Samsung 4nm process, has total area of 0.824 mm2 and power consumption of 284.31 mW. LPU-based servers achieve 1.33x and 1.32x energy efficiency over NVIDIA H100 and L4 servers, respectively.

cs.AR

Generative Model-based Simulation of Driver Behavior when Using Control Input Interface for Teleoperated Driving in Unstructured Canyon Terrains

Unmanned ground vehicles (UGVs) in unstructured environments mostly operate through teleoperation. To enable stable teleoperated driving in unstructured environments, some research has suggested driver assistance and evaluation methods that involve user studies, which can be costly and require lots of time and effort. A simulation model-based approach has been proposed to complement the user study; however, the models on teleoperated driving do not account for unstructured environments. Our proposed solution involves simulation models of teleoperated driving for drivers that utilize a deep generative model. Initially, we build a teleoperated driving simulator to imitate unstructured environments based on previous research and collect driving data from drivers. Then, we design and implement the simulation models based on a conditional variational autoencoder (CVAE). Our evaluation results demonstrate that the proposed teleoperated driving model can generate data by simulating the driver appropriately in unstructured canyon terrains.

cs.RO

Mobile Kink Solitons in a Van der Waals Charge-Density-Wave Layer

Kinks, point-like geometrical defects along dislocations, domain walls, and DNA, are stable and mobile, as solutions of a sine-Gordon wave equation. While they are widely investigated for crystal deformations and domain wall motions, electronic properties of individual kinks have received little attention. In this work, electronically and topologically distinct kinks are discovered along electronic domain walls in a correlated van der Waals insulator of 1$T$-TaS$_2$. Mobile kinks and antikinks are identified as trapped by pinning defects and imaged in scanning tunneling microscopy. Their atomic structures and in-gap electronic states are unveiled, which are mapped approximately into Su-Schrieffer-Heeger solitons. The twelve-fold degeneracy of the domain walls in the present system guarantees an extraordinarily large number of distinct kinks and antikinks to emerge. Such large degeneracy together with the robust geometrical nature may be useful for handling multilevel information in van der Waals materials architectures.

cond-mat.mtrl-sci

Machining feature recognition using descriptors with range constraints for mechanical 3D models

In machining feature recognition, geometric elements generated in a three-dimensional computer-aided design model are identified. This technique is used in manufacturability evaluation, process planning, and tool path generation. Here, we propose a method of recognizing 16 types of machining features using descriptors, often used in shape-based part retrieval studies. The base face is selected for each feature type, and descriptors express the base face's minimum, maximum, and equal conditions. Furthermore, the similarity in the three conditions between the descriptors extracted from the target face and those from the base face is calculated. If the similarity is greater than or equal to the threshold, the target face is determined as the base face of the feature. Machining feature recognition tests were conducted for two test cases using the proposed method, and all machining features included in the test cases were successfully recognized. Also, it was confirmed through an additional test that the proposed method in this study showed better feature recognition performance than the latest artificial neural network.

cs.CG

GoonDAE: Denoising-Based Driver Assistance for Off-Road Teleoperation

Because of the limitations of autonomous driving technologies, teleoperation is widely used in dangerous environments such as military operations. However, the teleoperated driving performance depends considerably on the driver's skill level. Moreover, unskilled drivers need extensive training time for teleoperations in unusual and harsh environments. To address this problem, we propose a novel denoising-based driver assistance method, namely GoonDAE, for real-time teleoperated off-road driving. The unskilled driver control input is assumed to be the same as the skilled driver control input but with noise. We designed a skip-connected long short-term memory (LSTM)-based denoising autoencoder (DAE) model to assist the unskilled driver control input by denoising. The proposed GoonDAE was trained with skilled driver control input and sensor data collected from our simulated off-road driving environment. To evaluate GoonDAE, we conducted an experiment with unskilled drivers in the simulated environment. The results revealed that the proposed system considerably enhanced driving performance in terms of driving stability.

cs.RO

Learned Threshold Pruning

This paper presents a novel differentiable method for unstructured weight pruning of deep neural networks. Our learned-threshold pruning (LTP) method learns per-layer thresholds via gradient descent, unlike conventional methods where they are set as input. Making thresholds trainable also makes LTP computationally efficient, hence scalable to deeper networks. For example, it takes $30$ epochs for LTP to prune ResNet50 on ImageNet by a factor of $9.1$. This is in contrast to other methods that search for per-layer thresholds via a computationally intensive iterative pruning and fine-tuning process. Additionally, with a novel differentiable $L_0$ regularization, LTP is able to operate effectively on architectures with batch-normalization. This is important since $L_1$ and $L_2$ penalties lose their regularizing effect in networks with batch-normalization. Finally, LTP generates a trail of progressively sparser networks from which the desired pruned network can be picked based on sparsity and performance requirements. These features allow LTP to achieve competitive compression rates on ImageNet networks such as AlexNet ($26.4\times$ compression with $79.1\%$ Top-5 accuracy) and ResNet50 ($9.1\times$ compression with $92.0\%$ Top-5 accuracy). We also show that LTP effectively prunes modern \textit{compact} architectures, such as EfficientNet, MobileNetV2 and MixNet.

cs.LG

Distinguishing a Mott Insulator from a Trivial Insulator with Atomic Adsorbates

In an electronic system with various interactions intertwined, revealing the origin of its many-body ground state is challenging and a direct experimental way to verify the correlated nature of an insulator has been lacking. Here we demonstrate a way to unambiguously distinguish a paradigmatic correlated insulator, a Mott insulator, from a trivial band insulator based on their distinct chemical behavior for a surface adsorbate using 1T-TaS2, which has been debated between a spin-frustrated Mott insulator or a spin-singlet trivial insulator. We start from the observation of different sizes of spectral gaps on different surface terminations and show that potassium adatoms on these two surface layers behave in totally different ways. This can be straightforwardly understood from distinct properties of a Mott and a band insulators due to the fundamental difference of a half and a full-filled orbital involved respectively. This work not only solves an outstanding problem in this particularly interesting material but also provides a simple touchstone to identify the correlated ground state of electrons experimentally.

cond-mat.str-el

Honeycomb-Lattice Mott insulator on Tantalum Disulphide

Effects of electron many-body interactions amplify in an electronic system with a narrow bandwidth opening a way to exotic physics. A narrow band in a two-dimensional (2D) honeycomb lattice is particularly intriguing as combined with Dirac bands and topological properties but the material realization of a strongly interacting honeycomb lattice described by the Kane-Mele-Hubbard model has not been identified. Here we report a novel approach to realize a 2D honeycomb-lattice narrow-band system with strongly interacting 5$d$ electrons. We engineer a well-known triangular lattice 2D Mott insulator 1T-TaS$_2$ into a honeycomb lattice utilizing an adsorbate superstructure. Potassium (K) adatoms at an optimum coverage deplete one-third of the unpaired $d$ electrons and the remaining electrons form a honeycomb lattice with a very small hopping. Ab initio calculations show extremely narrow Z$_2$ topological bands mimicking the Kane-Mele model. Electron spectroscopy detects an order of magnitude bigger charge gap confirming the substantial electron correlation as confirmed by dynamical mean field theory. It could be the first artificial Mott insulator with a finite spin Chern number.

cond-mat.str-el

LSQ+: Improving low-bit quantization through learnable offsets and better initialization

Unlike ReLU, newer activation functions (like Swish, H-swish, Mish) that are frequently employed in popular efficient architectures can also result in negative activation values, with skewed positive and negative ranges. Typical learnable quantization schemes [PACT, LSQ] assume unsigned quantization for activations and quantize all negative activations to zero which leads to significant loss in performance. Naively using signed quantization to accommodate these negative values requires an extra sign bit which is expensive for low-bit (2-, 3-, 4-bit) quantization. To solve this problem, we propose LSQ+, a natural extension of LSQ, wherein we introduce a general asymmetric quantization scheme with trainable scale and offset parameters that can learn to accommodate the negative activations. Gradient-based learnable quantization schemes also commonly suffer from high instability or variance in the final training performance, hence requiring a great deal of hyper-parameter tuning to reach a satisfactory performance. LSQ+ alleviates this problem by using an MSE-based initialization scheme for the quantization parameters. We show that this initialization leads to significantly lower variance in final performance across multiple training runs. Overall, LSQ+ shows state-of-the-art results for EfficientNet and MixNet and also significantly outperforms LSQ for low-bit quantization of neural nets with Swish activations (e.g.: 1.8% gain with W4A4 quantization and upto 5.6% gain with W2A2 quantization of EfficientNet-B0 on ImageNet dataset). To the best of our knowledge, ours is the first work to quantize such architectures to extremely low bit-widths.

cs.CV

Ordering Chaos: Memory-Aware Scheduling of Irregularly Wired Neural Networks for Edge Devices

Recent advances demonstrate that irregularly wired neural networks from Neural Architecture Search (NAS) and Random Wiring can not only automate the design of deep neural networks but also emit models that outperform previous manual designs. These designs are especially effective while designing neural architectures under hard resource constraints (memory, MACs, . . . ) which highlights the importance of this class of designing neural networks. However, such a move creates complication in the previously streamlined pattern of execution. In fact one of the main challenges is that the order of such nodes in the neural network significantly effects the memory footprint of the intermediate activations. Current compilers do not schedule with regard to activation memory footprint that it significantly increases its peak compared to the optimum, rendering it not applicable for edge devices. To address this standing issue, we present a memory-aware compiler, dubbed SERENITY, that utilizes dynamic programming to find a sequence that finds a schedule with optimal memory footprint. Our solution also comprises of graph rewriting technique that allows further reduction beyond the optimum. As such, SERENITY achieves optimal peak memory, and the graph rewriting technique further improves this resulting in 1.68x improvement with dynamic programming-based scheduler and 1.86x with graph rewriting, against TensorFlow Lite with less than one minute overhead.

cs.DC

QKD: Quantization-aware Knowledge Distillation

Quantization and Knowledge distillation (KD) methods are widely used to reduce memory and power consumption of deep neural networks (DNNs), especially for resource-constrained edge devices. Although their combination is quite promising to meet these requirements, it may not work as desired. It is mainly because the regularization effect of KD further diminishes the already reduced representation power of a quantized model. To address this short-coming, we propose Quantization-aware Knowledge Distillation (QKD) wherein quantization and KD are care-fully coordinated in three phases. First, Self-studying (SS) phase fine-tunes a quantized low-precision student network without KD to obtain a good initialization. Second, Co-studying (CS) phase tries to train a teacher to make it more quantizaion-friendly and powerful than a fixed teacher. Finally, Tutoring (TU) phase transfers knowledge from the trained teacher to the student. We extensively evaluate our method on ImageNet and CIFAR-10/100 datasets and show an ablation study on networks with both standard and depthwise-separable convolutions. The proposed QKD outperformed existing state-of-the-art methods (e.g., 1.3% improvement on ResNet-18 with W4A4, 2.6% on MobileNetV2 with W4A4). Additionally, QKD could recover the full-precision accuracy at as low as W3A3 quantization on ResNet and W6A6 quantization on MobilenetV2.

cs.CV