SearcharxivSearch

arXiv subjects

Woojoo Lee

Publications and source records attributed to Woojoo Lee.

At least 19 recordsLinked to original sources

Signatures of a light-induced exciton condensate exhibiting BEC-BCS crossover

Exciton condensates provide a platform to study quasiparticle pairing, Bose-Einstein condensation-Bardeen-Cooper-Schrieffer (BEC-BCS) crossover, and excitonic topological phenomena. Achieving a nonequilibrium exciton condensate allows the ultimate tunability of these emergent phenomena. Yet, evidence of a light-induced, nonequilibrium exciton condensate and its BEC-BCS crossover remains elusive. Here, we use time- and angle-resolved photoemission spectroscopy to demonstrate signatures of a non-equilibrium exciton condensate and its BEC-BCS crossover in monolayer MnBi2Te4. Following optical excitation, a distinctive hole-like dispersion representing excitons emerges and persists for >20 ps. Strikingly, energy-domain sharpening in the valence band occurs 2 ps after time zero and exhibits a sharp onset at a threshold pump fluence of 0.84 mJ/cm2. The delayed and strongly nonlinear response is difficult to reconcile with transient field effects or conventional carrier-induced band shifts but is consistent with a model of exciton condensation governed by a Berezinskii-Kosterlitz-Thouless transition. The estimated threshold exciton density agrees quantitatively with the Nelson-Kosterlitz critical density. At higher fluences, the exciton feature develops a camel-back-shaped dispersion, consistent with the BEC-BCS crossover in the condensate framework. Our work establishes ultrathin MnBi2Te4 as a model system for studying nonequilibrium exciton condensates with a connection to superconductivity and exciton-driven topological phases.

cond-mat.mes-hall

Bayesian optimization approach for tracking a moving target from far-field data in three dimensions

We investigate a three-dimensional inverse scattering problem for tracking a rigidly moving target from far-field data generated by a single incident field. Extending our recent two-dimensional study, we develop a Bayesian optimization framework for simultaneously tracking the target's location and orientation over successive time steps, with the translational and rotational motions modeled as independent stochastic processes. We derive analytical formulas for the far-field pattern under translations and rotations and use them to design a Bayesian optimization procedure tailored to the tracking problem. We further establish posterior consistency for the underlying probabilistic model. When the target shape is unknown, its shape is identified at the initial time using a fully connected neural network trained on a precomputed dataset. Numerical experiments validate the effectiveness of the proposed framework.

math.AP

HSF-S: Speed-Optimized Compilation and Acceleration for Hybrid Schrodinger-Feynman Quantum Circuit Emulation

Hybrid Schrodinger-Feynman (HSF) simulation offers an attractive memory-path tradeoff for exact quantum-circuit emulation, but its practical runtime is often dominated by exponential path growth from cross-boundary two-qubit gates. Existing GPU and FPGA quantum simulators are largely optimized for full-state Schrodinger execution and therefore do not align well with HSF's path-centric workflow. This paper presents HSF-S, a compiler-accelerator co-designed framework for exact HSF-based quantum circuit emulation. HSF-S lowers input circuits to an HSF-compatible basis, formulates a rank-aware effective path-cost model, and applies dependency-preserving reordering together with discounted-gain SWAP insertion to suppress recurring cross-boundary interactions while preserving exact circuit semantics. A regression-free selector guarantees that the compiled circuit never increases effective path cost relative to the naive lowered baseline. We further design a dedicated HSF-S accelerator and execution flow, and integrate them into a stand-alone processor for efficient per-path dual-slice evaluation and final accumulation without materializing the full state vector. Across 56 benchmark circuits, HSF-S matches reference amplitudes to within floating-point precision, reduces effective path cost by up to 90.0%, and substantially improves practical tractability, including representative timeout-to-sub-second reductions under a 1-hour budget. On the resulting compiled workloads, the HSF-S processor prototype delivers up to 4.34x additional speedup.

quant-ph

SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling

Diffusion inference remains costly for edge deployment, yet existing accelerators focus almost exclusively on score networks because standard drift is merely a trivial linear scaling. Kuramoto orientation diffusion replaces this trivial drift with locally coupled phase interactions, improving sampling efficiency but introducing a new hardware bottleneck: a center-dependent nonlinear 5 x 5 stencil evaluated at every reverse step. This kernel maps poorly to conventional CNN accelerators and matrix-oriented engines. We present SA-Kura, to our knowledge the first digital systolic-array accelerator dedicated to locally coupled Kuramoto drift. By reformulating pair-wise sinusoidal coupling into neighbor accumulation independent of the center phase followed by a single center-dependent multiply-subtract combination, SA-Kura eliminates in-PE transcendental units and enables regular systolic execution with register-level reuse. SA-Kura was implemented in synthesizable RTL, integrated into a lightweight RISC-V-based SoC, prototyped on FPGA, and evaluated through 45 nm CMOS synthesis and power analysis. For the drift kernel only, compared with software execution of the same kernel on the processor core in the same SoC platform, SA-Kura reduces latency and energy by 193x and 69.4x, respectively. Compared with a standalone Jetson Orin Nano CUDA implementation of the same kernel, it is 6.57x faster and achieves approximately 46.0x lower energy per pixel.

cs.AR

CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization

Contrast maximization (CMAX) is a direct geometric framework for event-based motion estimation, but its iterative warp-and-accumulate pipeline incurs input-dependent computation and frequent memory accesses, challenging real-time, low-power edge deployment. We present CMAX-CAMEL, a coarse-to-fine adaptive, memory-efficient, low-power edge processor for CMAX. CMAX-CAMEL combines a runtime-adaptive execution strategy with a memory-centric processor architecture. It adjusts coarse-to-fine execution according to the observed event distribution, prioritizing stages likely to improve estimation accuracy while suppressing low-value iterations and unnecessary stage transitions. Architecturally, a banked parallel memory organization sustains real-time throughput while reducing latency, and a subsampling-coupled accumulation structure lowers memory-access activity along the warp-and-accumulate dataflow. On a Virtex FPGA prototype operating at 200 MHz, CMAX-CAMEL improves estimation accuracy by up to 19% over fixed coarse-to-fine schedules, reduces processing latency by 53.3%, lowers effective memory accesses by 42%, and cuts total system energy by 52.2%, including adaptation overheads. These results show that CMAX-CAMEL is an HW-SW co-design that co-optimizes execution policy and data movement for real-time, low-power event-based motion estimation at the edge.

cs.AR

Band Renormalization in Monolayer MoS2 Induced by Multipole Screening

Dielectric screening plays a crucial role in shaping the electronic structure of two-dimensional (2D) materials. In 2D semiconductors, screened Coulomb interactions arising from the surrounding dielectric environment are known to induce band renormalization, which is typically understood as a rigid shift of the electronic bands. Here, we experimentally demonstrate that dielectric screening can also give rise to non-rigid, momentum-dependent band renormalization. Using temperature-dependent angle-resolved photoemission spectroscopy (ARPES), we observe pronounced changes in the electronic band structure of monolayer MoS2 on a highly oriented pyrolytic graphite (HOPG) substrate. The results indicate that temperature-driven variations in the effective interlayer separation modulate the dielectric screening experienced by monolayer MoS2. At room temperature, the screening behavior is well described by a momentum-independent monopole approximation, whereas at liquid-helium temperatures the screening evolves into a multipole-like regime, leading to momentum-dependent band shifts.

cond-mat.str-el

Millimeter-Scale, Atomically Controlled 2D Topological Insulators Revealed by Multimodal Spectroscopy

Quantum spin Hall insulators, or synonymously known as 2D topological insulators, are crucial 2D systems hosting topologically protected edge states. The working temperature of this topological quantum phase is dictated by the inverted bandgap. However, the previously identified large-gap 2D topological insulators are either extremely chemically unstable, or cannot be made with atomistic precision over macroscopic scales. Here, we establish two-quintuple-layer Bi2Te3 and MnBi2Te4/Bi2Te3 heterostructures as atomically controlled, millimeter-scale 2D topological insulators, enabled by precision layer-by-layer growth that yields a carpet-like morphology extending coherently over macroscopic distances. This carpet-like growth mode renders the films amenable to mechanical exfoliation and subsequent wet or dry transfer. Multimodal spectroscopies and microscopies reveal the integer-layer tuned electronic structure of (Bi2Te3)n with excellent agreement to theory. Photon-energy-dependent photoemission and time-resolved photoemission identify band inversion and band dynamics, respectively, while scanning tunneling spectroscopy resolves topological edge states, characteristic of the 2D topological insulator phase. Thickness- and photon-energy-dependent photoemission further validates MnBi2Te4/Bi2Te3 as a robust 2D topological insulator. The large inverted gaps of ~100 meV in (Bi2Te3)2 and ~150 meV in MnBi2Te4/Bi2Te3 suggest operation near ambient temperature. These results define a scalable materials platform for next-generation, low-loss quantum and energy-efficient devices.

cond-mat.mtrl-sci

TT-SEAL: TTD-Aware Selective Encryption for Adversarially-Robust and Low-Latency Edge AI

Cloud-edge AI must jointly satisfy model compression and security under tight device budgets. While Tensor-Train Decomposition (TTD) shrinks on-device models, prior selective-encryption studies largely assume dense weights, leaving its practicality under TTD compression unclear. We present TT-SEAL, a selective-encryption framework for TT-decomposed networks. TT-SEAL ranks TT cores with a sensitivity-based importance metric, calibrates a one-time robustness threshold, and uses a value-DP optimizer to encrypt the minimum set of critical cores with AES. Under TTD-aware, transfer-based threat models (and on an FPGA-prototyped edge processor) TT-SEAL matches the robustness of full (black-box) encryption while encrypting as little as 4.89-15.92% of parameters across ResNet-18, MobileNetV2, and VGG-16, and drives the share of AES decryption in end-to-end latency to low single digits (e.g., 58% -> 2.76% on ResNet-18), enabling secure, low-latency edge AI.

cs.CR

TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI

The growing demands of distributed learning on resource constrained edge devices underscore the importance of efficient on device model compression. Tensor Train Decomposition (TTD) offers high compression ratios with minimal accuracy loss, yet repeated singular value decompositions (SVDs) and matrix multiplications can impose significant latency and energy costs on low power processors. In this work, we present TT-Edge, a hardware software co designed framework aimed at overcoming these challenges. By splitting SVD into two phases--bidiagonalization and diagonalization--TT-Edge offloads the most compute intensive tasks to a specialized TTD Engine. This engine integrates tightly with an existing GEMM accelerator, thereby curtailing the frequent matrix vector transfers that often undermine system performance and energy efficiency. Implemented on a RISC-V-based edge AI processor, TT-Edge achieves a 1.7x speedup compared to a GEMM only baseline when compressing a ResNet 32 model via TTD, while reducing overall energy usage by 40.2 percent. These gains come with only a 4 percent increase in total power and minimal hardware overhead, enabled by a lightweight design that reuses GEMM resources and employs a shared floating point unit. Our experimental results on both FPGA prototypes and post-synthesis power analysis at 45 nm demonstrate that TT-Edge effectively addresses the latency and energy bottlenecks of TTD based compression in edge environments.

cs.DC

FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AI

Machine unlearning, driven by privacy regulations and the "right to be forgotten", is increasingly needed at the edge, yet server-centric or retraining-heavy methods are impractical under tight computation and energy budgets. We present FiCABU (Fisher-based Context-Adaptive Balanced Unlearning), a software-hardware co-design that brings unlearning to edge AI processors. FiCABU combines (i) Context-Adaptive Unlearning, which begins edits from back-end layers and halts once the target forgetting is reached, with (ii) Balanced Dampening, which scales dampening strength by depth to preserve retain accuracy. These methods are realized in a full RTL design of a RISC-V edge AI processor that integrates two lightweight IPs for Fisher estimation and dampening into a GEMM-centric streaming pipeline, validated on an FPGA prototype and synthesized in 45 nm for power analysis. Across CIFAR-20 and PinsFaceRecognition with ResNet-18 and ViT, FiCABU achieves random-guess forget accuracy while matching the retraining-free Selective Synaptic Dampening (SSD) baseline on retain accuracy, reducing computation by up to 87.52 percent (ResNet-18) and 71.03 percent (ViT). On the INT8 hardware prototype, FiCABU further improves retain preservation and reduces energy to 6.48 percent (CIFAR-20) and 0.13 percent (PinsFaceRecognition) of the SSD baseline. In sum, FiCABU demonstrates that back-end-first, depth-aware unlearning can be made both practical and efficient for resource-constrained edge AI devices.

cs.LG

LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices

On-device fine-tuning of CNNs is essential to withstand domain shift in edge applications such as Human Activity Recognition (HAR), yet full fine-tuning is infeasible under strict memory, compute, and energy budgets. We present LoRA-Edge, a parameter-efficient fine-tuning (PEFT) method that builds on Low-Rank Adaptation (LoRA) with tensor-train assistance. LoRA-Edge (i) applies Tensor-Train Singular Value Decomposition (TT-SVD) to pre-trained convolutional layers, (ii) selectively updates only the output-side core with zero-initialization to keep the auxiliary path inactive at the start, and (iii) fuses the update back into dense kernels, leaving inference cost unchanged. This design preserves convolutional structure and reduces the number of trainable parameters by up to two orders of magnitude compared to full fine-tuning. Across diverse HAR datasets and CNN backbones, LoRA-Edge achieves accuracy within 4.7% of full fine-tuning while updating at most 1.49% of parameters, consistently outperforming prior parameter-efficient baselines under similar budgets. On a Jetson Orin Nano, TT-SVD initialization and selective-core training yield 1.4-3.8x faster convergence to target F1. LoRA-Edge thus makes structure-aligned, parameter-efficient on-device CNN adaptation practical for edge platforms.

cs.CV

ASAP-FE: Energy-Efficient Feature Extraction Enabling Multi-Channel Keyword Spotting on Edge Processors

Multi-channel keyword spotting (KWS) has become crucial for voice-based applications in edge environments. However, its substantial computational and energy requirements pose significant challenges. We introduce ASAP-FE (Agile Sparsity-Aware Parallelized-Feature Extractor), a hardware-oriented front-end designed to address these challenges. Our framework incorporates three key innovations: (1) Half-overlapped Infinite Impulse Response (IIR) Framing: This reduces redundant data by approximately 25% while maintaining essential phoneme transition cues. (2) Sparsity-aware Data Reduction: We exploit frame-level sparsity to achieve an additional 50% data reduction by combining frame skipping with stride-based filtering. (3) Dynamic Parallel Processing: We introduce a parameterizable filter cluster and a priority-based scheduling algorithm that allows parallel execution of IIR filtering tasks, reducing latency and optimizing energy efficiency. ASAP-FE is implemented with various filter cluster sizes on edge processors, with functionality verified on FPGA prototypes and designs synthesized at 45 nm. Experimental results using TC-ResNet8, DS-CNN, and KWT-1 demonstrate that ASAP-FE reduces the average workload by 62.73% while supporting real-time processing for up to 32 channels. Compared to a conventional fully overlapped baseline, ASAP-FE achieves less than a 1% accuracy drop (e.g., 96.22% vs. 97.13% for DS-CNN), which is well within acceptable limits for edge AI. By adjusting the number of filter modules, our design optimizes the trade-off between performance and energy, with 15 parallel filters providing optimal performance for up to 25 channels. Overall, ASAP-FE offers a practical and efficient solution for multi-channel KWS on energy-constrained edge devices.

eess.AS

Asymptotic Error Bounds and Fractional-Bit Design for Fixed-Point Grover's Quantum Algorithm Emulation

Quantum computing (QC) emulators, which simulate quantum algorithms on classical hardware, are indispensable platforms for testing quantum algorithms before scalable quantum computers become widely available. A critical challenge in QC emulation is managing numerical errors from finite arithmetic precision, especially truncation errors in resource-efficient fixed-point arithmetic. Despite its importance, systematic studies quantifying how truncation errors impact quantum algorithm accuracy are limited. In this paper, we propose a rigorous quantitative framework analyzing truncation error propagation in fixed-point QC emulation, focusing on Grover's quantum search algorithm. First, we introduce a simplified two-value amplitude representation of quantum states during Grover's iterations and prove its theoretical validity. Using this representation, we derive explicit mathematical expressions characterizing truncation error accumulation across quantum gate operations. We quantify the overall emulation error by the $\ell_2$ distance between ideal and emulated probability distributions, obtaining asymptotic bounds scaling as $O(2^{n-f})$, where $n$ is the number of qubits and $f$ is fractional-bit precision. Extensive numerical simulations and empirical experiments on a practical fixed-point QC emulator confirm that observed errors precisely match our theoretical predictions. Finally, we provide a closed-form formula to determine the minimal fractional-bit precision required to achieve a specified error threshold, offering clear guidelines for emulator designers balancing accuracy and resource utilization.

quant-ph

HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices

Processing-in-Memory (PIM) architectures offer promising solutions for efficiently handling AI applications in energy-constrained edge environments. While traditional PIM designs enhance performance and energy efficiency by reducing data movement between memory and processing units, they are limited in edge devices due to continuous power demands and the storage requirements of large neural network weights in SRAM and DRAM. Hybrid PIM architectures, incorporating non-volatile memories like MRAM and ReRAM, mitigate these limitations but struggle with a mismatch between fixed computing resources and dynamically changing inference workloads. To address these challenges, this study introduces a Heterogeneous-Hybrid PIM (HH-PIM) architecture, comprising high-performance MRAM-SRAM PIM modules and low-power MRAM-SRAM PIM modules. We further propose a data placement optimization algorithm that dynamically allocates data based on computational demand, maximizing energy efficiency. FPGA prototyping and power simulations with processors featuring HH-PIM and other PIM types demonstrate that the proposed HH-PIM achieves up to $60.43$ percent average energy savings over conventional PIMs while meeting application latency requirements. These results confirm the suitability of HH-PIM for adaptive, energy-efficient AI processing in edge devices.

cs.AR

Standalone FPGA-Based QAOA Emulator for Weighted-MaxCut on Embedded Devices

Quantum computing QC emulation is crucial for advancing QC applications, especially given the scalability constraints of current devices. FPGA-based designs offer an efficient and scalable alternative to traditional large-scale platforms, but most are tightly integrated with high-performance systems, limiting their use in mobile and edge environments. This study introduces a compact, standalone FPGA-based QC emulator designed for embedded systems, leveraging the Quantum Approximate Optimization Algorithm (QAOA) to solve the Weighted-MaxCut problem. By restructuring QAOA operations for hardware compatibility, the proposed design reduces time complexity from O(N^2) to O(N), where N equals 2^n for n qubits. This reduction, coupled with a pipeline architecture, significantly minimizes resource consumption, enabling support for up to nine qubits on mid-tier FPGAs, roughly three times more than comparable designs. Additionally, the emulator achieved energy savings ranging from 1.53 times for two-qubit configurations to up to 852 times for nine-qubit configurations, compared to software-based QAOA on embedded processors. These results highlight the practical scalability and resource efficiency of the proposed design, providing a robust foundation for QC emulation in resource-constrained edge devices.

cs.ET

Bayesian optimization approach for tracking the location and orientation of a moving target using far-field data

We investigate the inverse scattering problem for tracking the location and orientation of a moving scatterer using a single incident field. We solve the problem by adopting the optimization approach with the objective function defined by the discrepancy in far-field data. We rigorously derive formulas for the far-field data under translation and rotation of the target and prove that the objective function is locally Lipschitz with respect to the orientation angle at the true angle. By integrating these formulas with the Bayesian optimization approach, we reduce the cost of objective function evaluations. For the instance of an unknown target, machine learning via fully connected neural networks is applied to identify the shape of the target. Numerical simulations for randomly generated shapes and trajectories demonstrate the effectiveness of the proposed method.

math.AP

Approximation of the first Steklov-Dirichlet eigenvalue on eccentric spherical shells in general dimensions

We study the first Steklov-Dirichlet eigenvalue on eccentric spherical shells in $\mathbb{R}^{n+2}$ with $n\geq 1$, imposing the Steklov condition on the outer boundary sphere, denoted by $\Gamma_S$, and the Dirichlet condition on the inner boundary sphere. The first eigenfunction admits a Fourier--Gegenbauer series expansion via the bispherical coordinates, where the Dirichlet-to-Neumann operator on $\Gamma_S$ can be recursively expressed in terms of the expansion coefficients arXiv:2309.09587. In this paper, we develop a finite section approach for the Dirichlet-to-Neumann operator to approximate the first Steklov--Dirichlet eigenvalue on eccentric spherical shells. We prove the exponential convergence of this approach by using the variational characterization of the first eigenvalue. Furthermore, based on the convergence result, we propose a numerical computation scheme as an extension of the two-dimensional result in [Hong et al., Ann. Mat. Pura Appl., 2022] to general dimensions. We provide numerical examples of the first Steklov-Dirichlet eigenvalue on eccentric spherical shells with various geometric configurations.

math.AP

Distinguishing Surface and Bulk Electromagnetism via Their Dynamics in an Intrinsic Magnetic Topological Insulator

The indirect exchange interaction between local magnetic moments via surface electrons has been long predicted to bolster the surface ferromagnetism in magnetic topological insulators (MTIs), which facilitates the quantum anomalous Hall effect. This unconventional effect is critical to determining the operating temperatures of future topotronic devices. However, the experimental confirmation of this mechanism remains elusive, especially in intrinsic MTIs. Here we combine time-resolved photoemission spectroscopy with time-resolved magneto-optical Kerr effect measurements to elucidate the unique electromagnetism at the surface of an intrinsic MTI MnBi2Te4. Theoretical modeling based on 2D Ruderman-Kittel-Kasuya-Yosida interactions captures the initial quenching of a surface-rooted exchange gap within a factor of two but over-estimates the bulk demagnetization by one order of magnitude. This mechanism directly explains the sizable gap in the quasi-2D electronic state and the nonzero residual magnetization in even-layer MnBi2Te4. Furthermore, it leads to efficient light-induced demagnetization comparable to state-of-the-art magnetophotonic crystals, promising an effective manipulation of magnetism and topological orders for future topotronics.

cond-mat.str-el