SearcharxivSearch

arXiv subjects

Weisheng Zhao

Publications and source records attributed to Weisheng Zhao.

At least 19 recordsLinked to original sources

Unbiased first-principles construction of complete tensorial spin Hamiltonians

Magnetic ground states are commonly predicted using spin Hamiltonians whose interaction terms are selected a priori, potentially overlooking the microscopic interactions that govern complex magnetic order. Here, we introduce a general framework for the unbiased first-principles construction of symmetry-complete tensorial spin Hamiltonians and its automated implementation in AMATIS. The framework constructs the Hamiltonian directly from density-functional theory while rigorously enforcing quantum spin algebra and crystallographic symmetry. Applied to representative two-dimensional van der Waals magnets, the framework reproduces established magnetic interactions and uncovers hidden physics beyond conventional spin models, including chiral interactions that stabilize metastable skyrmions, higher-rank tensorial interactions that reconstruct the magnetic phase diagram and establish stabilizing competing multi-Q phases, and an emergent p-wave altermagnetic electronic structure. Our results demonstrate that unbiased tensorial Hamiltonian construction provides a predictive alternative to the conventional practice of manually selecting spin-model interactions, enabling first-principles discovery of unconventional magnetic phases.

cond-mat.mtrl-sci

Automated Generation of Commensurate Magnetic Structures based on Spin Space Groups and Graph Theory

Magnetic structures with symmetry constraints are candidates for energetically favorable configurations. Enumerating these structures is essential for identifying experimental observations and provides unbiased, linearly stable, and optimally sampled reference configurations for energy fitting when extracting spin interactions. We present SpinGraph, an automated workflow generating symmetry-distinct magnetic configurations. SpinGraph is not only compatible with the more general spin space groups, but also allows precise control of the prescribed single- or multi-Q states superposition. For finite groups, we enumerate compatible subgroups directly. For spin space groups with a continuous or special spin-only part, the finite component is enumerated first and then combined with the compatible exact spin-only constraints. These actions are translated into graph constraints in real or Fourier space. The symmetry of every generated structure is re-evaluated after construction. Finally, we integrate SpinGraph with the magnetic analysis code AMATIS to perform a thorough calculation for the spin interactions of the insulating monolayer CrI3. SpinGraph complements the last part of AMATIS and realizes the fully automated workflow for the spin Hamiltonian construction, which is essential for studying phase transitions, spin textures, magnons, and spin dynamics.

cond-mat.mtrl-sci

Chiral Phonons and Giant Anisotropic Photoresponse in Quasi-1D van der Waals Semiconductor ZrSnS3

Low-dimensional van der Waals semiconductors with reduced symmetry provide a unique platform for exploring anisotropic physical properties. The quasi-one-dimensional family MXQ$_3$ (M = Hf, Zr; X = Sn; Q = S, Se) exhibits notable structural anisotropy, where zigzag atomic chains influence optical phenomena such as birefringence. This study investigates anisotropic lattice dynamics in ZrSnS$_3$ using angle- and polarization-dependent Raman spectroscopy. Temperature-dependent measurements reveal anharmonic phonon behavior, indicating strong phonon-phonon coupling. Density functional theory calculations show good agreement with the experimentally observed Raman spectra, validating the microscopic description of the lattice dynamics. We also observe a helicity-dependent intensity and a reversal in phonon intensity between lower- and higher-frequency modes under circularly polarized light, which is characteristic of chiral phonons governed by the polarization of the Zr/Sn chains. Our first-principles analysis further shows that angular-momentum-like phonon textures can emerge away from the $Γ$-point near mode-hybridization and avoided-crossing regions, providing microscopic insight into the observed helicity-dependent Raman signatures. Furthermore, we fabricate an optoelectronic device from a thin ZrSnS$_3$ nanowire, demonstrating a photoresponsivity of 50~mA/W under 520~nm laser excitation (1~mW/cm$^2$). The device exhibits a pronounced, power-scalable anisotropic photoresponse with a clear preferred polarization direction. These results highlight the coupling mechanisms between polarization, lattice vibrations, and charge carriers in ZrSnS$_3$, establishing it as a promising material for polarization-sensitive optoelectronics and directional quantum transport.

cond-mat.mtrl-sci

HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference

The deployment of large language models (LLMs) presents significant challenges due to their enormous memory footprints, low arithmetic intensity, and stringent latency requirements, particularly during the autoregressive decoding stage. Traditional compute-centric accelerators, such as GPUs, suffer from severe resource underutilization and memory bandwidth bottlenecks in these memory-bound workloads. To overcome these fundamental limitations, we propose HPIM, the first memory-centric heterogeneous Processing-In-Memory (PIM) accelerator that integrates SRAM-PIM and HBM-PIM subsystems designed specifically for LLM inference. HPIM employs a software-hardware co-design approach that combines a specialized compiler framework with a heterogeneous hardware architecture. It intelligently partitions workloads based on their characteristics: latency-critical attention operations are mapped to the SRAM-PIM subsystem to exploit its ultra-low latency and high computational flexibility, while weight-intensive GEMV computations are assigned to the HBM-PIM subsystem to leverage its high internal bandwidth and large storage capacity. Furthermore, HPIM introduces a tightly coupled pipeline strategy across SRAM-PIM and HBM-PIM subsystems to maximize intra-token parallelism, thereby significantly mitigating the serial dependency of the autoregressive decoding stage. Comprehensive evaluations using a cycle-accurate simulator demonstrate that HPIM significantly outperforms state-of-the-art accelerators, achieving a peak speedup of up to 23.1x compared to the NVIDIA A100 GPU. Moreover, HPIM exhibits superior performance over contemporary PIM-based accelerators, highlighting its potential as a highly practical and scalable solution for accelerating large-scale LLM inference.

cs.AR

Momentum-Resolved Tunneling Modulation Induced Giant Multistate Resistance in Antiferroelectric Multiferroic Junction

Multiferroic tunnel junctions (MFTJs), integrating ferroelectric and ferromagnetic functionalities within a single nanoscale device, hold significant promise for non-volatile, multi-state memory and innovative computing paradigms. In conventional MFTJs, tunneling resistance modulation relies primarily on ferroelectric (FE) polarization switching, which alters interfacial electric fields and shifts the Fermi level of adjacent ferromagnetic electrodes. However, achieving high tunnelelectroresistance (TER) through this approach demands strong built-in electric fields, which simultaneously hinder FE polarization switching, creating an intrinsic trade-off between reliable data reading and efficient writing. Here, we propose a dual mechanism that combines antiferroelectric (AFE) phase-transition modulation of the evanescent decay states with interfacial spin filtering based on $Fe_3GaTe_2$/bilayer-$In_2Se_3$/$Fe_3GaTe_2$ heterostructure. Beyond altering the electrostatic potential as in AFE-FE switching, the transitions between head-type and tail-type AFE states preserve the centrosymmetric potential profile yet fundamentally modulate the momentum-resolved distribution of evanescent decay rates across the Brillouin zone. When integrated with perfect spin filtering at the $Fe_3GaTe_2$/$α$-$In_2Se_3$ interface, this mechanism yields a giant TER (~$7.6\times10^3\%$), over 4 times that of conventional FE-based MFTJs, and a TMR exceeding $6.8\times10^5\%$, enhanced by two orders of magnitude over typical MFTJs. These mechanisms resolve the performance trade-off in MFTJs, enabling six distinct non-volatile resistance states at room temperature.

cond-mat.mes-hall

Two-Dimensional Superconductivity at the CaZrO3/KTaO3 (001) Heterointerfaces

Two-dimensional superconductivity at KTaO3 (KTO) heterointerfaces has sparked intensive investigations since its discovery, yet whether the (001)-oriented KTO interface hosts superconductivity remains to be elucidated. Here, we provide unambiguous evidence of superconductivity in two-dimensional electron gases (2DEGs) at CaZrO3/KTO(001) heterointerfaces, with a superconducting transition TC up to ~0.25 K. Notably, TC increases linearly with carrier density nS over the range of 4.5*10^13~10.3*10^13 cm^-2. Furthermore, superconductivity exhibits a pronounced dependence on crystallographic orientation, with TC rising from 0.25 K for (001) to 1.04 K for (110) and 2.22 K for (111), underscoring the crucial role of interfacial symmetry in the CaZrO3/KTO system. The two-dimensional nature of the superconducting state is corroborated by the Berezinskii-Kosterlitz-Thouless (BKT) transition and the large anisotropy of the upper critical field. For the CaZrO3/KTO(001) sample with nS=7.7*10^13 cm^-2, the estimated Ginzburg-Landau coherence length ξGL=146.4 nm is larger than the superconducting layer thickness dSC=10.1 nm by a factor of ~14.5, confirming significant two-dimensional confinement of the CaZrO3/KTO(001) superconductor. In addition, we demonstrate that the two-dimensional superconductivity at the CaZrO3/KTO(001) interface can be effectively tuned by applying a back gate voltage. Our findings reveal the existence of two-dimensional superconductivity at CaZrO3/KTO(001), providing a new platform for exploring two-dimensional superconductivity at oxide interfaces.

cond-mat.supr-con

An Ultra-Low Power and Fast Ising Machine using Voltage-Controlled Magnetoresistive Random Access Memory

Physics-inspired computing paradigms, such as Ising machines, are emerging as promising hardware alternatives to traditional von Neumann architectures for tackling computationally intensive combinatorial optimization problems (COPs). While quantum, optical, and electronic devices have garnered significant attention for their potential in realizing Ising machines, their translation into practical systems for industry-relevant applications remains challenging, with each approach facing specific limitations in power consumption and speed. To address this challenge, we report the first chip-level spintronic Ising machine using voltage-controlled magnetoresistive random access memory. The core of our design leverages magnetic tunnel junctions (MTJs) driven by the voltage-controlled magnetic anisotropy effect to realize the probabilistic update of Ising spins through a new mechanism. It enables a latency below 1 ns and an energy consumption under 40 fJ per spin update, achieving a 1000-times improvement over previous current-driven MTJ-based implementations. We map two real-world COPs in electronic design automation-global routing and layer assignment-onto the Ising model and demonstrate high-quality results with an energy efficiency of 25000 solutions per second per watt. This outperforms state-of-the-art quantum and graphics processing units by six and seven orders of magnitude, respectively. These results establish voltage-controlled spintronics as a compelling route towards next-generation physics-inspired machine intelligence, offering a paradigm for ultra-low-power, high-speed, and scalable computation.

physics.app-ph

Raman scattering fingerprints of the charge density wave state in one-dimensional NbTe$_4$

Charge-density waves (CDWs) are ordered quantum states of conduction electrons accompanied by periodic lattice distortions. Raman scattering (RS) spectroscopy is therefore well suited for probing CDW-induced structural modulations. We investigate the CDW state in quasi-one-dimensional NbTe$_4$ using RS spectroscopy. At $T$=5~K, the resonantly enhanced Raman spectrum exhibits 25 phonon modes. Polarization-dependent measurements reveal a strong coupling between phonon-mode symmetry and crystallographic symmetry, with modes polarized parallel or perpendicular to the crystallographic $c$-axis, along which the one-dimensional structure is elongated. Temperature-dependent RS measurements identify a transition between commensurate and incommensurate CDW phases, accompanied by pronounced thermal hysteresis, with transition temperatures of approximately 45~K upon cooling and 90~K upon warming. The hysteresis width depends on the warming rate, indicating a finite nucleation rate of CDW domains and suggesting potential relevance for memory-device applications.

cond-mat.mtrl-sci

Absence of magnetic order in epitaxial RuO2 revealed by X-ray linear dichroism

Recently, the topic of altermagnetism has attracted tremendous attention and RuO2 have been demonstrated to be one of the most promising altermagnetic candidates. However, disputes still remain on the existence of magnetic order in RuO2. Here in this work, we employ X-ray linear dichroism (XLD), a widely utilized technique for characterizing antiferromagnets, in conjunction with photoemission electron microscopy and multiple scattering calculation to provide clear evidence of the absence of magnetic order in epitaxial RuO2 films. The observed XLD signal is nearly invariant with temperature and independent on cooling field direction, in stark contrast to the substantial magnetic order-related XLD signal predicted by multiple scattering calculation. This finding strongly suggests a nonmagnetic origin for RuO2. Furthermore, we observed significantly distinct XLD signals at the Ru M3 and O K edges in RuO2 films grown on TiO2 substrate with different surface orientations, which can be attributed to the low-symmetry crystal field. These results unequivocally demonstrate the absence of magnetic order in RuO2 and establishes XLD measurement as a robust technique for probing the low-symmetry magnetic materials.

cond-mat.mtrl-sci

GCoDE: Efficient Device-Edge Co-Inference for GNNs via Architecture-Mapping Co-Search

Graph Neural Networks (GNNs) have emerged as the state-of-the-art graph learning method. However, achieving efficient GNN inference on edge devices poses significant challenges, limiting their application in real-world edge scenarios. This is due to the high computational cost of GNNs and limited hardware resources on edge devices, which prevent GNN inference from meeting real-time and energy requirements. As an emerging paradigm, device-edge co-inference shows potential for improving inference efficiency and reducing energy consumption on edge devices. Despite its potential, research on GNN device-edge co-inference remains scarce, and our findings show that traditional model partitioning methods are ineffective for GNNs. To address this, we propose GCoDE, the first automatic framework for GNN architecture-mapping Co-design and deployment on Device-Edge hierarchies. By abstracting the device communication process into an explicit operation, GCoDE fuses the architecture and mapping scheme in a unified design space for joint optimization. Additionally, GCoDE's system performance awareness enables effective evaluation of architecture efficiency across diverse heterogeneous systems. By analyzing the energy consumption of various GNN operations, GCoDE introduces an energy prediction method that improves energy assessment accuracy and identifies energy-efficient solutions. Using a constraint-based random search strategy, GCoDE identifies the optimal solution in 1.5 hours, balancing accuracy and efficiency. Moreover, the integrated co-inference engine in GCoDE enables efficient deployment and execution of GNN co-inference. Experimental results show that GCoDE can achieve up to 44.9x speedup and 98.2% energy reduction compared to existing approaches across diverse applications and system configurations.

cs.LG

TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices

Developing deep learning models on tiny devices (e.g. Microcontroller units, MCUs) has attracted much attention in various embedded IoT applications. However, it is challenging to efficiently design and deploy recent advanced models (e.g. transformers) on tiny devices due to their severe hardware resource constraints. In this work, we propose TinyFormer, a framework specifically designed to develop and deploy resource-efficient transformer models on MCUs. TinyFormer consists of SuperNAS, SparseNAS, and SparseEngine. Separately, SuperNAS aims to search for an appropriate supernet from a vast search space. SparseNAS evaluates the best sparse single-path transformer model from the identified supernet. Finally, SparseEngine efficiently deploys the searched sparse models onto MCUs. To the best of our knowledge, SparseEngine is the first deployment framework capable of performing inference of sparse transformer models on MCUs. Evaluation results on the CIFAR-10 dataset demonstrate that TinyFormer can design efficient transformers with an accuracy of 96.1% while adhering to hardware constraints of 1MB storage and 320KB memory. Additionally, TinyFormer achieves significant speedups in sparse inference, up to 12.2x comparing to the CMSIS-NN library. TinyFormer is believed to bring powerful transformers into TinyML scenarios and to greatly expand the scope of deep learning applications

cs.LG

CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures

Compute-in-memory (CIM) has emerged as a pivotal direction for accelerating workloads in the field of machine learning, such as Deep Neural Networks (DNNs). However, the effective exploitation of sparsity in CIM systems presents numerous challenges, due to the inherent limitations in their rigid array structures. Designing sparse DNN dataflows and developing efficient mapping strategies also become more complex when accounting for diverse sparsity patterns and the flexibility of a multi-macro CIM structure. Despite these complexities, there is still an absence of a unified systematic view and modeling approach for diverse sparse DNN workloads in CIM systems. In this paper, we propose CIMinus, a framework dedicated to cost modeling for sparse DNN workloads on CIM architectures. It provides an in-depth energy consumption analysis at the level of individual components and an assessment of the overall workload latency. We validate CIMinus against contemporary CIM architectures and demonstrate its applicability in two use-cases. These cases provide valuable insights into both the impact of sparsity patterns and the effectiveness of mapping strategies, bridging the gap between theoretical design and practical implementation.

cs.AR

Competition between Weak Localization and Antilocalization of Dirac-like Fermions in a Spin-Polarized Two-Dimensional Electron Gas at KTaO3 (111) Interface

Quantum transport phenomena in two-dimensional electron gases (2DEGs) at oxide interfaces have garnered significant interest owing to their potential in spintronic and quantum information technologies. Here, we systematically investigate the quantum conductance corrections of spin-polarized 2DEGs formed at the interfaces between two insulating oxides, ferromagnetic EuTiO3 (ETO) films and (111)-oriented KTaO3 (KTO) substrates. The anomalous Hall effect and hysteretic magnetoresistance provide clear evidence for long-range ferromagnetic order in the 2DEGs, which could be attributed to interfacial Eu doping in combination with the magnetic proximity effect of the ETO layer. The breaking of time-reversal symmetry by ferromagnetism in the 2DEGs, and with the assistance of spin-orbit coupling effect, gives rise to a nontrivial Berry phase. This results in a competition between weak localization (WL) and weak antilocalization (WAL) in the quantum transport of Dirac-like fermions at the KTO (111) interfaces. Notably, this competitive behavior can be effectively tuned by optical gating via a photoexcitation-induced shift of the Fermi level. Our findings demonstrate a controllable platform based on spin-polarized oxide 2DEGs for quantum transport, opening new avenues for spin-orbitronic and topological electronic applications.

cond-mat.str-el

An Event-Driven Spiking Compute-In-Memory Macro based on SOT-MRAM

The application of Magnetic Random-Access Memory (MRAM) in computing-in-memory (CIM) has gained significant attention. However, existing designs often suffer from high energy consumption due to their reliance on complex analog circuits for computation. In this work, we present a Spin-Orbit- Torque MRAM(SOT-MRAM)-based CIM macro that employs an event-driven spiking processing for high energy efficiency. The SOT-MRAM crossbar adopts a hybrid series-parallel cell structure to efficiently support matrix-vector multiplication (MVM). Signal information is (en) decoded as spikes using lightweight circuits, eliminating the need for conventional area- and powerintensive analog circuits. The SOT-MRAM macro is designed and evaluated in 28nm technology, and experimental results show that it achieves a peak energy efficiency of 243.6 TOPS/W, significantly outperforming existing designs.

cs.AR

Single femtosecond laser pulse-driven ferromagnetic switching

Light pulses offer a faster, more energy-efficient, and direct route to magnetic bit writing, pointing toward a hybrid memory and computing paradigm based on photon transmission and spin retention. Yet progress remains hindered, as deterministic, single-pulse optical toggle switching has so far been achieved only with ferrimagnetic materials, which require too specific a rare-earth composition and temperature conditions for technological use. In mainstream ferromagnet--central to spintronic memory and storage--such bistable switching is considered fundamentally difficult, as laser-induced heating does not inherently break time-reversal symmetry. Here, we report coherent magnetization switching in ferromagnets, driven by thermal anisotropy torque with single laser pulses. The toggle switching behavior is robust over a broad range of pulse durations, from femtoseconds to picoseconds, a prerequisite for practical applications. Furthermore, the phenomenon exhibits reproducibility in CoFeB/MgO-based magnetic tunnel junctions with a high magnetoresistance exceeding 110%, as well as the scalability down to nanoscales with remarkable energy efficiency (17 fJ per 100-nm-sized bit). These results mark a notable step toward integrating opto-spintronics into next-generation memory and storage technologies.

cond-mat.mes-hall

Field-Free Superconducting Diode Enabled by Geometric Asymmetry and Perpendicular Magnetization

The superconducting diode effect (SDE)- manifested as directional, dissipationless supercurrents - is pivotal for realizing energy-efficient superconducting logic and memory technologies. Achieving high-efficiency SDE without external magnetic fields, however, remains a fundamental challenge. Here, we report a strongly enhanced, field-free SDE in Pt/Co/Nb heterostructures, enabled by the interplay of engineered geometric asymmetry and stray fields from a perpendicularly magnetized Co layer. This configuration promotes directional vortex entry and spatially selective pinning, yielding diode efficiencies that exceed all previously reported field-free values. Temperature- and field-dependent transport measurements, supported by micromagnetic simulations, reveal that the enhanced nonreciprocity stems from three cooperative mechanisms: asymmetric vortex entry, localized magnetic pinning, and Lorentz-force imbalance. These findings establish a scalable, CMOS-compatible platform for high-performance superconducting rectifiers, offering new opportunities for cryogenic spintronics and quantum electronics.

cond-mat.supr-con

ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments

The device-edge co-inference paradigm effectively bridges the gap between the high resource demands of Graph Neural Networks (GNNs) and limited device resources, making it a promising solution for advancing edge GNN applications. Existing research enhances GNN co-inference by leveraging offline model splitting and pipeline parallelism (PP), which enables more efficient computation and resource utilization during inference. However, the performance of these static deployment methods is significantly affected by environmental dynamics such as network fluctuations and multi-device access, which remain unaddressed. We present ACE-GNN, the first Adaptive GNN Co-inference framework tailored for dynamic Edge environments, to boost system performance and stability. ACE-GNN achieves performance awareness for complex multi-device access edge systems via system-level abstraction and two novel prediction methods, enabling rapid runtime scheme optimization. Moreover, we introduce a data parallelism (DP) mechanism in the runtime optimization space, enabling adaptive scheduling between PP and DP to leverage their distinct advantages and maintain stable system performance. Also, an efficient batch inference strategy and specialized communication middleware are implemented to further improve performance. Extensive experiments across diverse applications and edge settings demonstrate that ACE-GNN achieves a speedup of up to 12.7x and an energy savings of 82.3% compared to GCoDE, as well as 11.7 better energy efficiency than Fograph.

cs.DC