SearcharxivSearch

arXiv subjects

Hanbo Sun

Publications and source records attributed to Hanbo Sun.

14 recordsLinked to original sources

Sliding ferroelectricity tunable conventional and anomalous spin Hall effects in bilayer 1T'-WTe2

The spin Hall effect, recognized for its high-speed, low-power, and highly controllable characteristics, is a key enabler for next-generation memory and logic devices. However, a primary challenge lies in achieving 180$^{\circ}$ magnetization switching without an external magnetic field in spin-orbit torque devices. Here, we propose a method to tune the conventional and anomalous spin Hall effects by the intrinsic sliding ferroelectricity. Importantly, the anomalous spin Hall effect can enable the field-free switching of perpendicular magnetization. We find a substantial anomalous spin Hall conductivity of $\sigma_{xy}^{y}$ = 45.62 ($\hbar$/e)S/cm and $\sigma_{yx}^{y}$ = 56.84 ($\hbar$/e)S/cm in monolayer 1T'-WTe$_2$. These values are significantly enhanced to $\sigma_{xy}^{y}$ = -96.77 ($\hbar$/e)S/cm and $\sigma_{yx}^{y}$ = 104.03 ($\hbar$/e)S/cm in the bilayer 1T'-WTe$_2$. More interestingly, the sliding ferroelectricity enables reversible switching of the signs and magnitudes for both the conventional and anomalous spin Hall conductivities. This originates from the fact that the sliding ferroelectric markedly shifts the relative spin Berry curvature contributions from the valence and conduction bands around the $\Gamma$-X path. Our findings not only reveal a strong coupling between sliding ferroelectricity and spin transport, but also propose a strategy for the nonvolatile electrical control of spintronic devices.

cond-mat.mtrl-sci

RTP-LLM: High-Performance Alibaba LLM Inference Engine

Large Language Models (LLMs) have revolutionized AI applications, but deploying them at scale presents significant challenges. We present RTP-LLM, a high-performance inference engine for industrial-scale LLM deployment, successfully deployed across Alibaba Group serving over 100 million users. RTP-LLM addresses fundamental bottlenecks through integrated design. It optimizes model loading via file-order-driven I/O and parallel I/O-communication overlapping. The Prefill-Decode Disaggregation architecture decouples compute-intensive prefill from memory-bound decode phases, combined with hierarchical multi-tiered KV cache management enabling efficient cache reuse. In addition, RTP-LLM incorporates modular speculative decoding supporting multiple algorithms, adaptive KV cache quantization, and decoupled multimodal processing, with support for multi-level parallelism. Comprehensive evaluations across diverse model architectures (8B-235B parameters) have been conducted, where both controlled benchmarks and real production workloads are used. The results demonstrate RTP-LLM's superior performance against vLLM and SGLang: 4.7x-6.3x model loading speedup, 35-37% TTFT P95 latency reduction with 215% cache reuse improvement in production traffic scheduling, 1.12x-2.48x and 1.86x-2.52x throughput improvements in speculative decoding and multimodal inference, respectively, and 35-40% batch latency reduction with 1.9x-3.0x TTFT improvement in quantized inference. RTP-LLM's production-proven architecture and open-source availability make it a comprehensive solution for industrial LLM deployment.

cs.OS

Sliding Ferroelectricity Induced and Switched Altermagnetism in GaSe-VPSe3-GaSe Sandwiched Heterostructure with Strong Magnetoelectric Effect

Magnetoelectric coupling is vital for exploring fundamental science and driving the development of high-density memory and energy-efficient spintronic devices. Altermagnets, which merge the benefits of ferromagnets and antiferromagnets, pave the way for unprecedented magnetoelectric coupling effects. However, the spin splitting in altermagnets is robustly protected by spin space group symmetry, posing a significant challenge for external manipulation. Here, we propose to utilize the coupling between the layer degree of freedom and the altermagnet to achieve an altermagnetic multiferroic with strong magnetoelectric coupling. In the GaSe-VPSe3-GaSe sandwiched structure, the magnetic order can be switched between altermagnetic and conventional antiferromagnetic by controllably breaking and restoring the combined spatial inversion and time-reversal symmetry using sliding ferroelectricity. Moreover, our systematic investigation of all pathways revealed that the transition from a ferroelectric CB stacking, through an antiferroelectric CC stacking, to a ferroelectric BC stacking is the most favorable, with an energy barrier of only 50.13 meV/f.u.. More importantly, we reveal that the microscopic mechanism of the magnetic phase transition stems from the interlayer covalent bonding of Se-Se or Se-P atomic pairs at the interface. Our findings unveil a new form of magnetoelectric coupling and lay the groundwork for designing miniature information processing and multiferroic memory devices based on altermagnetism.

cond-mat.mtrl-sci

Type III Valley Polarization and Anomalous Valley Hall Effect in Two-Dimensional Non-Janus and Janus Altermagnet Fe2WS2Se2

Exploiting the valley degree of freedom introduces a novel paradigm for advancing quantum information technology. Currently, the investigation on spontaneous valley polarization mainly focuses on two major types of systems. One type magnetic systems by breaking the time-reversal symmetry, the other is ferroelectric materials through breaking the inversion symmetry. Might there be additional scenarios? Here, we propose to realize spontaneous valley polarization by breaking the mirror symmetry in the altermagnets, named type III valley polarization. Through symmetry analysis and first-principles calculations, we confirm that this mechanism is feasible in Non-Janus Fe2WS2Se2. Monolayer Non-Janus and Janus Fe2WS2Se2 are stable Neel-type antiferromagnetic state with the direct band gap semiconductor. More interestingly, their magnetic anisotropy energy exhibits the rare biaxial anisotropy and a four-leaf clover shape in the xy plane, while the xz and yz planes show the common uniaxial anisotropy. This originated from the fourth-order single ion interactions. More importantly, the valley splitting is spontaneously generated in the Non-Janus Fe2WS2Se2 due to the Mxy symmetry breaking, without requiring the SOC effect. Both the Non-Janus and Janus Fe2WS2Se2 exhibit diverse valley polarization and anomalous valley Hall effect properties. In addition, the magnitude and direction of valley polarization can be effectively tuned by the biaxial strain and magnetic field. Our findings not only expand the realization system of spontaneous valley polarization, but also provide a theoretical basis for the high-density storage of valley degrees of freedom.

cond-mat.mtrl-sci

Valley Polarization and Anomalous Valley Hall Effect in Altermagnet Ti2Se2S with Multipiezo Properties

Recently, altermagnets demonstrate numerous newfangle physical phenomena due to their inherent antiferromagnetic coupling and spontaneous spin splitting, that are anticipated to enable innovative spintronic devices. However, the rare two-dimensional altermagnets have been reported, making it difficult to meet the requirements for high-performance spintronic devices on account of the growth big data. Here, we predict a stable monolayer Ti2Se2S with out-of-plane altermagnetic ground state and giant valley splitting. The electronic properties of altermagnet Ti2Se2S are highly dependent on the onsite electron correlation. Through symmetry analysis, we find that the valleys of X and Y points are protected by the mirror Mxy symmetry rather than the time-reversal symmetry. Therefore, the multipiezo effect, including piezovalley and piezomagnetism, can be induced by the uniaxial strain. The total valley splitting of monolayer Ti2Se2S can be as high as ~500 meV. Most interestingly, the direction of valley polarization can be effectively tuned by the uniaxial strain, based on this, we have defined logical "0", "+1", and "-1" states for data transmission and storage. In addition, we have designed a schematic diagram for observing the anomalous Hall effect in experimentally. Our findings have enriched the candidate materials of two-dimensional altermagnet for the ultra-fast and low power consumption device applications.

cond-mat.mtrl-sci

Multifield Induced Antiferromagnet Transformation into Altermagnet and Realized Anomalous Valley Hall Effect in Two-dimensional Materials

Altermagnetism, as a new category of collinear magnetism distinct from traditional ferromagnetism and antiferromagnetism, exhibits the spin splitting without net magnetization. Currently, researchers are focus on searching three-dimensional altermagnetism and exploring its novel physical properties. However, there is a lack of understanding of the physical origin of two-dimensional altermagnetic emergent behavior. Here, we propose an approach to realize the transition from Neel antiferromagnetism to altermagnetism in two-dimensional system using an electric field, Janus structure, and ferroelectric substrate. In monolayer VPSe3, we demonstrate that multiple-physical-fields cause the upper and lower Se atoms unequal to break PT symmetry, resulting in altermagnetic spin splitting. Noted that monolayer VPSe3 produces a spontaneous valley splitting of 2.91 meV at the conduction band minimum. The electric field can effectively tune the valley splitting magnitude, while the Janus structure not only changes the valley splitting magnitude, but also alters the direction. More interestingly, when the ferroelectric polarization of Al2S3 is upward, the direction of valley polarization is switched and the magnitude is almost unchanged. However, the valley splitting sigfinicantly increases under the downward. It is worth noting that the ferroelectric polarization can switch altermagnetic effect and realize anomalous valley Hall effect. Besides, we reveal the microscopic mechanism of valley splitting by an effective Hamiltonian. Our findings not only provide a method to designing altermagnet, but also enriches the valley physics.

cond-mat.mtrl-sci

Coexisting Triferroic and Multiple Types of Valley Polarization by Structural Phase Transition in Two-Dimensional Materials

The multiferroic materials, which coexist magnetism, ferroelectric, and ferrovalley, have broad practical application prospects in promoting the miniaturization and integration of spintronic and valleytronic devices. However, it is rare that there are triferroic orders and multiple types of valley polarization in a real material. Here, we propose a mechanism to realize triferroic order coexistence and multiple types of valley polarization by structural phase transition in two-dimensional (2D) materials. The 1T and 2H phase OsBr2 monolayers exhibit non-magnetic semiconductor and ferromagnetic semiconductor with valley polarization up to 175.49 meV, respectively. Interestingly, the 1T phase OsBr2 bilayer shows the tri-state valley polarization due to lattice symmetry breaking, while the valley polarization of 2H phase bilayer originates from the combined effect of time-reversal symmetry breaking and spin-orbit coupling. Furthermore, the valley polarization and ferroelectric polarization of 1T phase AB stackings and 2H phase AA stackings can be manipulated via interlayer sliding. Importantly, we have verified that the 2H phase can be transformed to 1T phase by Li+ ion intercalation, while the 2H phase can occur the structural phase transition into the 1T phase by infrared laser induction. Our work provides a feasible strategy for manipulating valley polarization and a design idea for nano-devices with nonvolatile multiferroic properties.

cond-mat.mtrl-sci

Multifield tunable valley splitting and anomalous valley Hall effect in two-dimensional antiferromagnetic MnBr

Compared to the ferromagnetic materials that realize the anomalous valley Hall effect by breaking time-reversal symmetry and spin-orbit coupling, the antiferromagnetic materials with the joint spatial inversion and time-reversal (PT) symmetry are rarely reported that achieve the anomalous valley Hall effect. Here, we predict that the antiferromagnetic monolayer MnBr possesses spontaneous valley polarization. The valley splitting of valence band maximum is 21.55 meV at K and K' points, which is originated from Mn-dx2-y2 orbital by analyzing the effective Hamiltonian. Importantly, monolayer MnBr has zero Berry curvature in the entire momentum space but non-zero spin-layer locked Berry curvature, which offers the condition for the anomalous valley Hall effect. In addition, the magnitude of valley splitting can be signally tuned by the onsite correlation, strain, magnetization rotation, electric field, and built-in electric field. The electric field and built-in electric field induce spin splitting due to breaking the P symmetry. Therefore, the spin-layer locked anomalous valley Hall effect can be observed in MnBr. More remarkably, the ferroelectric substrate Sc2CO2 can tune monolayer MnBr to realize the transition from metal to valley polarization semiconductor. Our findings not only extend the implementation of the anomalous valley Hall effect, but also provides a platform for designing low-power and non-volatile valleytronics devices.

cond-mat.mtrl-sci

Ferroelectric tuning of the valley polarized metal-semiconductor transition in Mn2P2S3Se3/Sc2CO2 van der Waals heterostructures and application to nonlinear Hall effect devices

In order to promote the development of the next generation of nano-spintronic devices, it is of great significance to tune the freedom of valley in two-dimensional (2D) materials. Here, we propose a mechanism for manipulating the valley and nonlinear Hall effect by the 2D ferroelectric substrate. The monolayer Mn2P2S3Se3 is a robust antiferromagnetic valley polarized semiconductor. Importantly, the valley polarized metal-semiconductor phase transition of Mn2P2S3Se3 can be effectively tuned by switching the ferroelectric polarization of Sc2CO2. We reveal the microscopic mechanism of phase transition, which origins from the charge transfer and band alignment. Additionally, we find that transformed polarization direction of Sc2CO2 flexibly manipulate the Berry curvature dipole. Based on this discovery, we present the detection valley polarized metal-semiconductor transition by the nonlinear Hall effect devices. These findings not only offer a scheme to tune the valley degree of freedom, but also provide promising platform to design the nonlinear Hall effect devices.

cond-mat.mtrl-sci

Coexisting Magnetism, Ferroelectric, and Ferrovalley Multiferroic in Stacking-Dependent Two-Dimensional Materials

The two-dimensional (2D) multiferroic materials have widespread of application prospects in facilitating the integration and miniaturization of nanodevices. However, it is rarely coupling between the magnetic, ferroelectric, and ferrovalley in one 2D material. Here, we propose a mechanism for manipulating magnetism, ferroelectric, and valley polarization by interlayer sliding in 2D bilayer material. Monolayer GdI2 exhibits a ferromagnetic semiconductor with the valley polarization up to 155.5 meV. More interestingly, the magnetism and valley polarization of bilayer GdI2 can be strongly coupled by sliding ferroelectricity, appearing these tunable and reversible. In addition, we uncover the microscopic mechanism of magnetic phase transition by spin Hamiltonian and electron hopping between layers. Our findings offer a new direction for investigating 2D multiferroic in the implication for next-generation electronic, valleytronic, and spintronic devices.

cond-mat.mtrl-sci

FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs

Transformer-based Large Language Models (LLMs) have made a significant impact on various domains. However, LLMs' efficiency suffers from both heavy computation and memory overheads. Compression techniques like sparsification and quantization are commonly used to mitigate the gap between LLM's computation/memory overheads and hardware capacity. However, existing GPU and transformer-based accelerators cannot efficiently process compressed LLMs, due to the following unresolved challenges: low computational efficiency, underutilized memory bandwidth, and large compilation overheads. This paper proposes FlightLLM, enabling efficient LLMs inference with a complete mapping flow on FPGAs. In FlightLLM, we highlight an innovative solution that the computation and memory overhead of LLMs can be solved by utilizing FPGA-specific resources (e.g., DSP48 and heterogeneous memory hierarchy). We propose a configurable sparse DSP chain to support different sparsity patterns with high computation efficiency. Second, we propose an always-on-chip decode scheme to boost memory bandwidth with mixed-precision support. Finally, to make FlightLLM available for real-world LLMs, we propose a length adaptive compilation method to reduce the compilation overhead. Implemented on the Xilinx Alveo U280 FPGA, FlightLLM achieves 6.0$\times$ higher energy efficiency and 1.8$\times$ better cost efficiency against commercial GPUs (e.g., NVIDIA V100S) on modern LLMs (e.g., LLaMA2-7B) using vLLM and SmoothQuant under the batch size of one. FlightLLM beats NVIDIA A100 GPU with 1.2$\times$ higher throughput using the latest Versal VHK158 FPGA.

cs.AR

Human Transcription Quality Improvement

High quality transcription data is crucial for training automatic speech recognition (ASR) systems. However, the existing industry-level data collection pipelines are expensive to researchers, while the quality of crowdsourced transcription is low. In this paper, we propose a reliable method to collect speech transcriptions. We introduce two mechanisms to improve transcription quality: confidence estimation based reprocessing at labeling stage, and automatic word error correction at post-labeling stage. We collect and release LibriCrowd - a large-scale crowdsourced dataset of audio transcriptions on 100 hours of English speech. Experiment shows the Transcription WER is reduced by over 50%. We further investigate the impact of transcription error on ASR model performance and found a strong correlation. The transcription quality improvement provides over 10% relative WER reduction for ASR models. We release the dataset and code to benefit the research community.

cs.CL

HTEC: Human Transcription Error Correction

High-quality human transcription is essential for training and improving Automatic Speech Recognition (ASR) models. Recent study~\cite{libricrowd} has found that every 1% worse transcription Word Error Rate (WER) increases approximately 2% ASR WER by using the transcriptions to train ASR models. Transcription errors are inevitable for even highly-trained annotators. However, few studies have explored human transcription correction. Error correction methods for other problems, such as ASR error correction and grammatical error correction, do not perform sufficiently for this problem. Therefore, we propose HTEC for Human Transcription Error Correction. HTEC consists of two stages: Trans-Checker, an error detection model that predicts and masks erroneous words, and Trans-Filler, a sequence-to-sequence generative model that fills masked positions. We propose a holistic list of correction operations, including four novel operations handling deletion errors. We further propose a variant of embeddings that incorporates phoneme information into the input of the transformer. HTEC outperforms other methods by a large margin and surpasses human annotators by 2.2% to 4.5% in WER. Finally, we deployed HTEC to assist human annotators and showed HTEC is particularly effective as a co-pilot, which improves transcription quality by 15.1% without sacrificing transcription velocity.

eess.AS

Enabling Efficient and Flexible FPGA Virtualization for Deep Learning in the Cloud

FPGAs have shown great potential in providing low-latency and energy-efficient solutions for deep neural network (DNN) inference applications. Currently, the majority of FPGA-based DNN accelerators in the cloud run in a time-division multiplexing way for multiple users sharing a single FPGA, and require re-compilation with $\sim$100 s overhead. Such designs lead to poor isolation and heavy performance loss for multiple users, which are far away from providing efficient and flexible FPGA virtualization for neither public nor private cloud scenarios. To solve these problems, we introduce a novel virtualization framework for instruction architecture set (ISA) based on DNN accelerators by sharing a single FPGA. We enable the isolation by introducing a two-level instruction dispatch module and a multi-core based hardware resources pool. Such designs provide isolated and runtime-programmable hardware resources, further leading to performance isolation for multiple users. On the other hand, to overcome the heavy re-compilation overheads, we propose a tiling-based instruction frame package design and two-stage static-dynamic compilation. Only the light-weight runtime information is re-compiled with $\sim$1 ms overhead, thus the performance is guaranteed for the private cloud. Our extensive experimental results show that the proposed virtualization design achieves 1.07-1.69x and 1.88-3.12x throughput improvement over previous static designs using the single-core and the multi-core architectures, respectively.

cs.DC