SearcharxivSearch

arXiv subjects

Xiaotian Sun

Publications and source records attributed to Xiaotian Sun.

12 recordsLinked to original sources

FlashAccel: Leveraging High-Bandwidth Flash (HBF) for High-Throughput LLM Inference

Large language model (LLM) inference is increasingly limited by the capacity of High-Bandwidth Memory (HBM) in GPUs, as model weights and KV cache grow rapidly. High-Bandwidth Flash (HBF) provides higher capacity than HBM while offering comparable bandwidth, making it a promising substrate for capacity-constrained LLM inference. However, its inherently high access latency, low bandwidth utilization, and lack of support for heterogeneous resource management make it difficult to integrate HBF into GPUs for LLM inference. We present FlashAccel, a co-designed system that enables efficient LLM inference using HBF. FlashAccel integrates HBF into HBM-based GPUs, providing architectural support to mitigate access latency. It improves bandwidth utilization through specialized data layouts for both model weights and KV cache, and introduces an HBF-aware storage management layer together with a programming model to organize persistent data in HBF and coordinate heterogeneous memory resources at the system level. Experimental results demonstrate that integrating six HBF stacks into the GPU enables FlashAccel to deliver an average improvement of 2.49$\times$ and 1.93$\times$ in throughput per GPU and energy efficiency over the HBM-only GPU under a 100ms latency constraint, respectively.

cs.AR

Early Prediction of In-Hospital ICU Mortality Using Innovative First-Day Data: A Review

The intensive care unit (ICU) manages critically ill patients, many of whom face a high risk of mortality. Early and accurate prediction of in-hospital mortality within the first 24 hours of ICU admission is crucial for timely clinical interventions, resource optimization, and improved patient outcomes. Traditional scoring systems, while useful, often have limitations in predictive accuracy and adaptability. Objective: This review aims to systematically evaluate and benchmark innovative methodologies that leverage data available within the first day of ICU admission for predicting in-hospital mortality. We focus on advancements in machine learning, novel biomarker applications, and the integration of diverse data types.

cs.LG

Pressure-Driven Metallicity in Ångström-Thickness 2D Bismuth and Layer-Selective Ohmic Contact to MoS2

Recent fabrication of two-dimensional (2D) metallic bismuth (Bi) via van der Waals (vdW) squeezing method opens a new avenue to ultrascaling metallic materials into the ångström-thickness regime [Nature 639, 354 (2025)]. However, freestanding 2D Bi is typically known to exhibit a semiconducting phase [Nature 617, 67 (2023), Phys. Rev. Lett. 131, 236801 (2023)], which contradicts with the experimentally observed metallicity in vdW-squeezed 2D Bi. Here we show that such discrepancy originates from the pressure-induced buckled-to-flat structural transition in 2D Bi, which changes the electronic structure from semiconducting to metallic phases. Based on the experimentally fabricated MoS2-Bi-MoS2 trilayer heterostructure, we demonstrate the concept of layer-selective Ohmic contact in which one MoS2 layer forms Ohmic contact to the sandwiched Bi monolayer while the opposite MoS2 layer exhibits a Schottky barrier. The Ohmic contact can be switched between the two sandwiching MoS2 monolayers by changing the polarity of an external gate field, thus enabling charge to be spatially injected into different MoS2 layers. The layer-selective Ohmic contact proposed here represents a layertronic generalization of metal/semiconductor contact, paving a way towards layertronic device application.

cond-mat.mtrl-sci

PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators

Various processing-in-memory (PIM) accelerators based on various devices, micro-architectures, and interfaces have been proposed to accelerate deep neural networks (DNNs). How to deploy DNNs onto PIM-based accelerators is the key to explore PIM's high performance and energy efficiency. The scale of DNN models, the diversity of PIM accelerators, and the complexity of deployment are far beyond the human deployment capability. Hence, an automatic deployment methodology is indispensable. In this work, we propose PIMCOMP, an end-to-end DNN compiler tailored for PIM accelerators, achieving efficient deployment of DNN models on PIM hardware. PIMCOMP can adapt to various PIM architectures by using an abstract configurable PIM accelerator template with a set of pseudo-instructions, which is a high-level abstraction of the hardware's fundamental functionalities. Through a generic multi-level optimization framework, PIMCOMP realizes an end-to-end conversion from a high-level DNN description to pseudo-instructions, which can be further converted to specific hardware intrinsics/primitives. The compilation addresses two critical issues in PIM-accelerated inference from a system perspective: resource utilization and dataflow scheduling. PIMCOMP adopts a flexible unfolding format to reshape and partition convolutional layers, adopts a weight-layout guided computation-storage-mapping approach to enhance resource utilization, and balances the system's computation, memory access, and communication characteristics. For dataflow scheduling, we design two scheduling algorithms with different inter-layer pipeline granularities to support varying application scenarios while ensuring high computational parallelism. Experiments demonstrate that PIMCOMP improves throughput, latency, and energy efficiency across various architectures. PIMCOMP is open-sourced at \url{https://github.com/sunxt99/PIMCOMP-NN}.

cs.AR

Bilayer TeO2: The First Oxide Semiconductor with Symmetric Sub-5-nm NMOS and PMOS

Wide bandgap oxide semiconductors are very promising channel candidates for next-generation electronics due to their large-area manufacturing, high-quality dielectrics, low contact resistance, and low leakage current. However, the absence of ultra-short gate length (Lg) p-type transistors has restricted their application in future complementary metal-oxide-semiconductor (CMOS) integration. Inspired by the successfully grown high-hole mobility bilayer (BL) beta tellurium dioxide (\b{eta}-TeO2), we investigate the performance of sub-5-nm-Lg BL \b{eta}-TeO2 field-effect transistors (FETs) by utilizing first-principles quantum transport simulation. The distinctive anisotropy of BL \b{eta}-TeO2 yields different transport properties. In the y-direction, both the sub-5-nm-Lg n-type and p-type BL \b{eta}-TeO2 FETs can fulfill the International Technology Roadmap for Semiconductors (ITRS) criteria for high-performance (HP) devices, which are superior to the reported oxide FETs (only n-type). Remarkably, we for the first time demonstrate the existence of the NMOS and PMOS symmetry in sub-5-nm-Lg oxide semiconductor FETs. As to the x-direction, the n-type BL \b{eta}-TeO2 FETs satisfy both the ITRS HP and low-power (LP) requirements with Lg down to 3 nm. Consequently, our work shed light on the tremendous prospects of BL \b{eta}-TeO2 for CMOS application.

cond-mat.mes-hall

Quantum Transport Simulation of Sub-1-nm Gate Length Monolayer MoS2 Transistors

Sub-1-nm gate length $MoS_2$ transistors have been experimentally fabricated, but their device performance limit remains elusive. Herein, we explore the performance limits of the sub-1-nm gate length monolayer (ML) $MoS_2$ transistors through ab initio quantum transport simulations. Our simulation results demonstrate that, through appropriate doping and dielectric engineering, the sub-1-nm devices can meet the requirement of extended 'ITRS'(International Technology Roadmap for Semiconductors) $L_g$=0.34 nm. Following device optimization, we achieve impressive maximum on-state current densities of 409 $μA / μm$ for n-type and 800 $μA / μm$ for p-type high-performance (HP) devices, while n-type and p-type low-power (LP) devices exhibit maximum on-state current densities of 75 $μA / μm$ and 187 $μA / μm$, respectively. We employed the Wentzel-Kramer-Brillouin (WKB) approximation to explain the physical mechanisms of underlap and spacer region optimization on transistor performance. The underlap and spacer regions primarily influence the transport properties of sub-1-nm transistors by respectively altering the width and body factor of the potential barriers. Compared to ML $MoS_2$ transistors with a 1 nm gate length, our sub-1-nm gate length HP and LP ML $MoS_2$ transistors exhibit lower energy-delay products. Hence the sub-1-nm gate length transistors have immense potential for driving the next generation of electronics.

physics.comp-ph

PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators

Processing-in-memory (PIM) has shown extraordinary potential in accelerating neural networks. To evaluate the performance of PIM accelerators, we present an ISA-based simulation framework including a dedicated ISA targeting neural networks running on PIM architectures, a compiler, and a cycleaccurate configurable simulator. Compared with prior works, this work decouples software algorithms and hardware architectures through the proposed ISA, providing a more convenient way to evaluate the effectiveness of software/hardware optimizations. The simulator adopts an event-driven simulation approach and has better support for hardware parallelism. The framework is open-sourced at https://github.com/wangxy-2000/pimsim-nn.

cs.AR

PIMSYN: Synthesizing Processing-in-memory CNN Accelerators

Processing-in-memory architectures have been regarded as a promising solution for CNN acceleration. Existing PIM accelerator designs rely heavily on the experience of experts and require significant manual design overhead. Manual design cannot effectively optimize and explore architecture implementations. In this work, we develop an automatic framework PIMSYN for synthesizing PIM-based CNN accelerators, which greatly facilitates architecture design and helps generate energyefficient accelerators. PIMSYN can automatically transform CNN applications into execution workflows and hardware construction of PIM accelerators. To systematically optimize the architecture, we embed an architectural exploration flow into the synthesis framework, providing a more comprehensive design space. Experiments demonstrate that PIMSYN improves the power efficiency by several times compared with existing works. PIMSYN can be obtained from https://github.com/lixixi-jook/PIMSYN-NN.

cs.AR

PIMCOMP: A Universal Compilation Framework for Crossbar-based PIM DNN Accelerators

Crossbar-based PIM DNN accelerators can provide massively parallel in-situ operations. A specifically designed compiler is important to achieve high performance for a wide variety of DNN workloads. However, some key compilation issues such as parallelism considerations, weight replication selection, and array mapping methods have not been solved. In this work, we propose PIMCOMP - a universal compilation framework for NVM crossbar-based PIM DNN accelerators. PIMCOMP is built on an abstract PIM accelerator architecture, which is compatible with the widely used Crossbar/IMA/Tile/Chip hierarchy. On this basis, we propose four general compilation stages for crossbar-based PIM accelerators: node partitioning, weight replicating, core mapping, and dataflow scheduling. We design two compilation modes with different inter-layer pipeline granularities to support high-throughput and low-latency application scenarios, respectively. Our experimental results show that PIMCMOP yields improvements of 1.6$\times$ and 2.4$\times$ in throughput and latency, respectively, relative to PUMA.

cs.AR

Eshelby-twisted 3D moire superlattices

Twisted bilayers of van der Waals materials have recently attracted great attention due to their tunable strongly correlated phenomena. Here, we investigate the chirality-specific physics in 3D moiré superlattices induced by Eshelby twist. Our direct DFT calculations reveal helical rotation leads to optical circular dichroism, and chirality-specific nonlinear Hall effect, even though there is no magnetization or magnetic field. Both these phenomena can be reversed by changing the structural chirality. This provides a way to constructing chirality-specific materials.

cond-mat.mtrl-sci

Valley pseudospin in monolayer MoSi2N4 and MoSi2As4

For a long time, two-dimensional (2D) hexagonal MoS2 was proposed as a promising material for valleytronic system. However, the limited size of growth and low carrier motilities in MoS2 restrict its further application. Very recently, a new kind of hexagonal 2D MXene, MoSi2N4, was successfully synthesized with large size, excellent ambient stability, and considerable hole mobility. In this paper, based on the first-principles calculations, we predict that the valley-contrast properties can be realized in monolayer MoSi2N4 and its derivative MoSi2As4. Beyond the traditional two-level valleys, the valleys in monolayer MoSi2As4 are multiple-folded, implying a new valley dimension. Such multiple-folded valleys can be described by a three-band low-power Hamiltonian. This study presents the theoretical advance and the potential applications of monolayer MoSi2N4 and MoSi2As4 in valleytronic devices, especially multiple information processing.

cond-mat.mtrl-sci

Spontaneous Valley Splitting and Valley Pseudospin Field Effect Transistor of Monolayer VAgP2Se6

Valleytronics is a rising topic to explore the emergent degree of freedom for charge carriers in energy band edges and has attracted a great interest due to many intriguing quantum phenomena and potential application in information processing industry. Creation of permanent valley polarization, i.e. unbalanced occupation at different valleys, is a chief challenge and also urgent question to be solved in valleytronics. Here we predict that the spin-orbit coupling and magnetic ordering allow spontaneous valley Zeeman-type splitting in pristine monolayer of VAgP2Se6 by using first-principles calculations. The Zeeman-type valley splitting can lead to permanent valley polarization after suitable doping. The Zeeman-type valley splitting is similar to the role of spin polarization in spintronics and is a vital requirement for practical devices in valleytronics. The nonequivalent valleys of VAgP2Se6 monolayer can emit or absorb circularly polarized photons with opposite chirality, and thus this material shows a great potential to work as a photonic spin filter and circularly-polarized-light resource. A valley pseudospin field effect transistor (VPFET) is designed based on the monolayer VAgP2Se6 akin to the spin field effect transistors. Beyond common transistors, VPFETs carry information of not only the electrons but also the valley pseudospins.

cond-mat.mtrl-sci