SearcharxivSearch

arXiv subjects

Zheyuan Chen

Publications and source records attributed to Zheyuan Chen.

4 recordsLinked to original sources

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU

Running language models in the browser presents a unique opportunity to build efficient, private, and portable AI applications, but requires contending with constrained memory availability and heterogeneous hardware targets. To realize this opportunity, we present Llamas on the Web (LlamaWeb), a WebGPU backend for llama$.$cpp that enables memory-efficient and performance-portable LLM inference across a wide range of model weight formats in the browser. Our design significantly reduces memory overhead through static memory planning and efficient model loading, addresses cross-device variability through a tunable kernel library, and introduces templated GPU kernels that support performant implementations of numerous quantization formats, enabling broad model support and extensibility to new formats. We evaluate LlamaWeb on 16 devices from 8 vendors, collecting data from 10 language models and four model weight formats. We compare LlamaWeb against existing browser-based LLM frameworks and find that LlamaWeb requires 29-33% less memory across several combinations of device, browser, and operating system. We also evaluate LlamaWeb's performance against these frameworks and find that it increases decode throughput by 45-69% across four GPUs from separate vendors. In addition, we compare LlamaWeb's performance against other llama$.$cpp backends, where it is competitive with and even beats vendor-specific backend performance on some devices.

cs.DC

sqlelf: a SQL-centric Approach to ELF Analysis

The exploration and understanding of Executable and Linkable Format (ELF) objects underpin various critical activities in computer systems, from debugging to reverse engineering. Traditional UNIX tooling like readelf, nm, and objdump have served the community reliably over the years. However, as the complexity and scale of software projects has grown, there arises a need for more intuitive, flexible, and powerful methods to investigate ELF objects. In this paper, we introduce sqlelf, an innovative tool that empowers users to probe ELF objects through the expressive power of SQL. By modeling ELF objects as relational databases, sqlelf offers the following advantages over conventional methods. Our evaluations demonstrate that sqlelf not only provides more nuanced and comprehensive insights into ELF objects but also significantly reduces the effort and time traditionally required for ELF exploration tasks

cs.SE

Imaging Stacking Order in Few-Layer Graphene

Few-layer graphene (FLG) has been predicted to exist in various crystallographic stacking sequences, which can strongly influence the electronic properties of FLG. We demonstrate an accurate and efficient method to characterize stacking order in FLG using the distinctive features of the Raman 2D-mode. Raman imaging allows us to visualize directly the spatial distribution of Bernal (ABA) and rhombohedral (ABC) stacking in tri- and tetra-layer graphene. We find that 15% of exfoliated graphene tri- and tetra-layers is comprised of micron-sized domains of rhombohedral stacking, rather than of usual Bernal stacking. These domains are stable and remain unchanged for temperatures exceeding $800^{\circ}$C.

cond-mat.mtrl-sci

Energy Transfer from Individual Semiconductor Nanocrystals to Graphene

Energy transfer from photoexcited zero-dimensional systems to metallic systems plays a prominent role in modern day materials science. A situation of particular interest concerns the interaction between a photoexcited dipole and an atomically thin metal. The recent discovery of graphene layers permits investigation of this phenomenon. Here we report a study of fluorescence from individual CdSe/ZnS nanocrystals in contact with single- and few-layer graphene sheets. The rate of energy transfer is determined from the strong quenching of the nanocrystal fluorescence. For single-layer graphene, we find a rate of ~ 4ns-1, in agreement with a model based on the dipole approximation and a tight-binding description of graphene. This rate increases significantly with the number of graphene layers, before approaching the bulk limit. Our study quantifies energy transfer to and fluorescence quenching by graphene, critical properties for novel applications in photovoltaic devices and as a molecular ruler.

physics.chem-ph