SearcharxivSearch

arXiv subjects

Hongyi Guan

Publications and source records attributed to Hongyi Guan.

7 recordsLinked to original sources

TOM: A Ternary Read-only Memory Accelerator for LLM-powered Edge Intelligence

The deployment of Large Language Models (LLMs) for real-time intelligence on edge devices is rapidly growing. However, conventional hardware architectures face a fundamental memory wall challenge, where limited on-device memory capacity and bandwidth severely constrain the size of deployable models and their inference speed, while also limiting on-device adaptation. To address this challenge, we propose TOM, a hybrid ROM-SRAM accelerator co-designed with ternary quantization, which balances extreme density with on-device tunability. TOM exploits the synergy between ternary quantization and ROM to achieve extreme memory density and bandwidth, while preserving flexibility through a hybrid ROM-SRAM architecture designed for QLoRA-based tunability. Specifically, we introduce: (1) a sparsity-aware ROM architecture that synthesizes ternary weights as standard-cell logic, eliminating area overhead from zero-valued bits; (2) a distributed processing architecture that co-locates high-density ROM banks with flexible SRAM-based QLoRA adapters and compute units; and (3) a workload-aware dynamic power gating scheme that exploits the logic-based nature of ROM to power down inactive banks, minimizing dynamic energy consumption. TOM achieves an inference throughput of 3,306 TPS using BitNet-2B model, demonstrating its effectiveness in delivering real-time, energy-efficient edge intelligence.

cs.AR

MiMo-V2-Flash Technical Report

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-V2-Flash adopts a hybrid attention architecture that interleaves Sliding Window Attention (SWA) with global attention, with a 128-token sliding window under a 5:1 hybrid ratio. The model is pre-trained on 27 trillion tokens with Multi-Token Prediction (MTP), employing a native 32k context length and subsequently extended to 256k. To efficiently scale post-training compute, MiMo-V2-Flash introduces a novel Multi-Teacher On-Policy Distillation (MOPD) paradigm. In this framework, domain-specialized teachers (e.g., trained via large-scale reinforcement learning) provide dense and token-level reward, enabling the student model to perfectly master teacher expertise. MiMo-V2-Flash rivals top-tier open-weight models such as DeepSeek-V3.2 and Kimi-K2, despite using only 1/2 and 1/3 of their total parameters, respectively. During inference, by repurposing MTP as a draft model for speculative decoding, MiMo-V2-Flash achieves up to 3.6 acceptance length and 2.6x decoding speedup with three MTP layers. We open-source both the model weights and the three-layer MTP weights to foster open research and community collaboration.

cs.CL

CuPyMag: GPU-Accelerated Finite-Element Micromagnetics with Magnetostriction

We introduce CuPyMag, an open-source, Python-based framework for large-scale micromagnetic simulations with magnetostriction. CuPyMag solves micromagnetics with finite elements in a GPU-resident workflow in which key operations, such as right-hand-side assembly, spatial derivatives, and volume averages, are tensorized using CuPy's BLAS-accelerated backend. Benchmark tests show that the GPU solvers in CuPyMag achieve a speedup of up to two orders of magnitude compared to the CPU codes. Its runtime grows linearly/sublinearly with problem size, demonstrating high efficiency. Additionally, CuPyMag uses the Gauss-Seidel projection method for time integration, which not only allows stable time steps (up to 11 ps) but also solves each governing equation with only 1-3 conjugate-gradient iterations without preconditioning. CuPyMag accounts for magnetoelastic coupling and far-field effects arising from the boundary of the magnetic body, both of which play an important role in magnetization reversal in the presence of local defects. CuPyMag solves these computationally intensive multiphysics simulations with a high-resolution mesh (up to 3M nodes) in under three hours on an NVIDIA H200 GPU. This acceleration enables micromagnetic simulations with non-trivial defect geometries and resolves nanoscale magnetic structures. It expands the scope of micromagnetic simulations towards realistic, large-scale problems that can guide experiments. More broadly, CuPyMag is developed using widely adopted Python libraries, which provide cross-platform compatibility, ease of installation, and accessibility for adaptations to diverse applications.

physics.comp-ph

Efficient Architecture for RISC-V Vector Memory Access

Vector processors frequently suffer from inefficient memory accesses, particularly for strided and segment patterns. While coalescing strided accesses is a natural solution, effectively gathering or scattering elements at fixed strides remains challenging. Naive approaches rely on high-overhead crossbars that remap any byte between memory and registers, leading to physical design issues. Segment operations require row-column transpositions, typically handled using either element-level in-place transposition (degrading performance) or large buffer-based bulk transposition (incurring high area overhead). In this paper, we present EARTH, a novel vector memory access architecture designed to overcome these challenges through shifting-based optimizations. For strided accesses, EARTH integrates specialized shift networks for gathering and scattering elements. After coalescing multiple accesses within the same cache line, data is routed between memory and registers through the shifting network with minimal overhead. For segment operations, EARTH employs a shifted register bank enabling direct column-wise access, eliminating dedicated segment buffers while providing high-performance, in-place bulk transposition. Implemented on FPGA with Chisel HDL based on an open-source RISC-V vector unit, EARTH enhances performance for strided memory accesses, achieving 4x-8x speedups in benchmarks dominated by strided operations. Compared to conventional designs, EARTH reduces hardware area by 9% and power consumption by 41%, significantly advancing both performance and efficiency of vector processors.

cs.AR

Carbon in GaN as a nonradiative recombination center

Trap-assisted nonradiative recombination has been shown to limit the efficiency of optoelectronic devices. While substitutional carbon ($\mathrm{C_N}$) has been suggested to be a nonradiative recombination center in GaN devices, a complete recombination cycle including the two charge-state transition levels has not been previously described. In this work, we investigate the trap-assisted recombination process due to $\mathrm{C_N}$ in GaN, including multiphonon emission (MPE), radiative recombination, trap-assisted Auger-Meitner (TAAM) recombination, as well as thermal emission of holes. Our study shows the key role of TAAM processes at the high carrier densities relevant for devices. We also reveal the carrier-density regimes where thermal emission and radiative recombination are expected to play an observable role. Our results highlight that carbon concentrations exceeding $\sim$10$^{17}$ cm$^{-3}$ can have a noticeable impact on device efficiency, not just in GaN active layers but also in InGaN and AlGaN. Our comprehensive formalism not only offers detailed results for carbon but provides a general framework for assessing the multiple processes that participate in trap-assisted recombination in semiconductors.

cond-mat.mtrl-sci

Magnetoelastic Interactions Reduce Hysteresis in Soft Magnets

The width of the magnetic hysteresis loop is often correlated with the material's magnetocrystalline anisotropy constant $\kappa_1$. Traditionally, a common approach to reduce the hysteresis width has been to develop alloys with $\kappa_1$ as close to zero as possible. However, contrary to this widely accepted view, we present evidence that magnetoelastic interactions governed by magnetostriction constants, elastic stiffness, and applied stresses play an important role in reducing magnetic hysteresis width, despite large $\kappa_1$ values. We use a nonlinear micromagnetics framework to systematically investigate the interplay between material constants $\lambda_{100}$, $c_{11}$, $c_{12}$, $\kappa_1$, applied or residual stresses $\sigma_{\mathrm{R}}$, and needle domains to collectively lower the energy barrier for magnetization reversal. A distinguishing feature of our work is that we correlate the energy barrier governing the growth of needle domains with the width of the hysteresis loop. This energy barrier approach enables us to capture the nuanced interplay between anisotropy constant, magnetostriction, and applied stresses, and their combined influence on magnetic hysteresis. We propose a mathematical relationship on the coercivity map: $\kappa_1 = \alpha(c_{11}-c_{12})(\lambda_{100}+\beta\sigma_{11})^2$ for which magnetic hysteresis can be minimized for a uniaxial residual stress $\sigma_\mathrm{R} = \sigma_{11}\hat{\mathbf{e}}_1\otimes\hat{\mathbf{e}}_1$ (and for some constants $\alpha$, $\beta$). These results serve as quantitative guidelines to design magnetic alloys with small hysteresis, and potentially guide the discovery of a new generation of soft magnets located beyond the $\kappa_1 \to 0$ region.

cond-mat.mtrl-sci

Superconductivity of Light-Elements Doped H$ {}_{3} $S

Pressurized hydrogen-rich compounds, which could be viewed as precompressed metallic hydrogen, exhibit high superconductivity, thereby providing a viable route toward the discovery of high-temperature superconductors. Of particular interest is to search for high-temperature superconductors with low stable pressure in terms of pressure-stabilized hydrides. In this work, with the aim of obtaining high-temperature superconducting compounds at low pressure, we attempt to study the doping effects for high-temperature superconductive $ \mathrm{H_3S} $ with supercells up to 64 atoms using first principle electronic structure simulations. As a result of various doping, we found that Na doping for $ \mathrm{H_3S} $ could lower the dynamically stable pressure by 40 GPa. The results also indicate P doping could enhance the superconductivity of $ \mathrm{H_3S} $ system, which is in agreement with previous calculations. Moreover, our work proposed an approach that could reasonably estimate the superconducting critical temperature ($ T_{c} $) of a compound containing a large number of atoms, saving the computational cost significantly for large-scale elements-doping superconductivity simulations.

cond-mat.supr-con