SearcharxivSearch

arXiv subjects

Mircea Stan

Publications and source records attributed to Mircea Stan.

5 recordsLinked to original sources

FlexPosit: Tunable Fractional Precision for LLM Inference Accelerators

Large language models (LLMs) offer remarkable capabilities but impose prohibitive compute and energy costs. Quantization governs the trade-offs between accuracy and hardware efficiency across granularity and bit-width. Finer granularity (e.g., group-wise) provides high accuracy but incurs scaling and control overhead, while coarser granularity (e.g., channel-wise) has lower overhead but loses accuracy at low precision. Meanwhile, mixed-precision quantization exposes rich accuracy-efficiency trade-offs algorithmically, but existing LLM accelerators remain limited to discrete precision modes, leaving the fractional design space between them unexplored. FlexPosit bridges these gaps through co-design of Posit-based quantization and a precision-tunable bit-serial architecture. Algorithmically, FlexPosit employs distribution-aware quantization with hardware-aligned, sensitivity-guided mixed-precision allocation, leveraging the Posit format's tapered precision to achieve group-wise-like accuracy with channel-wise-like regularity. Architecturally, FlexPosit is a unified bit-serial systolic array with lightweight per-column decoders, unified Processing Elements (PEs), and a global precision controller, enabling tunable fractional precision while preserving fully regular systolic dataflow. Across diverse LLMs, FlexPosit achieves near-FP16 accuracy with sub-5-bit fractional weights. It achieves up to 1.8x higher throughput and 1.2x lower energy than BitMoD (group-wise quantization), and 1.5x higher throughput and 2.0x lower energy than OliVe (channel-wise quantization), establishing a new Pareto frontier for precision-tunable LLM acceleration.

cs.AR

Hot-LEGO: Architect Microfluidic Cooling Equipped 3DICs with Pre-RTL Thermal Simulation

Microfluidic cooling has been recognized as one of the most promising solutions to achieve efficient thermal management for three-dimensional integrated circuits (3DICs). It enables more opportunities to architect 3DICs with different die configurations. It becomes increasingly important to perform thermal analysis in the early design phases to validate the architectural design decisions. This is even more critical for microfluidic cooling equipped 3DICs as the embedded cooling structures greatly influence the performance, power, and reliability of the stacked system. We exploited the existing architectural simulators and developed a Pre-register-transfer-level (Pre-RTL) thermal simulation methodology named Hot-LEGO that integrates these tools with their latest features such as support for microfluidic cooling and 3DIC stacking configurations. This methodology differs from existing ones by looking into the design granularity at a much finer level which enables the exploration of unique architecture combinations across the vertical stack. Though architectural-level simulators are not designed for signoff-calibre, it offers speed and agility which are imperative for early design space exploration. We claim that this ongoing work will speed up the co-design cycle of microfluidic cooling and offer a portable methodology for architects to perform exhaustive search for the optimal microarchitecture solutions in 3DICs.

cs.AR

Computing and Memory Technologies based on Magnetic Skyrmions

Solitonic magnetic excitations such as domain walls and, specifically, skyrmionics enable the possibility of compact, high density, ultrafast,all-electronic, low-energy devices, which is the basis for the emerging area of skyrmionics. The topological winding of skyrmion spins affects their overall lifetime, energetics and dynamical behavior. In this review, we discuss skyrmionics in the context of the present day solid state memory landscape, and show how their size, stability and mobility can be controlled by material engineering, as well as how they can be nucleated and detected. Ferrimagnetsnear their compensation points are important candidates for this application, leading to detailed exploration of amorphous CoGd as well as the study of emergent materials such as Mn$_4$N and Inverse Heusler alloys. Along with material properties, geometrical parameters such as film thickness, defect density and notches can be used to tune skyrmion properties, such as their size and stability. Topology, however, can be a double-edged sword, especially for isolated metastable skyrmions, as it brings stability at the cost of additional damping and deflective Magnus forces compared to domain walls. Skyrmion deformation in response to forces also makes them intrinsically slower than domain walls. We explore potential analog applications of skyrmions, including temporal memory at low density, and decorrelator for stochastic computing at a higher density that capitalizes on their interactions. We summarize the main challenges to achieve a skyrmionics technology, including maintaining positional stability with very high accuracy, electrical readout, especially for small ferrimagnetic skyrmions, deterministic nucleation and annihilation, and overall integration with digital circuits with the associated circuit overhead.

cond-mat.mtrl-sci

Temporal Memory with Magnetic Racetracks

Race logic is a relative timing code that represents information in a wavefront of digital edges on a set of wires in order to accelerate dynamic programming and machine learning algorithms. Skyrmions, bubbles, and domain walls are mobile magnetic configurations (solitons) with applications for Boolean data storage. We propose to use current-induced displacement of these solitons on magnetic racetracks as a native temporal memory for race logic computing. Locally synchronized racetracks can spatially store relative timings of digital edges and provide non-destructive read-out. The linear kinematics of skyrmion motion, the tunability and low-voltage asynchronous operation of the proposed device, and the elimination of any need for constant skyrmion nucleation make these magnetic racetracks a natural memory for low-power, high-throughput race logic applications.

physics.app-ph

Spin-torque switching in large size nano-magnet with perpendicular magnetic fields

DC current induced magnetization reversal and magnetization oscillation was observed in 500 nm large size Co90Fe10/Cu/Ni80Fe20 pillars. A perpendicular external field enhanced the coercive field separation between the reference layer (Co90Fe10) and free layer (Ni80Fe20) in the pseudo spin valve, allowing a large window of external magnetic field for exploring the free-layer reversal. The magnetization precession was manifested in terms of the multiple peaks on the differential resistance curves. Depending on the bias current and applied field, the regions of magnetic switching and magnetization precession on a dynamical stability diagram has been discussed in details. Micromagnetic simulations are shown to be in good agreement with experimental results and provide insight for synchronization of inhomogenieties in large sized device. The ability to manipulate spin-dynamics on large size devices could prove useful for increasing the output power of the spin-transfer nano-oscillators (STNOs).

cond-mat.mes-hall