Searcharxiv⌕ Search

arXiv subjects

Rishi Thotli

Publications and source records attributed to Rishi Thotli.

2 recordsLinked to original sources

An Energy-Efficient Approximate Posit Multiply-Divide Unit

In modern computing units, division operations are generally slower than other arithmetic operations and require more resources, such as area and power, than multiplication. To reduce the delay, fast division algorithms use an initial approximation of the reciprocal of the divisor and iteratively approach the correct value, followed by multiplication with the dividend. The hardware architecture and choice of algorithm can significantly alter the overall performance of the division unit. This paper proposes a reduced-accuracy division method for the posit number system, which is an alternative to the traditional floating-point system. The proposed design uses a Look-Up Table (LUT) and a single subtraction operation to perform approximate divisor reciprocation by leveragingthemathematicalsymmetriesofthepositnumbersystem.The paper also presents a hardware architecture that combines multiplication and division units. The reciprocal calculation has been incorporated into the posit Decoder, a common unit required to perform any hardware operation with posits. Compared to existing hardware implementations of division, the proposed method requires significantly fewer operations at the cost of perfect rounding for division. The proposed architecture was simulated using the Cadence RTL v7.1 E2 compiler at the TSMC 90 nm process node and achieves a Power Delay Product (PDP) reduction of 78.8% compared to an existing design that performs exact division, while only 46.33% of the area is required. The experimental results also demonstrate the effectiveness of the proposed system in improving the efficiency of multiplication in posit-based systems.

cs.AR↗

Closing the Gap Between Float and Posit Hardware Efficiency

The b-posit, or bounded posit, is a variation of the posit format designed for high performance computing (HPC) and AI applications. Unlike traditional floating-point formats (floats), posits use variable-length fields for exponent scaling and significand, providing better efficiency for the same bit width. However, this flexibility introduces high worst-case overhead in decode-encode logic, exceeding the cost of handling subnormals for floats. To address this, the b-posit restricts the regime field to a 6-bit limit, reducing variability in regime and fraction sizes. With an exponent size eS of 5 bits, the dynamic range is $2^{-192}$ to $2^{192}$ (about $10^{-58}$ to $10^{58}$) and the quire size is 800 bits, for any precision $n>12$. This constraint improves numerical properties and simplifies hardware implementation by allowing decode-encode operations with basic multiplexers. Our 32-bit b-posit decoder circuits achieve significant improvements: 79 percent less power consumption, 71 percent smaller area, and 60 percent reduced latency compared to standard posit decoders. The 32-bit b-posit encoder shows 68 percent lower power usage, 46 percent less area, and 44 percent shorter delay. The proposed b-posit hardware exhibits superior scalability with increasing bit widths, outperforming standard posit hardware at higher precisions, with even greater advantages at 64-bit. Notably, the b-posit decode-encode hardware matches or exceeds IEEE compliant 32-bit floating-point performance, offering faster and smaller area implementation, with slight increase in worst-case power due to higher speed. The b-posit hardware design provides the clean mathematical behavior and higher accuracy of posits versus IEEE floats without the power, area, or latency costs observed for the Posit Standard (2022). We believe the b-posit should influence future standard revisions.

cs.AR↗