SearcharxivSearch

arXiv subjects

Antony Joseph

Publications and source records attributed to Antony Joseph.

17 recordsLinked to original sources

SparseCol: A 1320 BTOPS/W Precision-scalable NPU Exploiting Training-free Structured Bit-level Sparsity and Dynamic Dataflow

Bit-serial computation enables sequential processing of data at the bit level, providing several advantages, such as scalable computational precision. This approach has gained significant attention, especially for exploiting bit-level sparsity in AI workloads. While current bit-serial processors leverage bit-level sparsity to eliminate the computation associated with zero bits, they face a fundamental trade-off: either they suffer from low memory-access and computation efficiency caused by irregular patterns of non-zero bits, or they incur substantial area overhead from complex online scheduling mechanisms required to reorganize bit-level data and preserve memory access and computation regularity. Therefore, we present the SparseCol processor, designed to harness extensive bit sparsity while maintaining high hardware utilization across various AI applications, including CNNs, RNNs, and transformers. In contrast to traditional methods, SparseCol exploits structured bit-level sparsity, denoted by bit-column sparsity, without requiring any re-training. Furthermore, SparseCol implements a dynamic dataflow architecture that tackles hardware under-utilization issues commonly found in existing bit-serial solutions. Fabricated in 16nm CMOS node, SparseCol delivers 1320 BTOPS/W (BTOPS represents Binary Tera-Operations Per Second, calculated as #W bits x #A bits TOPS) peak efficiency while maintaining accuracy, outperforming SotA sparse processors in terms of efficiency by 6.8x. Comprehensive evaluations on CNN classification tasks and transformer architectures demonstrate system-level efficiencies of 745.02 BTOPS/W and 850.5 BTOPS/W, respectively.

eess.SY

BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning Acceleration

Bit-serial computation facilitates bit-wise sequential data processing, offering numerous benefits, such as a reduced area footprint and dynamically-adaptive computational precision. It has emerged as a prominent approach, particularly in leveraging bit-level sparsity in Deep Neural Networks (DNNs). However, existing bit-serial accelerators exploit bit-level sparsity to reduce computations by skipping zero bits, but they suffer from inefficient memory accesses due to the irregular indices of the non-zero bits. As memory accesses typically are the dominant contributor to DNN accelerator performance, this paper introduces a novel computing approach called "bit-column-serial" and a compatible architecture design named "BitWave." BitWave harnesses the advantages of the "bit-column-serial" approach, leveraging structured bit-level sparsity in combination with dynamic dataflow techniques. This achieves a reduction in computations and memory footprints through redundant computation skipping and weight compression. BitWave is able to mitigate the performance drop or the need for retraining that is typically associated with sparsity-enhancing techniques using a post-training optimization involving selected weight bit-flips. Empirical studies conducted on four deep-learning benchmarks demonstrate the achievements of BitWave: (1) Maximally realize 13.25x higher speedup, 7.71x efficiency compared to state-of-the-art sparsity-aware accelerators. (2) Occupying 1.138 mm2 area and consuming 17.56 mW power in 16nm FinFet process node.

eess.SY

Unraveling the Effects of Cluster Transfer-Induced Breakups on $^{12}$C Fragmentation in Hadron Therapy

The capability of standard Geant4 PhysicsLists to address the fragmentation of $^{12}$C$-^{12}$C was assessed through a comparative analysis with experimental cross sections reported by Divay et al. and Dudouet et al. The standard PhysicsLists were found to be inadequate in explaining the fragmentation systematics. To address this limitation, the breakup component of fragmentation was systematically integrated into the standard PhysicsList, which successfully replicated the differential and double differential cross sections for $α$ production. This breakup component was modeled using fresco CDCC-CRC calculations. This novel physics process was then incorporated into the Geant4 framework, facilitating the calculation of dose distributions in water and tissue. The application of this method demonstrated a precise reproduction of the dose deposited at the Bragg peak region, corroborating the experimental data from Liedner et al., thereby enhancing the accurate visibility of dose tailing.

physics.med-ph

CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories

Deep neural networks (DNN) use a wide range of network topologies to achieve high accuracy within diverse applications. This model diversity makes it impossible to identify a single "dataflow" (execution schedule) to perform optimally across all possible layers and network topologies. Several frameworks support the exploration of the best dataflow for a given DNN layer and hardware. However, switching the dataflow from one layer to the next layer within one DNN model can result in hardware inefficiencies stemming from memory data layout mismatch among the layers. Unfortunately, all existing frameworks treat each layer independently and typically model memories as black boxes (one large monolithic wide memory), which ignores the data layout and can not deal with the data layout dependencies of sequential layers. These frameworks are not capable of doing dataflow cross-layer optimization. This work, hence, aims at cross-layer dataflow optimization, taking the data dependency and data layout reshuffling overheads among layers into account. Additionally, we propose to exploit the multibank memories typically present in modern DNN accelerators towards efficiently reshuffling data to support more dataflow at low overhead. These innovations are supported through the Cross-layer Memory-aware Dataflow Scheduler (CMDS). CMDS can model DNN execution energy/latency while considering the different data layout requirements due to the varied optimal dataflow of layers. Compared with the state-of-the-art (SOTA), which performs layer-optimized memory-unaware scheduling, CMDS achieves up to 5.5X energy reduction and 1.35X latency reduction with negligible hardware cost.

cs.AR

On the Microscopic Level Density Models for Nuclei Near Z=28 Shell Closure

A comprehensive test of level density models for explaining the decay of excited compound nuclei, 54 Mn, 56 Fe, 58 Co, 60 Ni, 61 Ni and 63 Cu, in the energy range of 28 - 36 MeV has been performed. The compound nuclei of interest in the desired ranges are populated using 6 Li based transfer reactions. The proton decay spectrum for each excitation energy bins has been measured. The measured proton spectrum has been reproduced using statistical model calculations with different level density models. A variance minimised approach has been employed for analysing the prediction capability of different level density models. This approach has been converged to Gogny Hartree-Fock-Bogoliubov(HFB) microscopic level density model and which is attributed as the most accurate model for the desired nuclei.

nucl-ex

Determination of photo-nuclear cross section of $^{61}$Ni($γ$,xp) reaction via surrogate ratio technique

The photo nuclear reaction cross section of $^{61}$Ni($γ$,xp) reaction have been measured by employing surrogate reaction technique. This indirect method is used for the first time to obtain the cross section of photo nuclear reaction. The compound nucleus $^{61}$Ni$^{*}$ was populated using the transfer reaction $^{59}$Co($^{6}$Li,$α$) at E$_{lab}=$ 40.5 MeV. To calculate the surrogate ratio, $^{60}$Ni($γ$,xp) was selected as reference reaction and the corresponding compound nucleus $^{60}$Ni$^{*}$ was populated using the transfer reaction $^{56}$Fe($^{6}$Li,d) at E$_{lab}=$ 35.9 MeV. The experimental cross section data of the reference reaction has been taken from EXFOR data libraries. Compound nuclear cross section calculations have been done using EMPIRE 3.2.3 code.

nucl-ex

An Indirect Measurement of $^6$Li(n,$γ$) Cross Sections

The $^6$Li(n,$γ$)$^7$Li cross sections in the neutron energy range of 0.6 to 4 MeV have been measured by the experimental implementation of the direct capture formalism. This was done by measuring the $γ$ transition probability experimentally and accounting for the spin factor by theoretical calculation. The electromagnetic transition probabilities from $^7$Li$^*$ analogous to the initial neutron capture states of $^6$Li$+n$ were measured by populating the J$_i$ states of $^7$Li through $^7$Li($p,p'$)$^7$Li$^*$ reaction. The impact of coupling of resonant states, above neutron separation threshold of $^7$Li, in the neutron capture, is observed from the capture $γ$ spectrum. The measured cross sections were reproduced through {\sc fresco} and Talys-1.95 Direct Capture Calculations.

nucl-ex

Impact of $^7$Be breakup on $^7$Li(p,n) Neutron Spectrum

The formation of continuum neutron distribution in $^7$Li(p,n) has been identified as due to the coupling of the $^7$Be breakup levels to the final state of the reaction. The continuum neutron spectra produced by $^7$Li(p,n) reaction has been estimated by measuring the double differential cross sections for continuum and resonant breakup of $^7$Be, through $^7$Li(p,n)$^7$Be$^*$ reaction at 21 MeV of proton energy. The breakup contributions from continuum and $5/2^-$, $7/2^-$ states of $^7$Be have been identified. The measured double differential cross sections have been reproduced through CDCC-CRC calculations. The cross sections were projected to neutron spectrum using Monte-Carlo approach and validated using experimentally measured $^3$He gated neutron spectra. $^7$Li(p,n) neutron spectrum at 20 MeV incident proton energy measured by McNaughton. et al. has been reproduced by adapting estimated model parameters for the reaction.

nucl-ex

A systematic study of the ground state properties of W, Os and Pt isotopes using HFB theory

A systematic study of the ground state properties of transitional nuclei W, Os and Pt is conducted with the help of Skyrme Hartree-Fock-Bogoliubov theory. Different Skyrme interactions are employed in the study. Two different bases, harmonic oscillator and transformed harmonic oscillator, are used in our investigation. 2n-separation energy, charge radii, neutron and proton rms radii, neutron skin thickness and deformation parameter have been estimated. The results obtained are in good agreement with the available experimental values.

nucl-th

A systematic study of alpha and cluster decay in Platinum isotopes

The feasibility of alpha and cluster decay from Pt isotopes has been investigated within the framework of Skyrme Hartree-Fock-Bogoliubov theory. Calculation has been carried out for various Skyrme forces. Harmonic oscillator and transformed harmonic oscillator basis are used to solve HFB equations. The role played by shell closure is withal analysed. Half-lives are estimated with the help of Universal Decay Law (UDL). Geiger-Nuttel plots are also plotted and successfully preserves its linear nature.

nucl-th

Alpha and cluster decay half lives in Tungsten isotopes: A microscopic analysis

Alpha and cluster decay half lives for W isotopes in the range between 2p drip line and beta stability line are studied. The sensitivity of different Skyrme parametrizations in predicting the alpha decay and probable cluster decay modes from W isotopes have been analysed. The half lives are calculated using UDL. Predicted half lives are compared with ELDM and also with available experimental values. The study also revealed the role of neutron shell closure in cluster decay process.

nucl-th

Lossy Compression via Sparse Linear Regression: Performance under Minimum-distance Encoding

We study a new class of codes for lossy compression with the squared-error distortion criterion, designed using the statistical framework of high-dimensional linear regression. Codewords are linear combinations of subsets of columns of a design matrix. Called a Sparse Superposition or Sparse Regression codebook, this structure is motivated by an analogous construction proposed recently by Barron and Joseph for communication over an AWGN channel. For i.i.d Gaussian sources and minimum-distance encoding, we show that such a code can attain the Shannon rate-distortion function with the optimal error exponent, for all distortions below a specified value. It is also shown that sparse regression codes are robust in the following sense: a codebook designed to compress an i.i.d Gaussian source of variance $σ^2$ with (squared-error) distortion $D$ can compress any ergodic source of variance less than $σ^2$ to within distortion $D$. Thus the sparse regression ensemble retains many of the good covering properties of the i.i.d random Gaussian ensemble, while having having a compact representation in terms of a matrix whose size is a low-order polynomial in the block-length.

cs.IT

Impact of regularization on Spectral Clustering

The performance of spectral clustering can be considerably improved via regularization, as demonstrated empirically in Amini et. al (2012). Here, we provide an attempt at quantifying this improvement through theoretical analysis. Under the stochastic block model (SBM), and its extensions, previous results on spectral clustering relied on the minimum degree of the graph being sufficiently large for its good performance. By examining the scenario where the regularization parameter $τ$ is large we show that the minimum degree assumption can potentially be removed. As a special case, for an SBM with two blocks, the results require the maximum degree to be large (grow faster than $\log n$) as opposed to the minimum degree. More importantly, we show the usefulness of regularization in situations where not all nodes belong to well-defined clusters. Our results rely on a `bias-variance'-like trade-off that arises from understanding the concentration of the sample Laplacian and the eigen gap as a function of the regularization parameter. As a byproduct of our bounds, we propose a data-driven technique \textit{DKest} (standing for estimated Davis-Kahan bounds) for choosing the regularization parameter. This technique is shown to work well through simulations and on a real data set.

stat.ML

Fast Sparse Superposition Codes have Exponentially Small Error Probability for R < C

For the additive white Gaussian noise channel with average codeword power constraint, sparse superposition codes are developed. These codes are based on the statistical high-dimensional regression framework. The paper [IEEE Trans. Inform. Theory 55 (2012), 2541 - 2557] investigated decoding using the optimal maximum-likelihood decoding scheme. Here a fast decoding algorithm, called adaptive successive decoder, is developed. For any rate R less than the capacity C communication is shown to be reliable with exponentially small error probability.

cs.IT

Variable Selection in High Dimensions with Random Designs and Orthogonal Matching Pursuit

The performance of Orthogonal Matching Pursuit (OMP) for variable selection is analyzed for random designs. When contrasted with the deterministic case, since the performance is here measured after averaging over the distribution of the design matrix, one can have far less stringent sparsity constraints on the coefficient vector. We demonstrate that for exact sparse vectors, the performance of the OMP is similar to known results on the Lasso algorithm [\textit{IEEE Trans. Inform. Theory} \textbf{55} (2009) 2183--2202]. Moreover, variable selection under a more relaxed sparsity assumption on the coefficient vector, whereby one has only control on the $\ell_1$ norm of the smaller coefficients, is also analyzed. As a consequence of these results, we also show that the coefficient estimate satisfies strong oracle type inequalities.

stat.ML

Toward Fast Reliable Communication at Rates Near Capacity with Gaussian Noise

For the additive Gaussian noise channel with average codeword power constraint, sparse superposition codes and adaptive successive decoding is developed. Codewords are linear combinations of subsets of vectors, with the message indexed by the choice of subset. A feasible decoding algorithm is presented. Communication is reliable with error probability exponentially small for all rates below the Shannon capacity.

cs.IT

Least Squares Superposition Codes of Moderate Dictionary Size, Reliable at Rates up to Capacity

For the additive white Gaussian noise channel with average codeword power constraint, new coding methods are devised in which the codewords are sparse superpositions, that is, linear combinations of subsets of vectors from a given design, with the possible messages indexed by the choice of subset. Decoding is by least squares, tailored to the assumed form of linear combination. Communication is shown to be reliable with error probability exponentially small for all rates up to the Shannon capacity.

cs.IT