SearcharxivSearch

arXiv subjects

Akash Kumar

Publications and source records attributed to Akash Kumar.

At least 91 records · Page 5Linked to original sources

Robust Empirical Risk Minimization with Tolerance

Developing simple, sample-efficient learning algorithms for robust classification is a pressing issue in today's tech-dominated world, and current theoretical techniques requiring exponential sample complexity and complicated improper learning rules fall far from answering the need. In this work we study the fundamental paradigm of (robust) $\textit{empirical risk minimization}$ (RERM), a simple process in which the learner outputs any hypothesis minimizing its training error. RERM famously fails to robustly learn VC classes (Montasser et al., 2019a), a bound we show extends even to `nice' settings such as (bounded) halfspaces. As such, we study a recent relaxation of the robust model called $\textit{tolerant}$ robust learning (Ashtiani et al., 2022) where the output classifier is compared to the best achievable error over slightly larger perturbation sets. We show that under geometric niceness conditions, a natural tolerant variant of RERM is indeed sufficient for $γ$-tolerant robust learning VC classes over $\mathbb{R}^d$, and requires only $\tilde{O}\left( \frac{VC(H)d\log \frac{D}{γδ}}{ε^2}\right)$ samples for robustness regions of (maximum) diameter $D$.

cs.LG

Robust mutual synchronization in long spin Hall nano-oscillator chains

Mutual synchronization of N serially connected spintronic nano-oscillators increases their coherence by a factor $N$ and their output power by $N^2$. Increasing the number of mutually synchronized nano-oscillators in chains is hence of great importance for better signal quality and also for emerging applications such as oscillator-based neuromorphic computing and Ising machines where larger N can tackle larger problems. Here we fabricate spin Hall nano-oscillator chains of up to 50 serially connected nano-constrictions in W/NiFe, W/CoFeB/MgO, and NiFe/Pt stacks and demonstrate robust and complete mutual synchronization of up to 21 nano-constrictions, reaching linewidths of below 200 kHz and quality factors beyond 79,000, while operating at 10 GHz. We also find a square increase in the peak power with the increasing number of mutually synchronized oscillators, resulting in a factor of 400 higher peak power in long chains compared to individual nano-constrictions. Although chains longer than 21 nano-constrictions also show complete mutual synchronization, it is not as robust and their signal quality does not improve as much as they prefer to break up into partially synchronized states. The low current and low field operation of these oscillators along with their wide frequency tunability (2-28 GHz) with both current and magnetic fields, make them ideal candidates for on-chip GHz-range applications and neuromorphic computing.

cond-mat.mes-hall

Bounding the List Color Function Threshold from Above

The chromatic polynomial of a graph $G$, denoted $P(G,m)$, is equal to the number of proper $m$-colorings of $G$ for each $m \in \mathbb{N}$. In 1990, Kostochka and Sidorenko introduced the list color function of graph $G$, denoted $P_{\ell}(G,m)$, which is a list analogue of the chromatic polynomial. The list color function threshold of $G$, denoted $τ(G)$, is the smallest $k \geq χ(G)$ such that $P_{\ell}(G,m) = P(G,m)$ whenever $m \geq k$. It is known that for every graph $G$, $τ(G)$ is finite, and in fact, $τ(G) \leq (|E(G)|-1)/\ln(1+ \sqrt{2}) + 1$. It is also known that when $G$ is a cycle or chordal graph, $G$ is enumeratively chromatic-choosable which means $τ(G) = χ(G)$. A recent paper of Kaul et al. suggests that understanding the list color function threshold of complete bipartite graphs is essential to the study of the extremal behavior of $τ$. In this paper we show that for any $n \geq 2$, $τ(K_{2,n}) \leq \lceil (n+2.05)/1.24 \rceil$ which gives an improvement on the general upper bound for $τ(G)$ when $G = K_{2,n}$. We also develop additional tools that allow us to show that $τ(K_{2,3}) = χ(K_{2,3})$ and $τ(K_{2,4}) = τ(K_{2,5}) = 3$.

math.CO

Large spin Hall conductivity in epitaxial thin films of kagome antiferromagnet Mn$_3$Sn at room temperature

Mn$_3$Sn is a non-collinear antiferromagnetic quantum material that exhibits a magnetic Weyl semimetallic state and has great potential for efficient memory devices. High-quality epitaxial $c$-plane Mn$_3$Sn thin films have been grown on a sapphire substrate using a Ru seed layer. Using spin pumping induced inverse spin Hall effect measurements on $c$-plane epitaxial Mn$_3$Sn/Ni$_{80}$Fe$_{20}$, we measure spin-diffusion length ($λ_{\rm Mn_3Sn}$), and spin Hall conductivity ($σ_{\rm{SH}}$) of Mn$_3$Sn thin films: $λ_{\rm Mn_3Sn}=0.42\pm 0.04$ nm and $σ_{\rm{SH}}=-702~\hbar/ e~Ω^{-1}$cm$^{-1}$. While $λ_{\rm Mn_3Sn}$ is consistent with earlier studies, $σ_{\rm{SH}}$ is an order of magnitude higher and of the opposite sign. The behavior is explained on the basis of excess Mn, which shifts the Fermi level in our films, leading to the observed behavior. Our findings demonstrate a technique for engineering $σ_{\rm{SH}}$ of Mn$_3$Sn films by employing Mn composition for functional spintronic devices.

cond-mat.mtrl-sci

Large spin-to-charge conversion at the two-dimensional interface of transition metal dichalcogenides and permalloy

Spin-to-charge conversion is an essential requirement for the implementation of spintronic devices. Recently, monolayers of semiconducting transition metal dichalcogenides (TMDs) have attracted considerable interest for spin-to-charge conversion due to their high spin-orbit coupling and lack of inversion symmetry in their crystal structure. However, reports of direct measurement of spin-to-charge conversion at TMD-based interfaces are very much limited. Here, we report on the room temperature observation of a large spin-to-charge conversion arising from the interface of Ni$_{80}$Fe$_{20}$ (Py) and four distinct large area ($\sim 5\times2$~mm$^2$) monolayer (ML) TMDs namely, MoS$_2$, MoSe$_2$, WS$_2$, and WSe$_2$. We show that both spin mixing conductance and the Rashba efficiency parameter ($λ_{IREE}$) scales with the spin-orbit coupling strength of the ML TMD layers. The $λ_{IREE}$ parameter is found to range between $-0.54$ and $-0.76$ nm for the four monolayer TMDs, demonstrating a large spin-to-charge conversion. Our findings reveal that TMD/ferromagnet interface can be used for efficient generation and detection of spin current, opening new opportunities for novel spintronic devices.

physics.app-ph

Combining Gradients and Probabilities for Heterogeneous Approximation of Neural Networks

This work explores the search for heterogeneous approximate multiplier configurations for neural networks that produce high accuracy and low energy consumption. We discuss the validity of additive Gaussian noise added to accurate neural network computations as a surrogate model for behavioral simulation of approximate multipliers. The continuous and differentiable properties of the solution space spanned by the additive Gaussian noise model are used as a heuristic that generates meaningful estimates of layer robustness without the need for combinatorial optimization techniques. Instead, the amount of noise injected into the accurate computations is learned during network training using backpropagation. A probabilistic model of the multiplier error is presented to bridge the gap between the domains; the model estimates the standard deviation of the approximate multiplier error, connecting solutions in the additive Gaussian noise space to actual hardware instances. Our experiments show that the combination of heterogeneous approximation and neural network retraining reduces the energy consumption for multiplications by 70% to 79% for different ResNet variants on the CIFAR-10 dataset with a Top-1 accuracy loss below one percentage point. For the more complex Tiny ImageNet task, our VGG16 model achieves a 53 % reduction in energy consumption with a drop in Top-5 accuracy of 0.5 percentage points. We further demonstrate that our error model can predict the parameters of an approximate multiplier in the context of the commonly used additive Gaussian noise (AGN) model with high accuracy. Our software implementation is available under https://github.com/etrommer/agn-approx.

cs.LG

End-to-End Semi-Supervised Learning for Video Action Detection

In this work, we focus on semi-supervised learning for video action detection which utilizes both labeled as well as unlabeled data. We propose a simple end-to-end consistency based approach which effectively utilizes the unlabeled data. Video action detection requires both, action class prediction as well as a spatio-temporal localization of actions. Therefore, we investigate two types of constraints, classification consistency, and spatio-temporal consistency. The presence of predominant background and static regions in a video makes it challenging to utilize spatio-temporal consistency for action detection. To address this, we propose two novel regularization constraints for spatio-temporal consistency; 1) temporal coherency, and 2) gradient smoothness. Both these aspects exploit the temporal continuity of action in videos and are found to be effective for utilizing unlabeled videos for action detection. We demonstrate the effectiveness of the proposed approach on two different action detection benchmark datasets, UCF101-24 and JHMDB-21. In addition, we also show the effectiveness of the proposed approach for video object segmentation on the Youtube-VOS which demonstrates its generalization capability The proposed approach achieves competitive performance by using merely 20% of annotations on UCF101-24 when compared with recent fully supervised methods. On UCF101-24, it improves the score by +8.9% and +11% at 0.5 f-mAP and v-mAP respectively, compared to supervised approach.

cs.CV

RAPID: AppRoximAte Pipelined Soft Multipliers and Dividers for High-Throughput and Energy-Efficiency

The rapid updates in error-resilient applications along with their quest for high throughput have motivated designing fast approximate functional units for Field-Programmable Gate Arrays (FPGAs). Studies that proposed imprecise functional techniques are posed with three shortcomings: first, most inexact multipliers and dividers are specialized for Application-Specific Integrated Circuit (ASIC) platforms. Second, state-of-the-art (SoA) approximate units are substituted, mostly in a single kernel of a multi-kernel application. Moreover, the end-to-end assessment is adopted on the Quality of Results (QoR), but not on the overall gained performance. Finally, existing imprecise components are not designed to support a pipelined approach, which could boost the operating frequency/throughput of, e.g., division-included applications. In this paper, we propose RAPID, the first pipelined approximate multiplier and divider architecture, customized for FPGAs. The proposed units efficiently utilize 6-input Look-up Tables (6-LUTs) and fast carry chains to implement Mitchell's approximate algorithms. Our novel error-refinement scheme not only has negligible overhead over the baseline Mitchell's approach but also boosts its accuracy to 99.4% for arbitrary size of multiplication and division. Experimental results demonstrate the efficiency of the proposed pipelined and non-pipelined RAPID multipliers and dividers over accurate counterparts. Moreover, the end-to-end evaluations of RAPID, deployed in three multi-kernel applications in the domains of bio-signal processing, image processing, and moving object tracking for Unmanned Air Vehicles (UAV) indicate up to 45% improvements in area, latency, and Area-Delay-Product (ADP), respectively, over accurate kernels, with negligible loss in QoR.

cs.AR

Exact recovery algorithm for Planted Bipartite Graph in Semi-random Graphs

The problem of finding the largest induced balanced bipartite subgraph in a given graph is NP-hard. This problem is closely related to the problem of finding the smallest Odd Cycle Transversal. In this work, we consider the following model of instances: starting with a set of vertices $V$, a set $S \subseteq V$ of $k$ vertices is chosen and an arbitrary $d$-regular bipartite graph is added on it; edges between pairs of vertices in $S \times (V \setminus S)$ and $(V \setminus S) \times (V \setminus S)$ are added with probability $p$. Since for $d=0$, the problem reduces to recovering a planted independent set, we don't expect efficient algorithms for $k=o(\sqrt{n})$. This problem is a generalization of the planted balanced biclique problem where the bipartite graph induced on $S$ is a complete bipartite graph; [Lev18] gave an algorithm for recovering $S$ in this problem when $k=Ω(\sqrt{n})$. Our main result is an efficient algorithm that recovers (w.h.p.) the planted bipartite graph when $k=Ω_p(\sqrt{n \log n})$ for a large range of parameters. Our results also hold for a natural semi-random model of instances, which involve the presence of a monotone adversary. Our proof shows that a natural SDP relaxation for the problem is integral by constructing an appropriate solution to it's dual formulation. Our main technical contribution is a new approach for constructing the dual solution where we calibrate the eigenvectors of the adjacency matrix to be the eigenvectors of the dual matrix. We believe that this approach may have applications to other recovery problems in semi-random models as well. When $k=Ω(\sqrt{n})$, we give an algorithm for recovering $S$ whose running time is exponential in the number of small eigenvalues in graph induced on $S$; this algorithm is based on subspace enumeration techniques due to the works of [KT07,ABS10,Kol11].

cs.DS

Video Action Detection: Analysing Limitations and Challenges

Beyond possessing large enough size to feed data hungry machines (eg, transformers), what attributes measure the quality of a dataset? Assuming that the definitions of such attributes do exist, how do we quantify among their relative existences? Our work attempts to explore these questions for video action detection. The task aims to spatio-temporally localize an actor and assign a relevant action class. We first analyze the existing datasets on video action detection and discuss their limitations. Next, we propose a new dataset, Multi Actor Multi Action (MAMA) which overcomes these limitations and is more suitable for real world applications. In addition, we perform a biasness study which analyzes a key property differentiating videos from static images: the temporal aspect. This reveals if the actions in these datasets really need the motion information of an actor, or whether they predict the occurrence of an action even by looking at a single frame. Finally, we investigate the widely held assumptions on the importance of temporal ordering: is temporal ordering important for detecting these actions? Such extreme experiments show existence of biases which have managed to creep into existing methods inspite of careful modeling.

cs.CV

dCSR: A Memory-Efficient Sparse Matrix Representation for Parallel Neural Network Inference

Reducing the memory footprint of neural networks is a crucial prerequisite for deploying them in small and low-cost embedded devices. Network parameters can often be reduced significantly through pruning. We discuss how to best represent the indexing overhead of sparse networks for the coming generation of Single Instruction, Multiple Data (SIMD)-capable microcontrollers. From this, we develop Delta-Compressed Storage Row (dCSR), a storage format for sparse matrices that allows for both low overhead storage and fast inference on embedded systems with wide SIMD units. We demonstrate our method on an ARM Cortex-M55 MCU prototype with M-Profile Vector Extension(MVE). A comparison of memory consumption and throughput shows that our method achieves competitive compression ratios and increases throughput over dense methods by up to $2.9 \times$ for sparse matrix-vector multiplication (SpMV)-based kernels and $1.06 \times$ for sparse matrix-matrix multiplication (SpMM). This is accomplished through handling the generation of index information directly in the SIMD unit, leading to an increase in effective memory bandwidth.

cs.DS

Ultrathin ferrimagnetic GdFeCo films with very low damping

Ferromagnetic materials dominate as the magnetically active element in spintronic devices, but come with drawbacks such as large stray fields, and low operational frequencies. Compensated ferrimagnets provide an alternative as they combine the ultrafast magnetization dynamics of antiferromagnets with a ferromagnet-like spin-orbit-torque (SOT) behavior. However to use ferrimagnets in spintronic devices their advantageous properties must be retained also in ultrathin films (t < 10 nm). In this study, ferrimagnetic Gdx(Fe87.5Co12.5)1-x thin films in the thickness range t = 2-20 nm were grown on high resistance Si(100) substrates and studied using broadband ferromagnetic resonance measurements at room temperature. By tuning their stoichiometry, a nearly compensated behavior is observed in 2 nm Gdx(Fe87.5Co12.5)1-x ultrathin films for the first time, with an effective magnetization of Meff = 0.02 T and a low effective Gilbert damping constant of α = 0.0078, comparable to the lowest values reported so far in 30 nm films. These results show great promise for the development of ultrafast and energy efficient ferrimagnetic spintronic devices.

cond-mat.mtrl-sci

Energy-efficient W$_{\text{100-x}}$Ta$_{\text{x}}$/CoFeB/MgO spin Hall nano-oscillators

We investigate a W-Ta alloying route to reduce the auto-oscillation current densities and the power consumption of nano-constriction based spin Hall nano oscillators. Using spin-torque ferromagnetic resonance (ST-FMR) measurements on microbars of W$_{\text{100-x}}$Ta$_{\text{x}}$(5 nm)/CoFeB(t)/MgO stacks with t = 1.4, 1.8, and 2.0 nm, we measure a substantial improvement in both the spin-orbit torque efficiency and the spin Hall conductivity. We demonstrate a 34\% reduction in threshold auto-oscillation current density, which translates into a 64\% reduction in power consumption as compared to pure W-based SHNOs. Our work demonstrates the promising aspects of W-Ta alloying for the energy-efficient operation of emerging spintronic devices.

cond-mat.mes-hall

Fabrication of voltage gated spin Hall nano-oscillators

We demonstrate an optimized fabrication process for electric field (voltage gate) controlled nano-constriction spin Hall nano-oscillators (SHNOs), achieving feature sizes of <30 nm with easy to handle ma-N 2401 e-beam lithography negative tone resist. For the nanoscopic voltage gates, we utilize a two-step tilted ion beam etching approach and through-hole encapsulation using 30 nm HfO x . The optimized tilted etching process reduces sidewalls by 75% compared to no tilting. Moreover, the HfO x encapsulation avoids any sidewall shunting and improves gate breakdown. Our experimental results on W/CoFeB/MgO/SiO 2 SHNOs show significant frequency tunability (6 MHz/V) even for moderate perpendicular magnetic anisotropy. Circular patterns with diameter of 45 nm are achieved with an aspect ratio better than 0.85 for 80% of the population. The optimized fabrication process allows incorporating a large number of individual gates to interface to SHNO arrays for unconventional computing and densely packed spintronic neural networks.

cond-mat.mes-hall

The complexity of testing all properties of planar graphs, and the role of isomorphism

Consider property testing on bounded degree graphs and let $\varepsilon>0$ denote the proximity parameter. A remarkable theorem of Newman-Sohler (SICOMP 2013) asserts that all properties of planar graphs (more generally hyperfinite) are testable with query complexity only depending on $\varepsilon$. Recent advances in testing minor-freeness have proven that all additive and monotone properties of planar graphs can be tested in $poly(\varepsilon^{-1})$ queries. Some properties falling outside this class, such as Hamiltonicity, also have a similar complexity for planar graphs. Motivated by these results, we ask: can all properties of planar graphs can be tested in $poly(\varepsilon^{-1})$ queries? Is there a uniform query complexity upper bound for all planar properties, and what is the "hardest" such property to test? We discover a surprisingly clean and optimal answer. Any property of bounded degree planar graphs can be tested in $\exp(O(\varepsilon^{-2}))$ queries. Moreover, there is a matching lower bound, up to constant factors in the exponent. The natural property of testing isomorphism to a fixed graph needs $\exp(Ω(\varepsilon^{-2}))$ queries, thereby showing that (up to polynomial dependencies) isomorphism to an explicit fixed graph is the hardest property of planar graphs. The upper bound is a straightforward adapation of the Newman-Sohler analysis that tracks dependencies on $\varepsilon$ carefully. The main technical contribution is the lower bound construction, which is achieved by a special family of planar graphs that are all mutually far from each other. We can also apply our techniques to get analogous results for bounded treewidth graphs. We prove that all properties of bounded treewidth graphs can be tested in $\exp(O(\varepsilon^{-1}\log \varepsilon^{-1}))$ queries. Moreover, testing isomorphism to a fixed forest requires $\exp(Ω(\varepsilon^{-1}))$ queries.

cs.DS

NMPO: Near-Memory Computing Profiling and Offloading

Real-world applications are now processing big-data sets, often bottlenecked by the data movement between the compute units and the main memory. Near-memory computing (NMC), a modern data-centric computational paradigm, can alleviate these bottlenecks, thereby improving the performance of applications. The lack of NMC system availability makes simulators the primary evaluation tool for performance estimation. However, simulators are usually time-consuming, and methods that can reduce this overhead would accelerate the early-stage design process of NMC systems. This work proposes Near-Memory computing Profiling and Offloading (NMPO), a high-level framework capable of predicting NMC offloading suitability employing an ensemble machine learning model. NMPO predicts NMC suitability with an accuracy of 85.6% and, compared to prior works, can reduce the prediction time by using hardware-dependent applications features by up to 3 order of magnitude.

cs.AR

Toward the Design of Fault-Tolerance- and Peak- Power-Aware Multi-Core Mixed-Criticality Systems

Mixed-Criticality (MC) systems have recently been devised to address the requirements of real-time systems in industrial applications, where the system runs tasks with different criticality levels on a single platform. In some workloads, a high-critically task might overrun and overload the system, or a fault can occur during the execution. However, these systems must be fault-tolerant and guarantee the correct execution of all high-criticality tasks by their deadlines to avoid catastrophic consequences, in any situation. Furthermore, in these MC systems, the peak power consumption of the system may increase, especially in an overload situation and exceed the processor Thermal Design Power (TDP) constraint. This may cause generating heat beyond the cooling capacity, resulting the system stop to avoid excessive heat and halting the processor. In this paper, we propose a technique for dependent dual-criticality tasks in fault-tolerant multi-core MC systems to manage peak power consumption and temperature. The technique develops a tree of possible task mapping and scheduling at design-time to cover all possible scenarios and reduce the low-criticality task drop rate in the high-criticality mode. At run-time, the system exploits the tree to select a proper schedule according to fault occurrences and criticality mode changes. Experimental results show that the average task schedulability is 74.14% on average for the proposed method, while the peak power consumption and maximum temperature are improved by 16.65% and 14.9 C on average, respectively, compared to a recent work. In addition, for a real-life application, our method reduces the peak power and maximum temperature by up to 20.06% and 5 C, respectively, compared to a state-of-the-art approach.

cs.DC

Random walks and forbidden minors III: poly(d/ε)-time partition oracles for minor-free graph classes

Consider the family of bounded degree graphs in any minor-closed family (such as planar graphs). Let d be the degree bound and n be the number of vertices of such a graph. Graphs in these classes have hyperfinite decompositions, where, for a sufficiently small \e > 0, one removes \edn edges to get connected components of size independent of n. An important tool for sublinear algorithms and property testing for such classes is the partition oracle, introduced by the seminal work of Hassidim-Kelner-Nguyen-Onak (FOCS 2009). A partition oracle is a local procedure that gives consistent access to a hyperfinite decomposition, without any preprocessing. Given a query vertex v, the partition oracle outputs the component containing v in time independent of n. All the answers are consistent with a single hyperfinite decomposition. The partition oracle of Hassidim et al. runs in time d^poly(d/\e) per query. They pose the open problem of whether poly(d/\e)-time partition oracles exist. Levi-Ron (ICALP 2013) give a refinement of the previous approach, to get a partition oracle that runs in time d^{\log(d/\e)-per query. In this paper, we resolve this open problem and give \poly(d/\e)-time partition oracles for bounded degree graphs in any minor-closed family. Unlike the previous line of work based on combinatorial methods, we employ techniques from spectral graph theory. We build on a recent spectral graph theoretical toolkit for minor-closed graph families, introduced by the authors to develop efficient property testers. A consequence of our result is a poly(d/\e)-query tester for any monotone and additive property of minor-closed families (such as bipartite planar graphs). Our result also gives poly(d/\e)-query algorithms for additive {\e}n-approximations for problems such as maximum matching, minimum vertex cover, maximum independent set, and minimum dominating set for these graph families.

cs.DS