SearcharxivSearch

arXiv subjects

Yu-Hsin Chen

Publications and source records attributed to Yu-Hsin Chen.

17 recordsLinked to original sources

Shubnikov-de Haas oscillations of two-dimensional electron gases in AlYN/GaN and AlScN/GaN heterostructures

AlYN and AlScN have recently emerged as promising nitride materials that can be integrated with GaN to form two-dimensional electron gases (2DEGs) at heterojunctions. Electron transport properties in these heterostructures have been enhanced through careful design and optimization of epitaxial growth conditions. In this work, we report for the first time Shubnikov-de Haas (SdH) oscillations of 2DEGs in AlYN/GaN and AlScN/GaN heterostructures, grown by metal-organic chemical vapor deposition. SdH oscillations provide direct access to key 2DEG parameters at the Fermi level: (1) carrier density, (2) electron effective mass (m* ~ 0.24 me for AlYN/GaN and m* ~ 0.25 me for AlScN/GaN), and (3) quantum scattering time (~ 68 fs for AlYN/GaN and ~ 70 fs for AlScN/GaN). These measurements of fundamental transport properties provide critical insights for advancing emerging nitride semiconductors for future high-frequency and power electronics.

cond-mat.mes-hall

XHEMTs on Ultrawide Bandgap Single-Crystal AlN Substrates

AlN has the largest bandgap in the wurtzite III-nitride semiconductor family, making it an ideal barrier for a thin GaN channel to achieve strong carrier confinement in field-effect transistors, analogous to silicon-on-insulator technology. Unlike SiO$_2$/Si/SiO$_2$, AlN/GaN/AlN can be grown fully epitaxially, enabling high carrier mobilities suitable for high-frequency applications. However, developing these heterostructures and related devices has been hindered by challenges in strain management, polarization effects, defect control and charge trapping. Here, the AlN single-crystal high electron mobility transistor (XHEMT) is introduced, a new nitride transistor technology designed to address these issues. The XHEMT structure features a pseudomorphic GaN channel sandwiched between AlN layers, grown on single-crystal AlN substrates. First-generation XHEMTs demonstrate RF performance on par with the state-of-the-art GaN HEMTs, achieving 5.92 W/mm output power and 65% peak power-added efficiency at 10 GHz under 17 V drain bias. These devices overcome several limitations present in conventional GaN HEMTs, which are grown on lattice-mismatched foreign substrates that introduce undesirable dislocations and exacerbated thermal resistance. With the recent availability of 100-mm AlN substrates and AlN's high thermal conductivity (340 W/m$\cdot$K), XHEMTs show strong potential for next-generation RF electronics.

cond-mat.mtrl-sci

Epitaxial high-K AlBN barrier GaN HEMTs

We report a polarization-induced 2D electron gas (2DEG) at an epitaxial AlBN/GaN heterojunction grown on a SiC substrate. Using this 2DEG in a long conducting channel, we realize ultra-thin barrier AlBN/GaN high electron mobility transistors that exhibit current densities of more than 0.25 A/mm, clean current saturation, a low pinch-off voltage of -0.43 V, and a peak transconductance of 0.14 S/mm. Transistor performance in this preliminary realization is limited by the contact resistance. Capacitance-voltage measurements reveal that introducing 7 % B in the epitaxial AlBN barrier on GaN boosts the relative dielectric constant of AlBN to 16, higher than the AlN dielectric constant of 9. Epitaxial high-K barrier AlBN/GaN HEMTs can thus extend performance beyond the capabilities of current GaN transistors.

physics.app-ph

Shubnikov-de Haas oscillations in coherently strained AlN/GaN/AlN quantum wells on bulk AlN substrates

We report the observation of Shubnikov-de Haas (SdH) oscillations in coherently strained, low-dislocation AlN/GaN/AlN quantum wells (QWs), including both undoped and $\delta$-doped structures. SdH measurements reveal a single subband occupation in the undoped GaN QW and two subband occupation in the $\delta$-doped GaN QW. More importantly, SdH oscillations enable direct measurement of critical two-dimensional electron gas (2DEG) parameters at the Fermi level: carrier density and ground state energy level, electron effective mass ($m^* \approx 0.289\,m_{\rm e}$ for undoped GaN QW and $m^* \approx 0.298\,m_{\rm e}$ for $\delta$-doped GaN QW), and quantum scattering time ($\tau_{\rm q} \approx 83.4 \, \text{fs}$ for undoped GaN QW and $\tau_{\rm q} \approx 130.6 \, \text{fs}$ for $\delta$-doped GaN QW). These findings provide important insights into the fundamental properties of 2DEGs that are strongly quantum confined in the thin GaN QWs, essential for designing nitride heterostructures for high-performance electronic applications.

cond-mat.mtrl-sci

Universal trimers with p-wave interactions and the faux-Efimov effect

An unusual class of $p$-wave universal trimers with symmetry $L^{\Pi}=1^{\pm}$ is identified, for both a two-component fermionic trimer with $s$- and $p$-wave scattering length close to unitarity and for a one-component fermionic trimer at $p$-wave unitarity. Moreover, fermionic trimers made of atoms with two internal spin components are found for $L^{\Pi}=1^{\pm}$, when the $p$-wave interaction between spin-up and spin-down fermions is close to unitarity and/or when the interaction between two spin-up fermions is close to the $p$-wave unitary limit. The universality of these $p$-wave universal trimers is tested here by considering van der Waals interactions in a Lennard-Jones potential with different numbers of two-body bound states; our calculations also determine the value of the scattering volume or length where the trimer state hits zero energy and can be observed as a recombination resonance. The faux-Efimov effect appears with trimer symmetry $L^{\Pi}=1^{-}$ when the two fermion interactions are close to $p$-wave unitarity and the lowest $1/R^2$ coefficient gets modified, thereby altering the usual Wigner threshold law for inelastic processes involving 3-body continuum channels.

physics.atom-ph

Ferroelectric AlBN Films by Molecular Beam Epitaxy

We report the properties of molecular beam epitaxy deposited AlBN thin films on a recently developed epitaxial nitride metal electrode Nb2N. While a control AlN thin film exhibits standard capacitive behavior, distinct ferroelectric switching is observed in the AlBN films with increasing Boron mole fraction. The measured remnant polarization Pr of 15 uC/cm2 and coercive field Ec of 1.45 MV/cm in these films are smaller than those recently reported on films deposited by sputtering, due to incomplete wake-up, limited by current leakage. Because AlBN preserves the ultrawide energy bandgap of AlN compared to other nitride hi-K dielectrics and ferroelectrics, and it can be epitaxially integrated with GaN and AlN semiconductors, its development will enable several opportunities for unique electronic, photonic, and memory devices.

cond-mat.mtrl-sci

PipeOrgan: Efficient Inter-operation Pipelining with Flexible Spatial Organization and Interconnects

Because of the recent trends in Deep Neural Networks (DNN) models being memory-bound, inter-operator pipelining for DNN accelerators is emerging as a promising optimization. Inter-operator pipelining reduces costly on-chip global memory and off-chip memory accesses by forwarding the output of a layer as the input of the next layer within the compute array, which is proven to be an effective optimization by previous works. However, the design space of inter-operator pipelining is huge, and the space is not yet fully explored. In particular, identifying the right depth and granularity of pipelining (or no pipelining at all) is significantly dependent on the layer shapes and data volumes of weights and activations, and these are different even within a domain. Moreover, works divide the substrate into large chunks and map one layer onto each chunk, which requires communicating halfway through or through the global buffer. However, for fine-grained inter-operation pipelining, placing the corresponding consumer of the next layer tile close to the producer tile of the current layer is a better way to exploit fine-grained spatial reuse. In order to support variable number of layers (ie the right depth) and support multiple spatial organizations of layers (in accordance with the pipelining granularity) on the substrate, we propose PipeOrgan, a new class of spatial data organization strategy for energy efficient and congestion-free communication between the PEs for various pipeline depth and granularity. PipeOrgan takes advantage of flexible spatial organization and can allocate layers to PEs based on the granularity of pipelining. We also propose changes to the conventional mesh topology to improve the performance of coarse-grained allocation. PipeOrgan achieves 1.95x performance improvement over the state-of-the-art pipelined dataflow on XR-bench workloads.

cs.AR

Deep denoising autoencoder-based non-invasive blood flow detection for arteriovenous fistula

Clinical guidelines underscore the importance of regularly monitoring and surveilling arteriovenous fistula (AVF) access in hemodialysis patients to promptly detect any dysfunction. Although phono-angiography/sound analysis overcomes the limitations of standardized AVF stenosis diagnosis tool, prior studies have depended on conventional feature extraction methods, restricting their applicability in diverse contexts. In contrast, representation learning captures fundamental underlying factors that can be readily transferred across different contexts. We propose an approach based on deep denoising autoencoders (DAEs) that perform dimensionality reduction and reconstruction tasks using the waveform obtained through one-level discrete wavelet transform, utilizing representation learning. Our results demonstrate that the latent representation generated by the DAE surpasses expectations with an accuracy of 0.93. The incorporation of noise-mixing and the utilization of a noise-to-clean scheme effectively enhance the discriminative capabilities of the latent representation. Moreover, when employed to identify patient-specific characteristics, the latent representation exhibited performance by surpassing an accuracy of 0.92. Appropriate light-weighted methods can restore the detection performance of the excessively reduced dimensionality version and enable operation on less computational devices. Our findings suggest that representation learning is a more feasible approach for extracting auscultation features in AVF, leading to improved generalization and applicability across multiple tasks. The manipulation of latent representations holds immense potential for future advancements. Further investigations in this area are promising and warrant continued exploration.

cs.LG

DREAM: A Dynamic Scheduler for Dynamic Real-time Multi-model ML Workloads

Emerging real-time multi-model ML (RTMM) workloads such as AR/VR and drone control involve dynamic behaviors in various granularity; task, model, and layers within a model. Such dynamic behaviors introduce new challenges to the system software in an ML system since the overall system load is not completely predictable, unlike traditional ML workloads. In addition, RTMM workloads require real-time processing, involve highly heterogeneous models, and target resource-constrained devices. Under such circumstances, developing an effective scheduler gains more importance to better utilize underlying hardware considering the unique characteristics of RTMM workloads. Therefore, we propose a new scheduler, DREAM, which effectively handles various dynamicity in RTMM workloads targeting multi-accelerator systems. DREAM quantifies the unique requirements for RTMM workloads and utilizes the quantified scores to drive scheduling decisions, considering the current system load and other inference jobs on different models and input frames. DREAM utilizes tunable parameters that provide fast and effective adaptivity to dynamic workload changes. In our evaluation of five scenarios of RTMM workload, DREAM reduces the overall UXCost, which is an equivalent metric of the energy-delay product (EDP) for RTMM defined in the paper, by 32.2% and 50.0% in the geometric mean (up to 80.8% and 97.6%) compared to state-of-the-art baselines, which shows the efficacy of our scheduling methodology.

cs.DC

Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation

Vision Transformers (ViTs) have emerged with superior performance on computer vision tasks compared to convolutional neural network (CNN)-based models. However, ViTs are mainly designed for image classification that generate single-scale low-resolution representations, which makes dense prediction tasks such as semantic segmentation challenging for ViTs. Therefore, we propose HRViT, which enhances ViTs to learn semantically-rich and spatially-precise multi-scale representations by integrating high-resolution multi-branch architectures with ViTs. We balance the model performance and efficiency of HRViT by various branch-block co-optimization techniques. Specifically, we explore heterogeneous branch designs, reduce the redundancy in linear layers, and augment the attention block with enhanced expressiveness. Those approaches enabled HRViT to push the Pareto frontier of performance and efficiency on semantic segmentation to a new level, as our evaluation results on ADE20K and Cityscapes show. HRViT achieves 50.20% mIoU on ADE20K and 83.16% mIoU on Cityscapes, surpassing state-of-the-art MiT and CSWin backbones with an average of +1.78 mIoU improvement, 28% parameter saving, and 21% FLOPs reduction, demonstrating the potential of HRViT as a strong vision backbone for semantic segmentation.

cs.CV

Efimov physics implications at $p$-wave fermionic unitarity

Efimov physics at $p$-wave unitarity for three equal mass fermions in multiple symmetries interacting via Lennard-Jones potentials is predicted to modify the long range interaction potential energy, but without producing a true Efimov effect. This analysis treats the following total orbital angular momenta and parities, $J^{\Pi}=0^{+}, 1^{+}, 1^{-}$ and $2^{-}$, for either three spin-polarized fermions ($\uparrow \uparrow \uparrow $), or two spin-up and one spin-down fermion ($\downarrow \uparrow \uparrow $). Our results for the long range interaction in some of those cases agree with previous work by Werner and Castin and by Blume {\it et al.}, namely in cases where the $s$-wave scattering length goes to infinity. The present results extend those calculated interaction energies to small and intermediate hyperradii comparable to the van der Waals length, and we consider additional unitarity scenarios where the $p$-wave scattering volume approaches infinity. The crucial role of the diagonal hyperradial adiabatic correction term is identified and characterized.

physics.atom-ph

Heterogeneous Dataflow Accelerators for Multi-DNN Workloads

Emerging AI-enabled applications such as augmented/virtual reality (AR/VR) leverage multiple deep neural network (DNN) models for sub-tasks such as object detection, hand tracking, and so on. Because of the diversity of the sub-tasks, the layers within and across the DNN models are highly heterogeneous in operation and shape. Such layer heterogeneity is a challenge for a fixed dataflow accelerator (FDA) that employs a fixed dataflow on a single accelerator substrate since each layer prefers different dataflows (computation order and parallelization) and tile sizes. Reconfigurable DNN accelerators (RDAs) have been proposed to adapt their dataflows to diverse layers to address the challenge. However, the dataflow flexibility in RDAs is enabled at the area and energy costs of expensive hardware structures (switches, controller, etc.) and per-layer reconfiguration. Alternatively, this work proposes a new class of accelerators, heterogeneous dataflow accelerators (HDAs), which deploys multiple sub-accelerators each supporting a different dataflow. HDAs enable coarser-grained dataflow flexibility than RDAs with higher energy efficiency and lower area cost comparable to FDAs. To exploit such benefits, hardware resource partitioning across sub-accelerators and layer execution schedule need to be carefully optimized. Therefore, we also present Herald, which co-optimizes hardware partitioning and layer execution schedule. Using Herald on a suite of AR/VR and MLPerf workloads, we identify a promising HDA architecture, Maelstrom, which demonstrates 65.3% lower latency and 5.0% lower energy than the best FDAs and 22.0% lower energy at the cost of 20.7% higher latency than a state-of-the-art RDA. The results suggest that HDA is an alternative class of Pareto-optimal accelerators to RDA with strength in energy, which can be a better choice than RDAs depending on the use cases.

cs.DC

Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices

A recent trend in DNN development is to extend the reach of deep learning applications to platforms that are more resource and energy constrained, e.g., mobile devices. These endeavors aim to reduce the DNN model size and improve the hardware processing efficiency, and have resulted in DNNs that are much more compact in their structures and/or have high data sparsity. These compact or sparse models are different from the traditional large ones in that there is much more variation in their layer shapes and sizes, and often require specialized hardware to exploit sparsity for performance improvement. Thus, many DNN accelerators designed for large DNNs do not perform well on these models. In this work, we present Eyeriss v2, a DNN accelerator architecture designed for running compact and sparse DNNs. To deal with the widely varying layer shapes and sizes, it introduces a highly flexible on-chip network, called hierarchical mesh, that can adapt to the different amounts of data reuse and bandwidth requirements of different data types, which improves the utilization of the computation resources. Furthermore, Eyeriss v2 can process sparse data directly in the compressed domain for both weights and activations, and therefore is able to improve both processing speed and energy efficiency with sparse models. Overall, with sparse MobileNet, Eyeriss v2 in a 65nm CMOS process achieves a throughput of 1470.6 inferences/sec and 2560.3 inferences/J at a batch size of 1, which is 12.6x faster and 2.5x more energy efficient than the original Eyeriss running MobileNet. We also present an analysis methodology called Eyexam that provides a systematic way of understanding the performance limits for DNN processors as a function of specific characteristics of the DNN model and accelerator design; it applies these characteristics as sequential steps to increasingly tighten the bound on the performance limits.

cs.DC

Efficient Processing of Deep Neural Networks: A Tutorial and Survey

Deep neural networks (DNNs) are currently widely used for many artificial intelligence (AI) applications including computer vision, speech recognition, and robotics. While DNNs deliver state-of-the-art accuracy on many AI tasks, it comes at the cost of high computational complexity. Accordingly, techniques that enable efficient processing of DNNs to improve energy efficiency and throughput without sacrificing application accuracy or increasing hardware cost are critical to the wide deployment of DNNs in AI systems. This article aims to provide a comprehensive tutorial and survey about the recent advances towards the goal of enabling efficient processing of DNNs. Specifically, it will provide an overview of DNNs, discuss various hardware platforms and architectures that support DNNs, and highlight key trends in reducing the computation cost of DNNs either solely via hardware design changes or via joint hardware design and DNN algorithm changes. It will also summarize various development resources that enable researchers and practitioners to quickly get started in this field, and highlight important benchmarking metrics and design considerations that should be used for evaluating the rapidly growing number of DNN hardware designs, optionally including algorithmic co-designs, being proposed in academia and industry. The reader will take away the following concepts from this article: understand the key design considerations for DNNs; be able to evaluate different DNN hardware implementations with benchmarks and comparison metrics; understand the trade-offs between various hardware architectures and platforms; be able to evaluate the utility of various DNN design techniques for efficient processing; and understand recent implementation trends and opportunities.

cs.CV

Towards Closing the Energy Gap Between HOG and CNN Features for Embedded Vision

Computer vision enables a wide range of applications in robotics/drones, self-driving cars, smart Internet of Things, and portable/wearable electronics. For many of these applications, local embedded processing is preferred due to privacy and/or latency concerns. Accordingly, energy-efficient embedded vision hardware delivering real-time and robust performance is crucial. While deep learning is gaining popularity in several computer vision algorithms, a significant energy consumption difference exists compared to traditional hand-crafted approaches. In this paper, we provide an in-depth analysis of the computation, energy and accuracy trade-offs between learned features such as deep Convolutional Neural Networks (CNN) and hand-crafted features such as Histogram of Oriented Gradients (HOG). This analysis is supported by measurements from two chips that implement these algorithms. Our goal is to understand the source of the energy discrepancy between the two approaches and to provide insight about the potential areas where CNNs can be improved and eventually approach the energy-efficiency of HOG while maintaining its outstanding performance accuracy.

cs.CV

Hardware for Machine Learning: Challenges and Opportunities

Machine learning plays a critical role in extracting meaningful information out of the zetabytes of sensor data collected every day. For some applications, the goal is to analyze and understand the data to identify trends (e.g., surveillance, portable/wearable electronics); in other applications, the goal is to take immediate action based the data (e.g., robotics/drones, self-driving cars, smart Internet of Things). For many of these applications, local embedded processing near the sensor is preferred over the cloud due to privacy or latency concerns, or limitations in the communication bandwidth. However, at the sensor there are often stringent constraints on energy consumption and cost in addition to throughput and accuracy requirements. Furthermore, flexibility is often required such that the processing can be adapted for different applications or environments (e.g., update the weights and model in the classifier). In many applications, machine learning often involves transforming the input data into a higher dimensional space, which, along with programmable weights, increases data movement and consequently energy consumption. In this paper, we will discuss how these challenges can be addressed at various levels of hardware design ranging from architecture, hardware-friendly algorithms, mixed-signal circuits, and advanced technologies (including memories and sensors).

cs.CV

Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning

Deep convolutional neural networks (CNNs) are indispensable to state-of-the-art computer vision algorithms. However, they are still rarely deployed on battery-powered mobile devices, such as smartphones and wearable gadgets, where vision algorithms can enable many revolutionary real-world applications. The key limiting factor is the high energy consumption of CNN processing due to its high computational complexity. While there are many previous efforts that try to reduce the CNN model size or amount of computation, we find that they do not necessarily result in lower energy consumption, and therefore do not serve as a good metric for energy cost estimation. To close the gap between CNN design and energy consumption optimization, we propose an energy-aware pruning algorithm for CNNs that directly uses energy consumption estimation of a CNN to guide the pruning process. The energy estimation methodology uses parameters extrapolated from actual hardware measurements that target realistic battery-powered system setups. The proposed layer-by-layer pruning algorithm also prunes more aggressively than previously proposed pruning methods by minimizing the error in output feature maps instead of filter weights. For each layer, the weights are first pruned and then locally fine-tuned with a closed-form least-square solution to quickly restore the accuracy. After all layers are pruned, the entire network is further globally fine-tuned using back-propagation. With the proposed pruning method, the energy consumption of AlexNet and GoogLeNet are reduced by 3.7x and 1.6x, respectively, with less than 1% top-5 accuracy loss. Finally, we show that pruning the AlexNet with a reduced number of target classes can greatly decrease the number of weights but the energy reduction is limited. Energy modeling tool and energy-aware pruned models available at http://eyeriss.mit.edu/energy.html

cs.CV