Searcharxiv⌕ Search

arXiv subjects

Han Cai

Publications and source records attributed to Han Cai.

At least 73 records · Page 4Linked to original sources

EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction

High-resolution dense prediction enables many appealing real-world applications, such as computational photography, autonomous driving, etc. However, the vast computational cost makes deploying state-of-the-art high-resolution dense prediction models on hardware devices difficult. This work presents EfficientViT, a new family of high-resolution vision models with novel multi-scale linear attention. Unlike prior high-resolution dense prediction models that rely on heavy softmax attention, hardware-inefficient large-kernel convolution, or complicated topology structure to obtain good performances, our multi-scale linear attention achieves the global receptive field and multi-scale learning (two desirable features for high-resolution dense prediction) with only lightweight and hardware-efficient operations. As such, EfficientViT delivers remarkable performance gains over previous state-of-the-art models with significant speedup on diverse hardware platforms, including mobile CPU, edge GPU, and cloud GPU. Without performance loss on Cityscapes, our EfficientViT provides up to 13.9$\times$ and 6.2$\times$ GPU latency reduction over SegFormer and SegNeXt, respectively. For super-resolution, EfficientViT delivers up to 6.4x speedup over Restormer while providing 0.11dB gain in PSNR. For Segment Anything, EfficientViT delivers 48.9x higher throughput on A100 GPU while achieving slightly better zero-shot instance segmentation performance on COCO.

cs.CV↗

Construction of Locally Repairable Array Codes with Optimal Repair Bandwidth under the Rack-Aware Storage Model

In this paper, we discuss codes for distributed storage systems with hierarchical repair properties. Specifically, we devote attention to the repair problem of the rack-aware storage model with locality, aiming to enhance the system's ability to repair a small number of erasures within each rack by locality and efficiently handling a rack erasure with a small repair bandwidth. By employing the regenerating coding technique, we construct a family of array codes with $(r,u-r+1)$-locality, where the $u$ nodes of each repair set are systematically organized into a rack. When the number of failures is less than $u - r + 1$, these failures can be repaired without counting the system bandwidth. In cases where the number of failures exceeds the locality, the failed nodes within a single rack can be recovered with optimal cross-rack bandwidth.

cs.IT↗

Optically Levitated Nanoparticles as Receiving Antennas for Low Frequency Wireless Communication

Low-frequency (LF) wireless communications play a crucial role in ensuring anti-interference, long-range, and efficient communication across various environments. However, in conventional LF communication systems, their antenna size is required to be inversely proportional to the wavelength, so that their mobility and flexibility are greatly limited. Here we introduce a novel prototype of LF receiving antennas based on optically levitated nanoparticles, which overcomes the size-frequency limitation to reduce the antenna size to the hundred-nanometer scale. These charged particles are extremely sensitive to external electric field as mechanical resonators, and their resonant frequencies are adjustable. The effectiveness of these antennas was experimentally demonstrated by using the frequency shift keying (2FSK) modulation scheme. The experimental results indicate a correlation between error rate and factors such as transmission rate, signal strength, and vacuum degree with a signal strength of approximately 0.1V/m and a bit error rate below 0.1%. This advancement in leveraging levitated particle mechanical resonators (LPMRs) as LF antennas marks a significant stride in long-distance communication technology.

physics.app-ph↗

Realization of all-band-flat photonic lattices

Flatbands play an important role in correlated quantum matter and have novel applications in photonic lattices. Synthetic magnetic fields and destructive interference in lattices are traditionally used to obtain flatbands. However, such methods can only obtain a few flatbands with most bands remaining dispersive. Here we realize all-band-flat photonic lattices of an arbitrary size by precisely controlling the coupling strengths between lattice sites to mimic those in Fock-state lattices. This allows us to go beyond the perturbative regime of strain engineering and group all eigenmodes in flatbands, which simultaneously achieves high band flatness and large usable bandwidth. We map out the distribution of each flatband in the lattices and selectively excite the eigenmodes with different chiralities. Our method paves a new way in controlling band structure and topology of photonic lattices.

physics.optics↗

Robust Multi-Sensor Multi-Target Tracking Using Possibility Labeled Multi-Bernoulli Filter

With the increasing complexity of multiple target tracking scenes, a single sensor may not be able to effectively monitor a large number of targets. Therefore, it is imperative to extend the single-sensor technique to Multi-Sensor Multi-Target Tracking (MSMTT) for enhanced functionality. Typical MSMTT methods presume complete randomness of all uncertain components, and therefore effective solutions such as the random finite set filter and covariance intersection method have been derived to conduct the MSMTT task. However, the presence of epistemic uncertainty, arising from incomplete information, is often disregarded within the context of MSMTT. This paper develops an innovative possibility Labeled Multi-Bernoulli (LMB) Filter based on the labeled Uncertain Finite Set (UFS) theory. The LMB filter inherits the high robustness of the possibility generalized labeled multi-Bernoulli filter with simplified computational complexity. The fusion of LMB UFSs is derived and adapted to develop a robust MSMTT scheme. Simulation results corroborate the superior performance exhibited by the proposed approach in comparison to typical probabilistic methods.

cs.IT↗

High-Temperature Superconductor Quantum Flux Parametron for Energy-Efficient Logic

As we rapidly advance through the information age, the power consumed by computers, data centers, and networks grows exponentially. This has inspired a race to develop alternative low-power computational technologies. A new adiabatic configuration of a decades-old superconducting digital logic device has darted into the lead called quantum flux parametrons (QFP). QFP operate with dissipation so low that they seemingly violate the laws of thermodynamics. In just a short span of time, they have gone from simple single NOT gates to complex processors containing thousands of gates. They are fabricated from elemental niobium superconductors cooled to just a few degrees above absolute zero. However, their efficiency is so great that for large high-performance computers with several gates, the energy savings are immense. For smaller computational platforms QFPs from high-temperature superconductors (high-Tc) are highly desirable. In this work, we take the first steps towards this goal with the demonstration of a high-T C QFP shift register. Our device is fabricated using focused helium ion beam lithography where the material is modified with an ion beam at the nanoscale to directly pattern these circuits into a high-T C thin film. We validate the correct logical operation at 25 K, over 6 times higher than niobium devices with an estimated bit energy of 0.1 attoJoule at 10 GHz.

physics.app-ph↗

Floquet superradiance lattices in thermal atoms

Floquet modulation has been widely used in optical lattices for coherent control of quantum gases, in particular for synthesizing artificial gauge fields and simulating topological matters. However, such modulation induces heating which can overwhelm the signal of quantum dynamics in ultracold atoms. Here we report that the thermal motion, instead of being a noise source, provides a new control knob in Floquet-modulated superradiance lattices, which are momentum-space tight-binding lattices of collectively excited states of atoms. The Doppler shifts combined with Floquet modulation provide effective forces along arbitrary directions in a lattice in frequency and momentum dimensions. Dynamic localization, dynamic delocalization and chiral edge currents can be simultaneously observed from a single transport spectrum of superradiance lattices in thermal atoms. Our work paves a way for simulating Floquet topological matters in room-temperature atoms and facilitates their applications in photonic devices.

cond-mat.quant-gas↗

Measuring Zak phase in room-temperature atoms

Cold atoms provide a flexible platform for synthesizing and characterizing topolog-ical matter, where geometric phases play a central role. However, cold atoms are intrinsically prone to thermal noise, which can overwhelm the topological response and hamper promised applications. On the other hand, geometric phases also de-termine the energy spectra of particles subjected to a static force, based on the po-larization relation between Wannier-Stark ladders and geometric Zak phases. By exploiting this relation, we develop a method to extract geometric phases from en-ergy spectra of room-temperature superradiance lattices, which are momentum-space lattices of timed Dicke states. In such momentum-space lattices the thermal motion of atoms, instead of being a source of noise, provides effective forces which lead to spectroscopic signatures of the Zak phases. We measure Zak phases direct-ly from the anti-crossings between Wannier-Stark ladders in the Doppler-broadened absorption spectra of superradiance lattices. Our approach paves the way of measuring topological invariants and developing their applications in room-temperature atoms.

cond-mat.quant-gas↗

A Bound on the Minimal Field Size of LRCs, and Cyclic MR Codes That Attain It

We prove a new lower bound on the field size of locally repairable codes (LRCs). Additionally, we construct maximally recoverable (MR) codes which are cyclic. While a known construction for MR codes has the same parameters, it produces non-cyclic codes. Furthermore, we prove both necessary conditions and sufficient conditions that specify when the known non-cyclic MR codes may be permuted to become cyclic, thus proving our construction produces cyclic MR codes with new parameters. Furthermore, using our new bound on the field size, we show that the new cyclic MR codes have optimal field size in certain cases. Other known LRCs are also shown to have optimal field size in certain cases.

cs.IT↗

Coherent control of quantum topological states of light in Fock-state lattices

Topological photonics provides a novel platform to explore topological physics beyond traditional electronic materials and stimulates promising applications in topologically protected light transport and lasers. Classical degrees of freedom such as polarizations and wavevectors are routinely used to synthesize topological light modes. Beyond the classical regime, inherent quantum nature of light gives birth to a wealth of fundamentally distinct topological states, which offer topological protection in quantum information processing. Here we implement such experiments on topological states of quantized light in a superconducting circuit, on which three resonators are tunably coupled to a gmon qubit. We construct one and two-dimensional Fock-state lattices where topological transport of zero-energy states, strain induced pseudo-Landau levels, valley Hall effect and Haldane chiral edge currents are demonstrated. Our study extends the topological states of light to the quantum regime, bridges topological phases of condensed matter physics with circuit quantum electrodynamics, and offers a new freedom in controlling the quantum states of multiple resonators.

quant-ph↗

Lite Pose: Efficient Architecture Design for 2D Human Pose Estimation

Pose estimation plays a critical role in human-centered vision applications. However, it is difficult to deploy state-of-the-art HRNet-based pose estimation models on resource-constrained edge devices due to the high computational cost (more than 150 GMACs per frame). In this paper, we study efficient architecture design for real-time multi-person pose estimation on edge. We reveal that HRNet's high-resolution branches are redundant for models at the low-computation region via our gradual shrinking experiments. Removing them improves both efficiency and performance. Inspired by this finding, we design LitePose, an efficient single-branch architecture for pose estimation, and introduce two simple approaches to enhance the capacity of LitePose, including Fusion Deconv Head and Large Kernel Convs. Fusion Deconv Head removes the redundancy in high-resolution branches, allowing scale-aware feature fusion with low overhead. Large Kernel Convs significantly improve the model's capacity and receptive field while maintaining a low computational cost. With only 25% computation increment, 7x7 kernels achieve +14.0 mAP better than 3x3 kernels on the CrowdPose dataset. On mobile platforms, LitePose reduces the latency by up to 5.0x without sacrificing performance, compared with prior state-of-the-art efficient pose estimation models, pushing the frontier of real-time multi-person pose estimation on edge. Our code and pre-trained models are released at https://github.com/mit-han-lab/litepose.

cs.CV↗

Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications

Deep neural networks (DNNs) have achieved unprecedented success in the field of artificial intelligence (AI), including computer vision, natural language processing and speech recognition. However, their superior performance comes at the considerable cost of computational complexity, which greatly hinders their applications in many resource-constrained devices, such as mobile phones and Internet of Things (IoT) devices. Therefore, methods and techniques that are able to lift the efficiency bottleneck while preserving the high accuracy of DNNs are in great demand in order to enable numerous edge AI applications. This paper provides an overview of efficient deep learning methods, systems and applications. We start from introducing popular model compression methods, including pruning, factorization, quantization as well as compact model design. To reduce the large design cost of these manual solutions, we discuss the AutoML framework for each of them, such as neural architecture search (NAS) and automated pruning and quantization. We then cover efficient on-device training to enable user customization based on the local data on mobile devices. Apart from general acceleration techniques, we also showcase several task-specific accelerations for point cloud, video and natural language processing by exploiting their spatial sparsity and temporal/token redundancy. Finally, to support all these algorithmic advancements, we introduce the efficient deep learning system design from both software and hardware perspectives.

cs.LG↗

Network Augmentation for Tiny Deep Learning

We introduce Network Augmentation (NetAug), a new training method for improving the performance of tiny neural networks. Existing regularization techniques (e.g., data augmentation, dropout) have shown much success on large neural networks by adding noise to overcome over-fitting. However, we found these techniques hurt the performance of tiny neural networks. We argue that training tiny models are different from large models: rather than augmenting the data, we should augment the model, since tiny models tend to suffer from under-fitting rather than over-fitting due to limited capacity. To alleviate this issue, NetAug augments the network (reverse dropout) instead of inserting noise into the dataset or the network. It puts the tiny model into larger models and encourages it to work as a sub-model of larger models to get extra supervision, in addition to functioning as an independent model. At test time, only the tiny model is used for inference, incurring zero inference overhead. We demonstrate the effectiveness of NetAug on image classification and object detection. NetAug consistently improves the performance of tiny models, achieving up to 2.2% accuracy improvement on ImageNet. On object detection, achieving the same level of performance, NetAug requires 41% fewer MACs on Pascal VOC and 38% fewer MACs on COCO than the baseline.

cs.CV↗

A New Cooperative Repair Scheme with k + 1 Helper Nodes for (n, k) Hadamard MSR codes with Small Sub-packetization

Cooperative repair model is an available technology to deal with multiple node failures in distributed storage systems. Recently, explicit constructions of cooperative MSR codes were given by Ye (IEEE Transactions on Information Theory, 2020) with sub-packetization level $(d-k+h)(d-k+1)^n$. Specifically, the sub-packetization level is $(h+1)2^n$ when $d=k+1$. In this paper, we propose a new cooperative repair scheme by means of the inter-instance and intra-instance pairing inherited from the perfect code which reduces the sub-packetization to $2^n$ when $(h+1)|2^n$ and $(2\ell+1)2^n$ when $h+1=(2\ell+1)2^m$ for $m\ge 0$, $\ell\ge 1$ with $d=k+1$ helper nodes. That is to say, the sub-packetization is $h + 1 $ times or $2^m$ times less than Ye's. It turned out to be the best result so far known.

cs.IT↗

A Class of Minimum Storage Cooperative Regenerating Codes with Low Access Property

In this paper, a new repair scheme for a modified construction of MDS codes is studied. The obtained repair scheme has optimal bandwidth for multiple failed nodes under the cooperative repair model. In addition, the repair scheme has relatively low access property, where the number of data accessed is less than two times the optimal value.

cs.IT↗

A Construction of Maximally Recoverable Codes with Order-Optimal Field Size

We construct maximally recoverable codes (corresponding to partial MDS codes) which are based on linearized Reed-Solomon codes. The new codes have a smaller field size requirement compared with known constructions. For certain asymptotic regimes, the constructed codes have order-optimal alphabet size, asymptotically matching the known lower bound.

cs.IT↗

TinyTL: Reduce Activations, Not Trainable Parameters for Efficient On-Device Learning

On-device learning enables edge devices to continually adapt the AI models to new data, which requires a small memory footprint to fit the tight memory constraint of edge devices. Existing work solves this problem by reducing the number of trainable parameters. However, this doesn't directly translate to memory saving since the major bottleneck is the activations, not parameters. In this work, we present Tiny-Transfer-Learning (TinyTL) for memory-efficient on-device learning. TinyTL freezes the weights while only learns the bias modules, thus no need to store the intermediate activations. To maintain the adaptation capacity, we introduce a new memory-efficient bias module, the lite residual module, to refine the feature extractor by learning small residual feature maps adding only 3.8% memory overhead. Extensive experiments show that TinyTL significantly saves the memory (up to 6.5x) with little accuracy loss compared to fine-tuning the full network. Compared to fine-tuning the last layer, TinyTL provides significant accuracy improvements (up to 34.1%) with little memory overhead. Furthermore, combined with feature extractor adaptation, TinyTL provides 7.3-12.9x memory saving without sacrificing accuracy compared to fine-tuning the full Inception-V3.

cs.CV↗

Unification of valley and anomalous Hall effects in a strained lattice

Two dimensional lattices are an important stage for studying many aspects of quantum physics, in particular the topological phases. The valley Hall and anomalous Hall effects are two representative topological phenomena. Here we show that they can be unified in a strained honeycomb lattice, where the hopping strengths between neighboring sites are designed by mimicking those between the Fock states in a three-mode Jaynes-Cummings model. Such a strain induces an effective magnetic field which results in quantized Landau levels. The eigenstates in the zeroth Landau level can be represented by the eigenstates of a large pseudo-spin. We find that the valley Hall current and the chiral edge current in the Haldane model correspond to the spin precession around different axes. Our study sheds light on connection between seemingly unrelated topological phases in condensed matter physics.

cond-mat.mes-hall↗