SearcharxivSearch

arXiv subjects

Wei-Po Wang

Publications and source records attributed to Wei-Po Wang.

3 recordsLinked to original sources

Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency

We investigate the statistical and computational limits of prompt tuning for transformer-based foundation models. Our key contributions are prompt tuning on \emph{single-head} transformers with only a \emph{single} self-attention layer: (i) is universal, and (ii) supports efficient (even almost-linear time) algorithms under the Strong Exponential Time Hypothesis (SETH). Statistically, we prove that prompt tuning on such simplest possible transformers are universal approximators for sequence-to-sequence Lipschitz functions. In addition, we provide an exponential-in-$dL$ and -in-$(1/ε)$ lower bound on the required soft-prompt tokens for prompt tuning to memorize any dataset with 1-layer, 1-head transformers. Computationally, we identify a phase transition in the efficiency of prompt tuning, determined by the norm of the \emph{soft-prompt-induced} keys and queries, and provide an upper bound criterion. Beyond this criterion, no sub-quadratic (efficient) algorithm for prompt tuning exists under SETH. Within this criterion, we showcase our theory by proving the existence of almost-linear time prompt tuning inference algorithms. These fundamental limits provide important necessary conditions for designing expressive and efficient prompt tuning methods for practitioners.

cs.LG

Outlier-Efficient Hopfield Layers for Large Transformer-Based Models

We introduce an Outlier-Efficient Modern Hopfield Model (termed $\mathrm{OutEffHop}$) and use it to address the outlier inefficiency problem of {training} gigantic transformer-based models. Our main contribution is a novel associative memory model facilitating \textit{outlier-efficient} associative memory retrievals. Interestingly, this memory model manifests a model-based interpretation of an outlier-efficient attention mechanism (${\rm Softmax}_1$): it is an approximation of the memory retrieval process of $\mathrm{OutEffHop}$. Methodologically, this allows us to introduce novel outlier-efficient Hopfield layers as powerful alternatives to traditional attention mechanisms, with superior post-quantization performance. Theoretically, the Outlier-Efficient Modern Hopfield Model retains and improves the desirable properties of standard modern Hopfield models, including fixed point convergence and exponential storage capacity. Empirically, we demonstrate the efficacy of the proposed model across large-scale transformer-based and Hopfield-based models (including BERT, OPT, ViT, and STanHop-Net), benchmarking against state-of-the-art methods like $\mathtt{Clipped\_Softmax}$ and $\mathtt{Gated\_Attention}$. Notably, $\mathrm{OutEffHop}$ achieves an average reduction of 22+\% in average kurtosis and 26+\% in the maximum infinity norm of model outputs across four models. Code is available at \href{https://github.com/MAGICS-LAB/OutEffHop}{GitHub}; models are on \href{https://huggingface.co/collections/magicslabnu/outeffhop-6610fcede8d2cda23009a98f}{Hugging Face Hub}; future updates are on \href{https://arxiv.org/abs/2404.03828}{arXiv}.

cs.LG

AnaBHEL (Analog Black Hole Evaporation via Lasers) Experiment: Concept, Design, and Status

Accelerating relativistic mirror has long been recognized as a viable setting where the physics mimics that of black hole Hawking radiation. In 2017, Chen and Mourou proposed a novel method to realize such a system by traversing an ultra-intense laser through a plasma target with a decreasing density. An international AnaBHEL (Analog Black Hole Evaporation via Lasers) Collaboration has been formed with the objectives of observing the analog Hawking radiation and shedding light on the information loss paradox. To reach these goals, we plan to first verify the dynamics of the flying plasma mirror and to characterize the correspondence between the plasma density gradient and the trajectory of the accelerating plasma mirror. We will then attempt to detect the analog Hawking radiation photons and measure the entanglement between the Hawking photons and their "partner particles". In this paper, we describe our vision and strategy of AnaBHEL using the Apollon laser as a reference, and we report on the progress of our R&D of the key components in this experiment, including the supersonic gas jet with a graded density profile, and the superconducting nanowire single-photon Hawking detector. In parallel to these hardware efforts, we performed computer simulations to estimate the potential backgrounds, and derive analytic expressions for modifications to the blackbody spectrum of Hawking radiation for a perfectly reflecting, point mirror, due to the semit-ransparency and finite-size effects specific to flying plasma mirrors. Based on this more realistic radiation spectrum, we estimate the Hawking photon yield to guide the design of the AnaBHEL experiment, which appears to be achievable.

gr-qc