SearcharxivSearch

arXiv subjects

Tong Pu

Publications and source records attributed to Tong Pu.

7 recordsLinked to original sources

ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity

Transformers, composed of QKV generation, attention computation, and FFNs, have become the dominant model across various domains due to their outstanding performance. However, their high computational cost hinders efficient hardware deployment. Sparsity offers a promising solution, yet most existing accelerators exploit only intra-row sparsity in attention, while few consider inter-row sparsity. Approaches leveraging inter-row sparsity often rely on costly global similarity estimation, which diminishes the acceleration benefits of sparsity, and typically apply sparsity to only one or two transformer components. Through careful analysis of the attention distribution and computation flow, we observe that local similarity allows end-to-end sparse acceleration with lower computational overhead. Motivated by this observation, we propose ESACT, an end-to-end sparse accelerator for compute-intensive Transformers. ESACT centers on the Sparsity Prediction with Local Similarity (SPLS) mechanism, which leverages HLog quantization to accurately predict local attention sparsity prior to QK generation, achieving efficient sparsity across all transformer components. To support efficient hardware realization, we introduce three architectural innovations. Experimental results on 26 benchmarks demonstrate that SPLS reduces total computation by 52.03% with less than 1% accuracy loss. ESACT achieves an end-to-end energy efficiency of 3.29 TOPS/W, and improves attention-level energy efficiency by 2.95x and 2.26x over SOTA attention accelerators SpAtten and Sanger, respectively.

cs.LG

On multivariate contribution measures of systemic risk with applications in cryptocurrency market

Conditional risk measures and their associated risk contribution measures are commonly employed in finance and actuarial science for evaluating systemic risk and quantifying the effects of risk interactions. This paper introduces various types of contribution ratio measures based on the MCoVaR, MCoES, and MMME studied in Ortega-Jim\'enez et al. (2021) and Das & Fasen-Hartmann (2018) to assess the relative effects of a single risk when other risks in a group are in distress. The properties of these contribution risk measures are examined, and sufficient conditions for comparing these measures between two sets of random vectors are established using univariate and multivariate stochastic orders and statistically dependent notions. Numerical examples are presented to validate these conditions. Finally, a real dataset from the cryptocurrency market is used to analyze the spillover effects through our proposed contribution measures.

q-fin.RM

On Vulnerability Conditional Risk Measures: Comparisons and Applications in Cryptocurrency Market

We introduce a novel class of systemic risk measures, the Vulnerability Conditional risk measures, which try to capture the "tail risk" of a risky position in scenarios where one or more market participants is experiencing financial distress. Various theoretical properties of Vulnerability Conditional risk measures, along with a series of related contribution measures, have been considered in this paper. We further introduce the backtesting procedures of VCoES and MCoES. Through numerical examples, we validate our theoretical insights and further apply our newly proposed risk measures to the empirical analysis of cryptocurrencies, demonstrating their practical relevance and utility in capturing systemic risk.

q-fin.RM

On Joint Marginal Expected Shortfall and Associated Contribution Risk Measures

Systemic risk is the risk that a company- or industry-level risk could trigger a huge collapse of another or even the whole institution. Various systemic risk measures have been proposed in the literature to quantify the domino and (relative) spillover effects induced by systemic risks such as the well-known CoVaR, CoES, MES and CoD risk measures, and associated contribution measures. This paper proposes another new type of systemic risk measure, called the joint marginal expected shortfall (JMES), to measure whether the MES of one entity's risk-taking adds to another one or the overall risk conditioned on the event that the entity is already in some specified distress level. We further introduce two useful systemic risk contribution measures based on the difference function or relative ratio function of the JMES and the conventional ES, respectively. Some basic properties of these proposed measures are studied such as monotonicity, comonotonic additivity, non-identifiability and non-elicitability. For both risk measures and two different vectors of bivariate risks, we establish sufficient conditions imposed on copula structure, stress levels, and stochastic orders to compare these new measures. We further provide some numerical examples to illustrate our main findings. A real application in analyzing the risk contagion among several stock market indices is implemented to show the performances of our proposed measures compared with other commonly used measures including CoVaR, CoES, MES, and their associated contribution measures.

q-fin.RM

FGraDA: A Dataset and Benchmark for Fine-Grained Domain Adaptation in Machine Translation

Previous research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real-world scenarios. One representative of such challenging scenarios is to deploy a translation system for a conference with a specific topic, e.g., global warming or coronavirus, where there are usually extremely less resources due to the limited schedule. To motivate wider investigation in such a scenario, we present a real-world fine-grained domain adaptation task in machine translation (FGraDA). The FGraDA dataset consists of Chinese-English translation task for four sub-domains of information technology: autonomous vehicles, AI education, real-time networks, and smart phone. Each sub-domain is equipped with a development set and test set for evaluation purposes. To be closer to reality, FGraDA does not employ any in-domain bilingual training data but provides bilingual dictionaries and wiki knowledge base, which can be easier obtained within a short time. We benchmark the fine-grained domain adaptation task and present in-depth analyses showing that there are still challenging problems to further improve the performance with heterogeneous resources.

cs.CL

Generalized Location-Scale Mixtures of Elliptical Distributions: Definitions and Stochastic Comparisons

This paper proposes a unified class of generalized location-scale mixture of multivariate elliptical distributions and studies integral stochastic orderings of random vectors following such distributions. Given a random vector $\boldsymbol{Z}$, independent of $\boldsymbol{X}$ and $\boldsymbol{Y}$, the scale parameter of this class of distributions is mixed with a function $\alpha(\boldsymbol{Z})$ and its skew parameter is mixed with another function $\beta(\boldsymbol{Z})$. Sufficient (and necessary) conditions are established for stochastically comparing different random vectors stemming from this class of distributions by means of several stochastic orders including the usual stochastic order, convex order, increasing convex order, supermodular order, and some related linear orders. Two insightful assumptions for the density generators of elliptical distributions, aiming to control the generators' tail, are provided to make stochastic comparisons among mixed-elliptical vectors. Some applications in applied probability and actuarial science are also provided as illustrations on the main findings.

math.ST

An Identity for Expectations and Characteristic Function of Matrix Variate Skew-normalDistribution with Applications to Associated Stochastic Orderings

We establish an identity for E f (Y) -E f (X), when X and Y both have matrix variateskew-normal distributions and the function f fulfills some weak conditions. Thecharacteristic function of matrix variate skew normal distribution is then derived. Finally,we make use of it to derive some necessary and sucient conditions for the comparisonof matrix variate skew-normal distributions under six di erent orders, such as usualstochastic order, convex order, increasing convex order, upper orthant order,directionally convex order and supermodular order.

math.ST