SearcharxivSearch

arXiv subjects

Zhenlin Yang

Publications and source records attributed to Zhenlin Yang.

3 recordsLinked to original sources

Testing Clustered Equal Predictive Ability with Unknown Clusters

We develop tests of clustered equal predictive ability (C-EPA) in panels where the clusters are unknown and estimated by the Panel Kmeans algorithm. To address the challenge of testing hypotheses that depend on data-driven clusters, we adopt a selective conditional inference framework. Specifically, we first derive a Wald-type test for pairwise equality and show that the limiting distribution of its square root conditional on the estimated clusters is that of a truncated $\chi$ variable. We characterize the associated truncation set by quadratic inequalities in the data space. Then, for the C-EPA hypothesis, we propose a $p$-value combination method by aggregating the evidence against the pairwise equality and overall EPA null hypotheses. The Monte Carlo results show accurate size control and good finite-sample power of the proposed tests. An empirical application to exchange-rate forecasting, using both traditional time-series models and machine-learning methods, illustrates the practical relevance of our procedure.

econ.EM

Taming the Titans: A Survey of Efficient LLM Inference Serving

Large Language Models (LLMs) for Generative AI have achieved remarkable progress, evolving into sophisticated and versatile tools widely adopted across various domains and applications. However, the substantial memory overhead caused by their vast number of parameters, combined with the high computational demands of the attention mechanism, poses significant challenges in achieving low latency and high throughput for LLM inference services. Recent advancements, driven by groundbreaking research, have significantly accelerated progress in this field. This paper provides a comprehensive survey of these methods, covering fundamental instance-level approaches, in-depth cluster-level strategies, emerging scenario directions, and other miscellaneous but important areas. At the instance level, we review model placement, request scheduling, decoding length prediction, storage management, and the disaggregation paradigm. At the cluster level, we explore GPU cluster deployment, multi-instance load balancing, and cloud service solutions. For emerging scenarios, we organize the discussion around specific tasks, modules, and auxiliary methods. To ensure a holistic overview, we also highlight several niche yet critical areas. Finally, we outline potential research directions to further advance the field of LLM inference serving.

cs.CL

Equal Predictive Ability Tests Based on Panel Data with Applications to OECD and IMF Forecasts

We propose two types of equal predictive ability (EPA) tests with panels to compare the predictions made by two forecasters. The first type, namely $S$-statistics, focuses on the overall EPA hypothesis which states that the EPA holds on average over all panel units and over time. The second, called $C$-statistics, focuses on the clustered EPA hypothesis where the EPA holds jointly for a fixed number of clusters of panel units. The asymptotic properties of the proposed tests are evaluated under weak and strong cross-sectional dependence. An extensive Monte Carlo simulation shows that the proposed tests have very good finite sample properties even with little information about the cross-sectional dependence in the data. The proposed framework is applied to compare the economic growth forecasts of the OECD and the IMF, and to evaluate the performance of the consumer price inflation forecasts of the IMF.

econ.EM