SearcharxivSearch

arXiv subjects

Nan Zou

Publications and source records attributed to Nan Zou.

13 recordsLinked to original sources

HiLoRA: Hierarchical Low-Rank Adaptation for Personalized Federated Learning

Vision Transformers (ViTs) have been widely adopted in vision tasks due to their strong transferability. In Federated Learning (FL), where full fine-tuning is communication heavy, Low-Rank Adaptation (LoRA) provides an efficient and communication-friendly way to adapt ViTs. However, existing LoRA-based federated tuning methods overlook latent client structures in real-world settings, limiting shared representation learning and hindering effective adaptation to unseen clients. To address this, we propose HiLoRA, a hierarchical LoRA framework that places adapters at three levels: root, cluster, and leaf, each designed to capture global, subgroup, and client-specific knowledge, respectively. Through cross-tier orthogonality and cascaded optimization, HiLoRA separates update subspaces and aligns each tier with its residual personalized objective. In particular, we develop a LoRA-Subspace Adaptive Clustering mechanism that infers latent client groups via subspace similarity analysis, thereby facilitating knowledge sharing across structurally aligned clients. Theoretically, we establish a tier-wise generalization analysis that supports HiLoRA's design. Experiments on ViT backbones with CIFAR-100 and DomainNet demonstrate consistent improvements in both personalization and generalization.

cs.CV

On the Distributed Estimation for Scalar-on-Function Regression Models

This paper proposes distributed estimation procedures for three scalar-on-function regression models: the functional linear model (FLM), the functional non-parametric model (FNPM), and the functional partial linear model (FPLM). The framework addresses two key challenges in functional data analysis, namely the high computational cost of large samples and limitations on sharing raw data across institutions. Monte Carlo simulations show that the distributed estimators substantially reduce computation time while preserving high estimation and prediction accuracy for all three models. When block sizes become too small, the FPLM exhibits overfitting, leading to narrower prediction intervals and reduced empirical coverage probability. An example of an empirical study using the \textit{tecator} dataset further supports these findings.

stat.CO

Serving Large Language Models on Huawei CloudMatrix384

The rapid evolution of large language models (LLMs), driven by growing parameter scales, adoption of mixture-of-experts (MoE) architectures, and expanding context lengths, imposes unprecedented demands on AI infrastructure. Traditional AI clusters face limitations in compute intensity, memory bandwidth, inter-chip communication, and latency, compounded by variable workloads and strict service-level objectives. Addressing these issues requires fundamentally redesigned hardware-software integration. This paper introduces Huawei CloudMatrix, a next-generation AI datacenter architecture, realized in the production-grade CloudMatrix384 supernode. It integrates 384 Ascend 910 NPUs and 192 Kunpeng CPUs interconnected via an ultra-high-bandwidth Unified Bus (UB) network, enabling direct all-to-all communication and dynamic pooling of resources. These features optimize performance for communication-intensive operations, such as large-scale MoE expert parallelism and distributed key-value cache access. To fully leverage CloudMatrix384, we propose CloudMatrix-Infer, an advanced LLM serving solution incorporating three core innovations: a peer-to-peer serving architecture that independently scales prefill, decode, and caching; a large-scale expert parallelism strategy supporting EP320 via efficient UB-based token dispatch; and hardware-aware optimizations including specialized operators, microbatch-based pipelining, and INT8 quantization. Evaluation with the DeepSeek-R1 model shows CloudMatrix-Infer achieves state-of-the-art efficiency: prefill throughput of 6,688 tokens/s per NPU and decode throughput of 1,943 tokens/s per NPU (<50 ms TPOT). It effectively balances throughput and latency, sustaining 538 tokens/s per NPU even under stringent 15 ms latency constraints, while INT8 quantization maintains model accuracy across benchmarks.

cs.DC

On a penalised likelihood approach for joint modelling of longitudinal covariates and partly interval-censored data -- an application to the Anti-PD1 brain collaboration trial

This article considers the joint modeling of longitudinal covariates and partly-interval censored time-to-event data. Longitudinal time-varying covariates play a crucial role in obtaining accurate clinically relevant predictions using a survival regression model. However, these covariates are often measured at limited time points and may be subject to measurement error. Further methodological challenges arise from the fact that, in many clinical studies, the event times of interest are interval-censored. A model that simultaneously accounts for all these factors is expected to improve the accuracy of survival model estimations and predictions. In this article, we consider joint models that combine longitudinal time-varying covariates with the Cox model for time-to-event data which is subject to interval censoring. The proposed model employs a novel penalised likelihood approach for estimating all parameters, including the random effects. The covariance matrix of the estimated parameters can be obtained from the penalised log-likelihood. The performance of the model is compared to an existing method under various scenarios. The simulation results demonstrated that our new method can provide reliable inferences when dealing with interval-censored data. Data from the Anti-PD1 brain collaboration clinical trial in advanced melanoma is used to illustrate the application of the new method.

stat.ME

Estimating POT Second-order Parameter for Bias Correction

The stable tail dependence function provides a full characterization of the extremal dependence structures. Unfortunately, the estimation of the stable tail dependence function often suffers from significant bias, whose scale relates to the Peaks-Over-Threshold (POT) second-order parameter. For this second-order parameter, this paper introduces a penalized estimator that discourages it from being too close to zero. This paper then establishes this estimator's asymptotic consistency, uses it to correct the bias in the estimation of the stable tail dependence function, and illustrates its desirable empirical properties in the estimation of the extremal dependence structures.

stat.ME

POT-flavored estimator of Pickands dependence function

This work proposes an estimator with both Peak-Over-Threshold and Block-Maxima flavors, uses it to estimate the Pickands dependence function of bivariate time series, and illustrates how it brings down the asymptotic bias and the overall mean squared error.

stat.ME

The Bootstrap for Dynamical Systems

Despite their deterministic nature, dynamical systems often exhibit seemingly random behaviour. Consequently, a dynamical system is usually represented by a probabilistic model of which the unknown parameters must be estimated using statistical methods. When measuring the uncertainty of such parameter estimation, the bootstrap stands out as a simple but powerful technique. In this paper, we develop the bootstrap for dynamical systems and establish not only its consistency but also its second-order efficiency via a novel \textit{continuous} Edgeworth expansion for dynamical systems. This is the first time such continuous Edgeworth expansions have been studied. Moreover, we verify the theoretical results about the bootstrap using computer simulations.

math.DS

Bootstrap Seasonal Unit Root Test under Periodic Variation

Both seasonal unit roots and periodic variation can be prevalent in seasonal data. When testing seasonal unit roots under periodic variation, the validity of the existing methods, such as the HEGY test, remains unknown. This paper analyzes the behavior of the augmented HEGY test and the unaugmented HEGY test under periodic variation. It turns out that the asymptotic null distributions of the HEGY statistics testing the single roots at $1$ or $-1$ when there is periodic variation are identical to the asymptotic null distributions when there is no periodic variation. On the other hand, the asymptotic null distributions of the statistics testing any coexistence of roots at $1$, $-1$, $i$, or $-i$ when there is periodic variation are non-standard and are different from the asymptotic null distributions when there is no periodic variation. Therefore, when periodic variation exists, HEGY tests are not directly applicable to the joint tests for any concurrence of seasonal unit roots. As a remedy, bootstrap is proposed; in particular, the augmented HEGY test with seasonal independent and identically distributed (iid) bootstrap and the unaugmented HEGY test with seasonal block bootstrap are implemented. The consistency of these bootstrap procedures is established. The finite-sample behavior of these bootstrap tests is illustrated via simulation and prevails over their competitors'. Finally, these bootstrap tests are applied to detect the seasonal unit roots in various economic time series.

stat.ME

Multiple block sizes and overlapping blocks for multivariate time series extremes

Block maxima methods constitute a fundamental part of the statistical toolbox in extreme value analysis. However, most of the corresponding theory is derived under the simplifying assumption that block maxima are independent observations from a genuine extreme value distribution. In practice however, block sizes are finite and observations from different blocks are dependent. Theory respecting the latter complications is not well developed, and, in the multivariate case, has only recently been established for disjoint blocks of a single block size. We show that using overlapping blocks instead of disjoint blocks leads to a uniform improvement in the asymptotic variance of the multivariate empirical distribution function of rescaled block maxima and any smooth functionals thereof (such as the empirical copula), without any sacrifice in the asymptotic bias. We further derive functional central limit theorems for multivariate empirical distribution functions and empirical copulas that are uniform in the block size parameter, which seems to be the first result of this kind for estimators based on block maxima in general. The theory allows for various aggregation schemes over multiple block sizes, leading to substantial improvements over the single block length case and opens the door to further methodology developments. In particular, we consider bias correction procedures that can improve the convergence rates of extreme-value estimators and shed some new light on estimation of the second-order parameter when the main purpose is bias correction.

math.ST

On Second Order Conditions in the Multivariate Block Maxima and Peak over Threshold Method

Second order conditions provide a natural framework for establishing asymptotic results about estimators for tail related quantities. Such conditions are typically tailored to the estimation principle at hand, and may be vastly different for estimators based on the block maxima (BM) method or the peak-over-threshold (POT) approach. In this paper we provide details on the relationship between typical second order conditions for BM and POT methods in the multivariate case. We show that the two conditions typically imply each other, but with a possibly different second order parameter. The latter implies that, depending on the data generating process, one of the two methods can attain faster convergence rates than the other. The class of multivariate Archimax copulas is examined in detail; we find that this class contains models for which the second order parameter is smaller for the BM method and vice versa. The theory is illustrated by a small simulation study.

math.ST

First-principles investigation on diffusion mechanism of alloying elements in dilute Zr alloys

Impurity diffusion in Zr is potentially important for many applications of Zr alloys, and in particular for their use of nuclear reactor cladding. However, significant uncertainty presently exists about which elements are vacancy vs. interstitial diffusers, which can inhibit understanding and prediction of their behavior under different temperature, irradiation, and alloying conditions. Therefore, first-principles calculations based on density functional theory (DFT) have been employed to predict the temperature-dependent dilute impurity diffusion coefficients for 14 substitutional alloying elements in hexagonal closed packed (HCP) Zr. Vacancy-mediated diffusion was modeled with the eight-frequency model. Interstitial contributions to diffusion are estimated from interstitial formation and select migration energies. Formation energies for each impurity in nine high-symmetry interstitial sites were determined, including significant effects of thermal expansion. The dominant diffusion mechanism of each solute in HCP Zr was identified in terms of the calculated vacancy-mediated activation energy, lower and upper bounds of interstitial activation energy, and the formation entropy, suggesting a rough relation with the metallic radii of solutes. It is predicted that Cr, Cu, V, Zn, Mo, W, Au, Ag, Al, Nb, Ta and Ti all diffuse predominantly by an interstitial mechanism, while Hf, Zr, and Sn are likely to be predominantly vacancy-mediated diffusers at low temperature and interstitial diffusers at high temperature, although the identification of mechanisms for these elements at high-temperature is quite uncertain.

cond-mat.mtrl-sci

Linear Process Bootstrap Unit Root Test

One of the most widely applied unit root test, Phillips-Perron test, enjoys in general highpowers, but suffers from size distortions when moving average noise exists. As a remedy, thispaper proposes a nonparametric bootstrap unit root test that specifically targets moving aver-age noise. Via a bootstrap functional central limit theorem, the consistency of this bootstrapapproach is established under general assumptions which allows a large family of non-linear timeseries. In simulation, this bootstrap test alleviates the size distortions of the Phillips-Perrontest while preserving its high powers.

stat.ME

On spurious regressions with trending variables

This paper examines three types of spurious regressions where both the dependent and independent variables contain deterministic trends, stochastic trends, or breaking trends. We show that the problem of spurious regression disappears if the trend functions are included as additional regressors. In the presence of autocorrelation, we show that using a Feasible General Least Square (FGLS) estimator can help alleviate or eliminate the problem. Our theoretical results are clearly reflected in finite samples. As an illustration, we apply our methods to revisit the seminal study of Yule (1926).

stat.ME