SearcharxivSearch

arXiv subjects

Xianmin Liu

Publications and source records attributed to Xianmin Liu.

4 recordsLinked to original sources

Knowledge-Aware Evolution for Task-Free Streaming Federated Continual Learning with Arbitrary Class Overlap

Federated Continual Learning (FCL) leverages inter-client collaboration to better balance new knowledge acquisition and old knowledge retention on non-stationary data. However, existing FCL methods struggle to adapt to streaming scenarios where sequential and ephemerally accessible data chunks lack task identifiers and exhibit arbitrary class overlap, leading to confusion between old and new knowledge and an inability to sustain local inference on all encountered classes. To address this, we propose FedKACE with three components: 1) an adaptive mechanism that determines when to switch the inference model from the local to the global one to improve client-side inference performance; 2) a responsive gradient-balanced replay scheme that utilizes the ratio of the squared L2 gradient norms to balance client-specific knowledge between new acquisition and old retention; 3) a holistic buffer maintenance strategy that preserves highly informative and boundary-significant samples to enhance knowledge retention under class overlap.Experiments across multiple scenarios and theoretical analysis demonstrate the effectiveness of FedKACE.

cs.LG

Personalized Federated Learning on Data with Dynamic Heterogeneity under Limited Storage

Recently, a large number of data sources opened up by informatization intensify the data heterogeneity, the faster speed of data generation and the gradual implementation of data regulations limit the storage time of data. In personalized Federated Learning (pFL), clients train customized models to meet their personal objectives. However, due to the time-varying local data heterogeneity and the inaccessibility of previous data, existing pFL methods not only fail to solve the catastrophic forgetting of local models, but also difficult to estimate the degree of collaboration between clients. To address this issue, our core idea is a low consumption and high-quality generative replay architecture. Specifically, we decouple the generator by category to reduce the generation error of each category while mitigating catastrophic forgetting, use local model to improving the quality of generated data and reducing the update frequency of generator, and propose a local data reconstruction scheme to reduce data generation while adjusting the proportion of data categories. Based on above, we propose our pFL framework, pFedGRP, to achieve personalized aggregation and local knowledge transfer. Comprehensive experiments on five datasets with multiple settings show the superiority of pFedGRP over eight baseline methods.

cs.DC

Complexity and Efficient Algorithms for Data Inconsistency Evaluating and Repairing

Data inconsistency evaluating and repairing are major concerns in data quality management. As the basic computing task, optimal subset repair is not only applied for cost estimation during the progress of database repairing, but also directly used to derive the evaluation of database inconsistency. Computing an optimal subset repair is to find a minimum tuple set from an inconsistent database whose remove results in a consistent subset left. Tight bound on the complexity and efficient algorithms are still unknown. In this paper, we improve the existing complexity and algorithmic results, together with a fast estimation on the size of optimal subset repair. We first strengthen the dichotomy for optimal subset repair computation problem, we show that it is not only APXcomplete, but also NPhard to approximate an optimal subset repair with a factor better than $17/16$ for most cases. We second show a $(2-0.5^{\tinyσ-1})$-approximation whenever given $σ$ functional dependencies, and a $(2-η_k+\frac{η_k}{k})$-approximation when an $η_k$-portion of tuples have the $k$-quasi-Tur$\acute{\text{a}}$n property for some $k>1$. We finally show a sublinear estimator on the size of optimal \textit{S}-repair for subset queries, it outputs an estimation of a ratio $2n+εn$ with a high probability, thus deriving an estimation of FD-inconsistency degree of a ratio $2+ε$. To support a variety of subset queries for FD-inconsistency evaluation, we unify them as the $\subseteq$-oracle which can answer membership-query, and return $p$ tuples uniformly sampled whenever given a number $p$. Experiments are conducted on range queries as an implementation of $\subseteq$-oracle, and results show the efficiency of our FD-inconsistency degree estimator.

cs.DB

Recognizing the Tractability in Big Data Computing

Due to the limitation on computational power of existing computers, the polynomial time does not works for identifying the tractable problems in big data computing. This paper adopts the sublinear time as the new tractable standard to recognize the tractability in big data computing, and the random-access Turing machine is used as the computational model to characterize the problems that are tractable on big data. First, two pure-tractable classes are first proposed. One is the class $\mathrm{PL}$ consisting of the problems that can be solved in polylogarithmic time by a RATM. The another one is the class $\mathrm{ST}$ including all the problems that can be solved in sublinear time by a RATM. The structure of the two pure-tractable classes is deeply investigated and they are proved $\mathrm{PL^i} \subsetneq \mathrm{PL^{i+1}}$ and $\mathrm{PL} \subsetneq \mathrm{ST}$. Then, two pseudo-tractable classes, $\mathrm{PTR}$ and $\mathrm{PTE}$, are proposed. $\mathrm{PTR}$ consists of all the problems that can solved by a RATM in sublinear time after a PTIME preprocessing by reducing the size of input dataset. $\mathrm{PTE}$ includes all the problems that can solved by a RATM in sublinear time after a PTIME preprocessing by extending the size of input dataset. The relations among the two pseudo-tractable classes and other complexity classes are investigated and they are proved that $\mathrm{PT} \subseteq \mathrm{P}$, $\sqcap'\mathrm{T^0_Q} \subsetneq \mathrm{PTR^0_Q}$ and $\mathrm{PT_P} = \mathrm{P}$.

cs.CC