SearcharxivSearch

arXiv subjects

Xiaobo Zhao

Publications and source records attributed to Xiaobo Zhao.

8 recordsLinked to original sources

MUFFLe: Efficient Model Update Compression via Generalized Deduplication for Federated Learning

Federated learning is well suited to edge environments but is often limited by the uplink cost of transmitting model updates. This Work-in-Progress paper presents MUFFLe, a communication-efficient update compression scheme that integrates generalized deduplication (GD) into the FedAvg pipeline. MUFFLe deduplicates repeated patterns across the update vector, yielding a fixed-rate, variable-count compression scheme. Preliminary experiments on IID MNIST with 20 clients show that MUFFLe reaches the target accuracy of $92.93\%$ with 38~MB cumulative uplink communication, compared with 75~MB for 8-bit quantization, 86~MB for Top-$k$ sparsification, and 310~MB for uncompressed FedAvg. These results demonstrate the feasibility of applying GD to communication-efficient federated learning.

cs.LG

EntroGD: Scalable Generalized Deduplication for Efficient Direct Analytics on Compressed IoT Data

Massive data streams from IoT and cyber-physical systems must be processed under strict bandwidth, latency, and resource constraints. Generalized Deduplication (GD) is a promising lossless compression framework, as it supports random access and direct analytics on compressed data. However, existing GD algorithms exhibit quadratic complexity $\mathcal{O}(nd^{2})$, which limits their scalability for high-dimensional datasets. This paper proposes \textbf{EntroGD}, an entropy-guided GD framework that decouples analytical fidelity from compression efficiency to achieve linear complexity $\mathcal{O}(nd)$. EntroGD adopts a two-stage design, first constructing compact condensed samples to preserve information critical for analytics, and then applying entropy-based bit selection to maximize compression. Experiments on 18 IoT datasets show that EntroGD reduces configuration time by up to $53.5\times$ compared to state-of-the-art GD compressors. Moreover, by enabling analytics with access to only $2.6\%$ of the original data volume, EntroGD accelerates clustering by up to $31.6\times$ with negligible loss in accuracy. Overall, EntroGD provides a scalable and system-efficient solution for direct analytics on compressed IoT data.

cs.DB

Data-driven Piecewise Affine Decision Rules for Stochastic Programming with Covariate Information

Focusing on stochastic programming (SP) with covariate information, this paper proposes an empirical risk minimization (ERM) method embedded within a nonconvex piecewise affine decision rule (PADR), which aims to learn the direct mapping from features to optimal decisions. We establish the nonasymptotic consistency result of our PADR-based ERM model for unconstrained problems and asymptotic consistency result for constrained ones. To solve the nonconvex and nondifferentiable ERM problem, we develop an enhanced stochastic majorization-minimization algorithm and establish the asymptotic convergence to (composite strong) directional stationarity along with complexity analysis. We show that the proposed PADR-based ERM method applies to a broad class of nonconvex SP problems with theoretical consistency guarantees and computational tractability. Our numerical study demonstrates the superior performance of PADR-based ERM methods compared to state-of-the-art approaches under various settings, with significantly lower costs, less computation time, and robustness to feature dimensions and nonlinearity of the underlying dependency.

math.OC

dreaMLearning: Data Compression Assisted Machine Learning

Despite rapid advancements, machine learning, particularly deep learning, is hindered by the need for large amounts of labeled data to learn meaningful patterns without overfitting and immense demands for computation and storage, which motivate research into architectures that can achieve good performance with fewer resources. This paper introduces dreaMLearning, a novel framework that enables learning from compressed data without decompression, built upon Entropy-based Generalized Deduplication (EntroGeDe), an entropy-driven lossless compression method that consolidates information into a compact set of representative samples. DreaMLearning accommodates a wide range of data types, tasks, and model architectures. Extensive experiments on regression and classification tasks with tabular and image data demonstrate that dreaMLearning accelerates training by up to 8.8x, reduces memory usage by 10x, and cuts storage by 42%, with a minimal impact on model performance. These advancements enhance diverse ML applications, including distributed and federated learning, and tinyML on resource-constrained edge devices, unlocking new possibilities for efficient and scalable learning.

cs.LG

Error Analysis of an Approximate Optimal Policy for a Non-stationary Inventory System with Setup Costs

In this paper, we consider a finite horizon non-stationary inventory system with setup costs. We detail an algorithm to find an approximate optimal policy and the basic idea is to use numerical procedure for computing integrals involved in the standard method. We provide analytical error bounds, which converge to zero, between the costs of an approximate optimal policy found in this paper and the optimal policy, which are the main contributions of this paper. To the best of our knowledge, the error bound results are not found in the literature on finite horizon inventory systems with setup costs. A convergence result for the approximate optimal policy is also provided. In this paper, algorithms to find the above error bounds are also provided. Several numerical examples show that the performance of Algorithms in this paper is satisfactory.

math.OC

Reliable IoT Storage: Minimizing Bandwidth Use in Storage Without Newcomer Nodes

This letter characterizes the optimal policies for bandwidth use and storage for the problem of distributed storage in Internet of Things (IoT) scenarios, where lost nodes cannot be replaced by new nodes as is typically assumed in Data Center and Cloud scenarios. We develop an information flow model that captures the overall process of data transmission between IoT devices, from the initial preparation stage (generating redundancy from the original data) to the different repair stages with fewer and fewer devices. Our numerical results show that in a system with 10 nodes, the proposed optimal scheme can save as much as 10.3% of bandwidth use, and as much as 44% storage use with respect to the closest suboptimal approach.

cs.NI

Optimal Scheduling of Electric Vehicles Charging in low-Voltage Distribution Systems

Uncoordinated charging of large-scale electric vehicles (EVs) will have a negative impact on the secure and economic operation of the power system, especially at the distribution level. Given that the charging load of EVs can be controlled to some extent, research on the optimal charging control of EVs has been extensively carried out. In this paper, two possible smart charging scenarios in China are studied: centralized optimal charging operated by an aggregator and decentralized optimal charging managed by individual users. Under the assumption that the aggregators and individual users only concern the economic benefits, new load peaks will arise under time of use (TOU) pricing which is extensively employed in China. To solve this problem, a simple incentive mechanism is proposed for centralized optimal charging while a rolling-update pricing scheme is devised for decentralized optimal charging. The original optimal charging models are modified to account for the developed schemes. Simulated tests corroborate the efficacy of optimal scheduling for charging EVs in various scenarios.

eess.SY

Transmission comb of a distributed Bragg reflector induced by two surface dielectric gratings

With transfer matrix theory, we study the transmission of a distributed Bragg reflector (DBR) with two dielectric gratings on top and on the bottom. Owing to the diffraction of the two gratings, the transmission shows a comb-like spectrum which red shifts with increasing the grating period during the forbidden band of the DBR. The number density of the comb peaks increases with increasing the number of the DBR cells, while the ratio of the average full width at half maximum (FWHM) of the transmission peaks in the transmission comb to the corresponding average free spectral range, being about 0.04 and 0.02 for the TE and TM incident waves, is almost invariant. The average FWHM of the TM waves is about half of the TE waves, and both they could be narrower than 0.1 nm. In addition, the transmission comb peaks of the TE and TM waves can be fully separated during certain waveband. We further prove that the transmission comb is robust against the randomness of the heights of the DBR layers, even when a 15\% randomness is added to their heights. Therefore, the proposed structure is a candidate for a multichannel narrow-band filter or a multichannel polarizer.

physics.optics