SearcharxivSearch

arXiv subjects

Bangbang Ren

Publications and source records attributed to Bangbang Ren.

6 recordsLinked to original sources

Sensing and Storing Less: A MARL-based Solution for Energy Saving in Edge Internet of Things

As the number of Internet of Things (IoT) devices continuously grows and application scenarios constantly enrich, the volume of sensor data experiences an explosive increase. However, substantial data demands considerable energy during computation and transmission. Redundant deployment or mobile assistance is essential to cover the target area reliably with fault-prone sensors. Consequently, the ``butterfly effect" may appear during the IoT operation, since unreasonable data overlap could result in many duplicate data. To this end, we propose Senses, a novel online energy saving solution for edge IoT networks, with the insight of sensing and storing less at the network edge by adopting Muti-Agent Reinforcement Learning (MARL). Senses achieves data de-duplication by dynamically adjusting sensor coverage at the sensor level. For exceptional cases where sensor coverage cannot be altered, Senses conducts data partitioning and eliminates redundant data at the controller level. Furthermore, at the global level, considering the heterogeneity of IoT devices, Senses balances the operational duration among the devices to prolong the overall operational duration of edge IoT networks. We evaluate the performance of Senses through testbed experiments and simulations. The results show that Senses saves 11.37% of energy consumption on control devices and prolongs 20% overall operational duration of the IoT device network.

cs.NI

Communication and Computation Efficient Split Federated Learning in O-RAN

The hierarchical architecture of Open Radio Access Network (O-RAN) has enabled a new Federated Learning (FL) paradigm that trains models using data from non- and near-real-time (near-RT) Radio Intelligent Controllers (RICs). However, the ever-increasing model size leads to longer training time, jeopardizing the deadline requirements for both non-RT and near-RT RICs. To address this issue, split federated learning (SFL) offers an approach by offloading partial model layers from near-RT-RIC to high-performance non-RT-RIC. Nonetheless, its deployment presents two challenges: (i) Frequent data/gradient transfers between near-RT-RIC and non-RT-RIC in SFL incur significant communication cost in O-RAN. (ii) Proper allocation of computational and communication resources in O-RAN is vital to satisfying the deadline and affects SFL convergence. Therefore, we propose SplitMe, an SFL framework that exploits mutual learning to alternately and independently train the near-RT-RIC's model and the non-RT-RIC's inverse model, eliminating frequent transfers. The ''inverse'' of the inverse model is derived via a zeroth-order technique to integrate the final model. Then, we solve a joint optimization problem for SplitMe to minimize overall resource costs with deadline-aware selection of near-RT-RICs and adaptive local updates. Our numerical results demonstrate that SplitMe remarkably outperforms FL frameworks like SFL, FedAvg and O-RANFed regarding costs and convergence.

cs.LG

DHO$_2$: Accelerating Distributed Hybrid Order Optimization via Model Parallelism and ADMM

Scaling deep neural network (DNN) training to more devices can reduce time-to-solution. However, it is impractical for users with limited computing resources. FOSI, as a hybrid order optimizer, converges faster than conventional optimizers by taking advantage of both gradient information and curvature information when updating the DNN model. Therefore, it provides a new chance for accelerating DNN training in the resource-constrained setting. In this paper, we explore its distributed design, namely DHO$_2$, including distributed calculation of curvature information and model update with partial curvature information to accelerate DNN training with a low memory burden. To further reduce the training time, we design a novel strategy to parallelize the calculation of curvature information and the model update on different devices. Experimentally, our distributed design can achieve an approximate linear reduction of memory burden on each device with the increase of the device number. Meanwhile, it achieves $1.4\times\sim2.1\times$ speedup in the total training time compared with other distributed designs based on conventional first- and second-order optimizers.

cs.LG

Analytic Personalized Federated Meta-Learning

Analytic Federated Learning (AFL) is an enhanced gradient-free federated learning (FL) paradigm designed to accelerate training by updating the global model in a single step with closed-form least-square (LS) solutions. However, the obtained global model suffers performance degradation across clients with heterogeneous data distribution. Meta-learning is a common approach to tackle this problem by delivering personalized local models for individual clients. Yet, integrating meta-learning with AFL presents significant challenges: First, conventional AFL frameworks cannot support deep neural network (DNN) training which can influence the fast adaption capability of meta-learning for complex FL tasks. Second, the existing meta-learning method requires gradient information, which is not involved in AFL. To overcome the first challenge, we propose an AFL framework, namely FedACnnL, in which a layer-wise DNN collaborative training method is designed by modeling the training of each layer as a distributed LS problem. For the second challenge, we further propose an analytic personalized federated meta-learning framework, namely pFedACnnL. It generates a personalized model for each client by analytically solving a local objective which bridges the gap between the global model and the individual data distribution. FedACnnL is theoretically proven to require significantly shorter training time than the conventional FL frameworks on DNN training while the reduction ratio is $83\%\sim99\%$ in the experiment. Meanwhile, pFedACnnL excels at test accuracy with the vanilla FedACnnL by $4\%\sim8\%$ and it achieves state-of-the-art (SOTA) model performance in most cases of convex and non-convex settings compared with previous SOTA frameworks.

cs.DC

Embedding the Minimum Cost SFC with End-to-end Delay Constraint

Many network applications, especially the multimedia applications, often deliver flows with high QoS, like end-to-end delay constraint. Flows of these applications usually need to traverse a series of different network functions orderly before reaching to the host in the customer end, which is called the service function chain (SFC). The emergence of network function virtualization (NFV) increases the deployment flexibility of such network functions. In this paper, we present heuristics to embed the SFC for a given flow considering: i) bounded end-to-end delay along the path, and ii) minimum cost of the SFC embedding, where cost and delay can be independent metrics and be attached to both links and nodes. This problem of embedding SFC is NP-hard, which can be reduced to the Knapsack problem. We then design a greedy algorithm which is applied to a multilevel network. The simulation results demonstrate that the multilevel greedy algorithm can efficiently solve the NP-hard problem.

cs.NI

PPtaxi: Non-stop Package Delivery via Multi-hop Ridesharing

City-wide package delivery becomes popular due to the dramatic rise of online shopping. It places a tremendous burden on the traditional logistics industry, which relies on dedicated couriers and is labor-intensive. Leveraging the ridesharing systems is a promising alternative, yet existing solutions are limited to one-hop ridesharing or need consignment warehouses as relays. In this paper, we propose a new package delivery scheme which takes advantage of multi-hop ridesharing and is entirely consignment free. Specifically, a package is assigned to a taxi which is guided to deliver the package all along to its destination while transporting successive passengers. We tackle it with a two-phase solution, named \textbf{PPtaxi}. In the first phase, we use the Multivariate Gauss distribution and Bayesian inference to predict the passenger orders. In the second phase, both the computation efficiency and solution effectiveness are considered to plan package delivery routes. We evaluate \textbf{PPtaxi} with a real-world dataset from an online taxi-taking platform and compare it with multiple benchmarks. The results show that the successful delivery rate of packages with our solution can reach $95\%$ on average during the daytime, and is at most $46.9\%$ higher than those of the benchmarks.

cs.DC