SearcharxivSearch

arXiv subjects

Zhuang Wang

Publications and source records attributed to Zhuang Wang.

32 records · Page 2Linked to original sources

MXDAG: A Hybrid Abstraction for Cluster Applications

Distributed applications, such as database queries and distributed training, consist of both compute and network tasks. DAG-based abstraction primarily targets compute tasks and has no explicit network-level scheduling. In contrast, Coflow abstraction collectively schedules network flows among compute tasks but lacks the end-to-end view of the application DAG. Because of the dependencies and interactions between these two types of tasks, it is sub-optimal to only consider one of them. We argue that co-scheduling of both compute and network tasks can help applications towards the globally optimal end-to-end performance. However, none of the existing abstractions can provide fine-grained information for co-scheduling. We propose MXDAG, an abstraction to treat both compute and network tasks explicitly. It can capture the dependencies and interactions of both compute and network tasks leading to improved application performance.

cs.DC

Efficient and Less Centralized Federated Learning

With the rapid growth in mobile computing, massive amounts of data and computing resources are now located at the edge. To this end, Federated learning (FL) is becoming a widely adopted distributed machine learning (ML) paradigm, which aims to harness this expanding skewed data locally in order to develop rich and informative models. In centralized FL, a collection of devices collaboratively solve a ML task under the coordination of a central server. However, existing FL frameworks make an over-simplistic assumption about network connectivity and ignore the communication bandwidth of the different links in the network. In this paper, we present and study a novel FL algorithm, in which devices mostly collaborate with other devices in a pairwise manner. Our nonparametric approach is able to exploit network topology to reduce communication bottlenecks. We evaluate our approach on various FL benchmarks and demonstrate that our method achieves 10X better communication efficiency and around 8% increase in accuracy compared to the centralized approach.

cs.DC

Shufflecast: An Optical, Data-rate Agnostic and Low-Power Multicast Architecture for Next-Generation Compute Clusters

An optical circuit-switched network core has the potential to overcome the inherent challenges of a conventional electrical packet-switched core of today's compute clusters. As optical circuit switches (OCS) directly handle the photon beams without any optical-electrical-optical (O/E/O) conversion and packet processing, OCS-based network cores have the following desirable properties: a) agnostic to data-rate, b) negligible/zero power consumption, c) no need of transceivers, d) negligible forwarding latency, and e) no need for frequent upgrade. Unfortunately, OCS can only provide point-to-point (unicast) circuits. They do not have built-in support for one-to-many (multicast) communication, yet multicast is fundamental to a plethora of data-intensive applications running on compute clusters nowadays. In this paper, we propose Shufflecast, a novel optical network architecture for next-generation compute clusters that can support high-performance multicast satisfying all the properties of an OCS-based network core. Shufflecast leverages small fanout, inexpensive, passive optical splitters to connect the Top-of-rack (ToR) switch ports, ensuring data-rate agnostic, low-power, physical-layer multicast. We thoroughly analyze Shufflecast's highly scalable data plane, light-weight control plane, and graceful failure handling. Further, we implement a complete prototype of Shufflecast in our testbed and extensively evaluate the network. Shufflecast is more power-efficient than the state-of-the-art multicast mechanisms. Also, Shufflecast is more cost-efficient than a conventional packet-switched network. By adding Shufflecast alongside an OCS-based unicast network, an all-optical network core with the aforementioned desirable properties supporting both unicast and multicast can be realized.

cs.NI

Improved regularity of harmonic diffeomorphic extensions on quasihyperbolic domains

Let $\mathbb{X}$ be a Jordan domain satisfying hyperbolic growth conditions. Assume that $\varphi$ is a homeomorphism from the boundary $\partial \mathbb{X}$ of $\mathbb{X}$ onto the unit circle. Denote by $h$ the harmonic diffeomorphic extension of $\varphi $ from $\mathbb{X}$ onto the unit disk. We establish the optimal Orlicz-Sobolev regularity and weighted Sobolev estimate of $h.$ These generalize the Sobolev regularity of $h$ by Koski-Onninen [21, Theorem 3.1].

math.CV

MergeComp: A Compression Scheduler for Scalable Communication-Efficient Distributed Training

Large-scale distributed training is increasingly becoming communication bound. Many gradient compression algorithms have been proposed to reduce the communication overhead and improve scalability. However, it has been observed that in some cases gradient compression may even harm the performance of distributed training. In this paper, we propose MergeComp, a compression scheduler to optimize the scalability of communication-efficient distributed training. It automatically schedules the compression operations to optimize the performance of compression algorithms without the knowledge of model architectures or system parameters. We have applied MergeComp to nine popular compression algorithms. Our evaluations show that MergeComp can improve the performance of compression algorithms by up to 3.83x without losing accuracy. It can even achieve a scaling factor of distributed training up to 99% over high-speed networks.

cs.DC

Characterization of trace spaces on regular trees via dyadic norms

In this paper, we study the traces of Orlicz-Sobolev spaces on a regular rooted tree. After giving a dyadic decomposition of the boundary of the regular tree, we present a characterization on the trace spaces of those first order Orlicz-Sobolev spaces whose Young function is of the form $t^p\log^\lambda(e+t)$, based on integral averages on dyadic elements of the dyadic decomposition.

math.FA

Traces of Newton-Sobolev, Hajlasz-Sobolev and BV functions on metric spaces

We study the boundary traces of Newton-Sobolev, Hajlasz-Sobolev, and BV (bounded variation) functions. Assuming less regularity of the domain than is usually done in the literature, we show that all of these function classes achieve the same "boundary values", which in particular implies that the trace spaces coincide provided that they exist. Many of our results seem to be new even in Euclidean spaces but we work in a more general complete metric space equipped with a doubling measure and supporting a Poincare inequality.

math.MG

Controlled diffeomorphic extension of homeomorphisms

Let $\Omega$ be an internal chord-arc Jordan domain and $\varphi:\mathbb S\rightarrow\partial\Omega$ be a homeomorphism. We show that $\varphi$ has finite dyadic energy if and only if $\varphi$ has a diffeomorphic extension $h: \mathbb D\rightarrow \Omega$ which has finite energy.

math.CV

Delay-Energy Joint Optimization for Task Offloading in Mobile Edge Computing

Mobile-edge computing (MEC) has been envisioned as a promising paradigm to meet ever-increasing resource demands of mobile users, prolong battery lives of mobile devices, and shorten request response delays experienced by users. An MEC environment consists of many MEC servers and ubiquitous access points interconnected into an edge cloud network. Mobile users can offload their computing-intensive tasks to one or multiple MEC servers for execution to save their batteries. Due to large numbers of MEC servers deployed in MEC, selecting a subset of servers to serve user tasks while satisfying delay requirements of their users is challenging. In this paper, we formulate a novel delay-energy joint optimization problem through jointly considering the CPU-cycle frequency scheduling at mobile devices, server selection to serve user offloading tasks, and task allocations to the selected servers. To this end, we first formulate the problem as a mixed-integer nonlinear programming, due to the hardness to solve this nonlinear programming, we instead then relax the problem into a nonlinear programming problem that can be solved in polynomial time. We also show how to derive a feasible solution to the original problem from the solution of this relaxed solution. We finally conduct experiments to evaluate the performance of the proposed algorithm. Experimental results demonstrate that the proposed algorithm is promising.

cs.NI

Traces of weighted function spaces: dyadic norms and Whitney extensions

The trace spaces of Sobolev spaces and related fractional smoothness spaces have been an active area of research since the work of Nikolskii, Aronszajn, Slobodetskii, Babich and Gagliardo among others in the 1950's. In this paper we review the literature concerning such results for a variety of weighted smoothness spaces. For this purpose, we present a characterization of the trace spaces (of fractional order of smoothness), based on integral averages on dyadic cubes, which is well adapted to extending functions using the Whitney extension operator.

math.CA

Fair Packet Scheduling in Network on Chip

Interconnection networks of parallel systems are used for servicing traf- fic generated by different applications, often belonging to different users. When multiple traffic flows contend for channel bandwidth, the scheduling algorithm regulating the access to that channel plays a key role in ensur- ing that each flow obtains the required quality of service. Fairness is a highly desirable property for a scheduling algorithm. We show that using the Relative Fairness Bound as a fairness measure may lead to decrease in throughput and increase in latency. We propose an alternative metric to evaluate the fairness and avoid the drawback of Relative Fairness Bound.

cs.DC