SearcharxivSearch

arXiv subjects

Zhengdao Wang

Publications and source records attributed to Zhengdao Wang.

At least 19 recordsLinked to original sources

Noise-Adaptive Layerwise Learning Rates: Accelerating Geometry-Aware Optimization for Deep Neural Network Training

Geometry-aware optimization algorithms, such as Muon, have achieved remarkable success in training deep neural networks (DNNs). These methods leverage the underlying geometry of DNNs by selecting appropriate norms for different layers and updating parameters via norm-constrained linear minimization oracles (LMOs). However, even within a group of layers associated with the same norm, the local curvature can be heterogeneous across layers and vary dynamically over the course of training. For example, recent work shows that sharpness varies substantially across transformer layers and throughout training, yet standard geometry-aware optimizers impose fixed learning rates to layers within the same group, which may be inefficient for DNN training. In this paper, we introduce a noise-adaptive layerwise learning rate scheme on top of geometry-aware optimization algorithms and substantially accelerate DNN training compared to methods that use fixed learning rates within each group. Our method estimates gradient variance in the dual norm induced by the chosen LMO on the fly, and uses it to assign time-varying noise-adaptive layerwise learning rates within each group. We provide a theoretical analysis showing that our algorithm achieves a sharp convergence rate. Empirical results on transformer architectures such as LLaMA and GPT demonstrate that our approach achieves faster convergence than state-of-the-art optimizers.

cs.LG

Reinforcement Learning Architectures: SAC, TAC, and ESAC

The trend is to implement intelligent agents capable of analyzing available information and utilize it efficiently. This work presents a number of reinforcement learning (RL) architectures; one of them is designed for intelligent agents. The proposed architectures are called selector-actor-critic (SAC), tuner-actor-critic (TAC), and estimator-selector-actor-critic (ESAC). These architectures are improved models of a well known architecture in RL called actor-critic (AC). In AC, an actor optimizes the used policy, while a critic estimates a value function and evaluate the optimized policy by the actor. SAC is an architecture equipped with an actor, a critic, and a selector. The selector determines the most promising action at the current state based on the last estimate from the critic. TAC consists of a tuner, a model-learner, an actor, and a critic. After receiving the approximated value of the current state-action pair from the critic and the learned model from the model-learner, the tuner uses the Bellman equation to tune the value of the current state-action pair. ESAC is proposed to implement intelligent agents based on two ideas, which are lookahead and intuition. Lookahead appears in estimating the values of the available actions at the next state, while the intuition appears in maximizing the probability of selecting the most promising action. The newly added elements are an underlying model learner, an estimator, and a selector. The model learner is used to approximate the underlying model. The estimator uses the approximated value function, the learned underlying model, and the Bellman equation to estimate the values of all actions at the next state. The selector is used to determine the most promising action at the next state, which will be used by the actor to optimize the used policy. Finally, the results show the superiority of ESAC compared with the other architectures.

cs.LG

Polarized Low-Density Parity-Check Codes on the BSC

The connections between variable nodes and check nodes have a great influence on the performance of low-density parity-check (LDPC) codes. Inspired by the unique structure of polar code's generator matrix, we proposed a new method of constructing LDPC codes that achieves a polarization effect. The new code, named as polarized LDPC codes, is shown to achieve lower or no error floor in the binary symmetric channel (BSC)

cs.IT

PINE: Universal Deep Embedding for Graph Nodes via Partial Permutation Invariant Set Functions

Graph node embedding aims at learning a vector representation for all nodes given a graph. It is a central problem in many machine learning tasks (e.g., node classification, recommendation, community detection). The key problem in graph node embedding lies in how to define the dependence to neighbors. Existing approaches specify (either explicitly or implicitly) certain dependencies on neighbors, which may lead to loss of subtle but important structural information within the graph and other dependencies among neighbors. This intrigues us to ask the question: can we design a model to give the maximal flexibility of dependencies to each node's neighborhood. In this paper, we propose a novel graph node embedding (named PINE) via a novel notion of partial permutation invariant set function, to capture any possible dependence. Our method 1) can learn an arbitrary form of the representation function from the neighborhood, withour losing any potential dependence structures, and 2) is applicable to both homogeneous and heterogeneous graph embedding, the latter of which is challenged by the diversity of node types. Furthermore, we provide theoretical guarantee for the representation capability of our method for general homogeneous and heterogeneous graphs. Empirical evaluation results on benchmark data sets show that our proposed PINE method outperforms the state-of-the-art approaches on producing node vectors for various learning tasks of both homogeneous and heterogeneous graphs.

cs.LG

Optimization of a Two-Hop Network with Energy Conferencing Relays

This paper considers a two-hop network consisting of a source, two parallel half-duplex relay nodes, and two destinations. While the destinations have an adequate power supply, the source and relay nodes rely on harvested energy for data transmission. Different from all existing works, the two relay nodes can also transfer their harvested energy to each other. For such a system, an optimization problem is formulated with the objective of maximizing the total data rate and conserving the source and relays transmission energy, where any extra energy saved in the current transmission cycle can be used in the next cycle. It turns out that the optimal solutions for this problem can be either found in a closed form or through one-dimensional searches, depending on the scenario. Simulation results based on both the average data rate and the outage probability show that energy cooperation between the two relays consistently improves the system performance.

eess.SP

On the Sublinear Convergence of Randomly Perturbed Alternating Gradient Descent to Second Order Stationary Solutions

The alternating gradient descent (AGD) is a simple but popular algorithm which has been applied to problems in optimization, machine learning, data ming, and signal processing, etc. The algorithm updates two blocks of variables in an alternating manner, in which a gradient step is taken on one block, while keeping the remaining block fixed. When the objective function is nonconvex, it is well-known the AGD converges to the first-order stationary solution with a global sublinear rate. In this paper, we show that a variant of AGD-type algorithms will not be trapped by "bad" stationary solutions such as saddle points and local maximum points. In particular, we consider a smooth unconstrained optimization problem, and propose a perturbed AGD (PA-GD) which converges (with high probability) to the set of second-order stationary solutions (SS2) with a global sublinear rate. To the best of our knowledge, this is the first alternating type algorithm which takes $\mathcal{O}(\text{polylog}(d)/ε^{7/3})$ iterations to achieve SS2 with high probability [where polylog$(d)$ is polynomial of the logarithm of dimension $d$ of the problem].

math.OC

A Nonconvex Splitting Method for Symmetric Nonnegative Matrix Factorization: Convergence Analysis and Optimality

Symmetric nonnegative matrix factorization (SymNMF) has important applications in data analytics problems such as document clustering, community detection and image segmentation. In this paper, we propose a novel nonconvex variable splitting method for solving SymNMF. The proposed algorithm is guaranteed to converge to the set of Karush-Kuhn-Tucker (KKT) points of the nonconvex SymNMF problem. Furthermore, it achieves a global sublinear convergence rate. We also show that the algorithm can be efficiently implemented in parallel. Further, sufficient conditions are provided which guarantee the global and local optimality of the obtained solutions. Extensive numerical results performed on both synthetic and real data sets suggest that the proposed algorithm converges quickly to a local minimum solution.

math.OC

Achievable Rates and Training Optimization for Uplink Multiuser Massive MIMO Systems

We study the performance of uplink transmission in a large-scale (massive) MIMO system, where all the transmitters have single antennas and the base station has a large number of antennas. Specifically, we first derive the rates that are possible through minimum mean-squared error (MMSE) channel estimation and three linear receivers: maximum ratio combining (MRC), zero-forcing (ZF), and MMSE. Based on the derived rates, we quantify the amount of energy savings that are possible through increased number of base-station antennas or increased coherence interval. We also analyze achievable total degrees of freedom (DoF) of such a system without assuming channel state information at the receiver, which is shown to be the same as that of a point-to-point MIMO channel. Linear receiver is sufficient to achieve total DoF when the number of users is less than the number of antennas. When the number of users is equal to or larger than the number of antennas, nonlinear processing is necessary to achieve the full degrees of freedom. Finally, the training period and optimal training energy allocation under the average and peak power constraints are optimized jointly to maximize the achievable sum rate when either MRC or ZF receiver is adopted at the receiver.

cs.IT

Many Access for Small Packets Based on Precoding and Sparsity-aware Recovery

Modern mobile terminals produce massive small data packets. For these short-length packets, it is inefficient to follow the current multiple access schemes to allocate transmission resources due to heavy signaling overhead. We propose a non-orthogonal many-access scheme that is well suited for the future communication systems equipped with many receive antennas. The system is modeled as having a block-sparsity pattern with unknown sparsity level (i.e., unknown number of transmitted messages). Block precoding is employed at each single-antenna transmitter to enable the simultaneous transmissions of many users. The number of simultaneously served active users is allowed to be even more than the number of receive antennas. Sparsity-aware recovery is designed at the receiver for joint user detection and symbol demodulation. To reduce the effects of channel fading on signal recovery, normalized block orthogonal matching pursuit (BOMP) algorithm is introduced, and based on its approximate performance analysis, we develop interference cancellation based BOMP (ICBOMP) algorithm. The ICBOMP performs error correction and detection in each iteration of the normalized BOMP. Simulation results demonstrate the effectiveness of the proposed scheme in small packet services, as well as the advantages of ICBOMP in improving signal recovery accuracy and reducing computational cost.

cs.IT

Joint Optimization of Power Allocation and Training Duration for Uplink Multiuser MIMO Communications

In this paper, we consider a multiuser multiple-input multiple-output (MU-MIMO) communication system between a base station equipped with multiple antennas and multiple mobile users each equipped with a single antenna. The uplink scenario is considered. The uplink channels are acquired by the base station through a training phase. Two linear processing schemes are considered, namely maximum-ratio combining (MRC) and zero-forcing (ZF). We optimize the training period and optimal training energy under the average and peak power constraint so that an achievable sum rate is maximized.

cs.IT

Multiple Access for Small Packets Based on Precoding and Sparsity-Aware Detection

Modern mobile terminals often produce a large number of small data packets. For these packets, it is inefficient to follow the conventional medium access control protocols because of poor utilization of service resources. We propose a novel multiple access scheme that employs block-spreading based precoding at the transmitters and sparsity-aware detection schemes at the base station. The proposed scheme is well suited for the emerging massive multiple-input multiple-output (MIMO) systems, as well as conventional cellular systems with a small number of base-station antennas. The transmitters employ precoding in time domain to enable the simultaneous transmissions of many users, which could be even more than the number of receive antennas at the base station. The system is modeled as a linear system of equations with block-sparse unknowns. We first adopt the block orthogonal matching pursuit (BOMP) algorithm to recover the transmitted signals. We then develop an improved algorithm, named interference cancellation BOMP (ICBOMP), which takes advantage of error correction and detection coding to perform perfect interference cancellation during each iteration of BOMP algorithm. Conditions for guaranteed data recovery are identified. The simulation results demonstrate that the proposed scheme can accommodate more simultaneous transmissions than conventional schemes in typical small-packet transmission scenarios.

cs.IT

A Novel Uplink Data Transmission Scheme For Small Packets In Massive MIMO System

Intelligent terminals often produce a large number of data packets of small lengths. For these packets, it is inefficient to follow the conventional medium access control (MAC) protocols because they lead to poor utilization of service resources. We propose a novel multiple access scheme that targets massive multiple-input multiple-output (MIMO) systems based on compressive sensing (CS). We employ block precoding in the time domain to enable the simultaneous transmissions of many users, which could be even more than the number of receive antennas at the base station. We develop a block-sparse system model and adopt the block orthogonal matching pursuit (BOMP) algorithm to recover the transmitted signals. Conditions for data recovery guarantees are identified and numerical results demonstrate that our scheme is efficient for uplink small packet transmission.

cs.IT

Multiple-Antenna Interference Network with Receive Antenna Joint Processing and Real Interference Alignment

In this paper, the degrees of freedom (DoF) regions of constant coefficient multiple antenna interference channels are investigated. First, we consider a $K$-user Gaussian interference channel with $M_k$ antennas at transmitter $k$, $1\le k\le K$, and $N_j$ antennas at receiver $j$, $1\le j\le K$, denoted as a $(K,[M_k],[N_j])$ channel. Relying on a result of simultaneous Diophantine approximation, a real interference alignment scheme with joint receive antenna processing is developed. The scheme is used to obtain an achievable DoF region. The proposed DoF region includes two previously known results as special cases, namely 1) the total DoF of a $K$-user interference channel with $N$ antennas at each node, $(K, [N], [N])$ channel, is $NK/2$; and 2) the total DoF of a $(K, [M], [N])$ channel is at least $KMN/(M+N)$. We next explore constant-coefficient interference networks with $K$ transmitters and $J$ receivers, all having $N$ antennas. Each transmitter emits an independent message and each receiver requests an arbitrary subset of the messages. Employing the novel joint receive antenna processing, the DoF region for this set-up is obtained. We finally consider wireless X networks where each node is allowed to have an arbitrary number of antennas. It is shown that the joint receive antenna processing can be used to establish an achievable DoF region, which is larger than what is possible with antenna splitting. As a special case of the derived achievable DoF region for constant coefficient X network, the total DoF of wireless X networks with the same number of antennas at all nodes and with joint antenna processing is tight while the best inner bound based on antenna splitting cannot meet the outer bound. Finally, we obtain a DoF region outer bound based on the technique of transmitter grouping.

cs.IT

Performance of Uplink Multiuser Massive MIMO Systems

We study the performance of uplink transmission in a large-scale (massive) MIMO system, where all the transmitters have single antennas and the receiver (base station) has a large number of antennas. Specifically, we analyze achievable degrees of freedom of the system without assuming channel state information at the receiver. Also, we quantify the amount of power saving that is possible with increasing number of receive antennas.

cs.IT

Multiple-Antenna Interference Channel with Receive Antenna Joint Processing and Real Interference Alignment

We consider a constant $K$-user Gaussian interference channel with $M$ antennas at each transmitter and $N$ antennas at each receiver, denoted as a $(K,M,N)$ channel. Relying on a result on simultaneous Diophantine approximation, a real interference alignment scheme with joint receive antenna processing is developed. The scheme is used to provide new proofs for two previously known results, namely 1) the total degrees of freedom (DoF) of a $(K, N, N)$ channel is $NK/2$; and 2) the total DoF of a $(K, M, N)$ channel is at least $KMN/(M+N)$. We also derive the DoF region of the $(K,N,N)$ channel, and an inner bound on the DoF region of the $(K,M,N)$ channel.

cs.IT

Joint Viterbi Decoding and Decision Feedback Equalization for Monobit Digital Receivers

In ultra-wideband (UWB) communication systems with impulse radio (IR) modulation, the bandwidth is usually 1GHz or more. To process the received signal digitally, high sampling rate analog-digital-converters (ADC) are required. Due to the high complexity and large power consumption, monobit ADC is appropriate. The optimal monobit receiver has been derived. But it is not efficient to combat intersymbol interference (ISI). Decision feedback equalization (DFE) is an effect way dealing with ISI. In this paper, we proposed a algorithm that combines Viterbi decoding and DFE together for monobit receivers. In this way, we suppress the impact of ISI effectively, thus improving the bit error rate (BER) performance. By state expansion, we achieve better performance. The simulation results show that the algorithm has about 1dB SNR gain compared to separate demodulation and decoding method and 1dB loss compared to the BER performance in the channel without ISI. Compare to the full resolution detection in fading channel without ISI, it has 3dB SNR loss after state expansion.

cs.IT

Degrees of Freedom Region for an Interference Network with General Message Demands

We consider a single hop interference network with $K$ transmitters and $J$ receivers, all having $M$ antennas. Each transmitter emits an independent message and each receiver requests an arbitrary subset of the messages. This generalizes the well-known $K$-user $M$-antenna interference channel, where each message is requested by a unique receiver. For our setup, we derive the degrees of freedom (DoF) region. The achievability scheme generalizes the interference alignment schemes proposed by Cadambe and Jafar. In particular, we achieve general points in the DoF region by using multiple base vectors and aligning all interferers at a given receiver to the interferer with the largest DoF. As a byproduct, we obtain the DoF region for the original interference channel. We also discuss extensions of our approach where the same region can be achieved by considering a reduced set of interference alignment constraints, thus reducing the time-expansion duration needed. The DoF region for the considered system depends only on a subset of receivers whose demands meet certain characteristics. The geometric shape of the DoF region is also discussed.

cs.IT

Real Interference Alignment and Degrees of Freedom Region of Wireless X Networks

We consider a single hop wireless X network with $K$ transmitters and $J$ receivers, all with single antenna. Each transmitter conveys for each receiver an independent message. The channel is assumed to have constant coefficients. We develop interference alignment scheme for this setup and derived several achievable degrees of freedom regions. We show that in some cases, the derived region meets a previous outer bound and are hence the DoF region. For our achievability schemes, we divide each message into streams and use real interference alignment on the streams. Several previous results on the DoF region and total DoF for various special cases can be recovered from our result.

cs.IT