Searcharxiv⌕ Search

arXiv subjects

P. R. Kumar

Publications and source records attributed to P. R. Kumar.

At least 37 records · Page 2Linked to original sources

Learning from Few Samples: Transformation-Invariant SVMs with Composition and Locality at Multiple Scales

Motivated by the problem of learning with small sample sizes, this paper shows how to incorporate into support-vector machines (SVMs) those properties that have made convolutional neural networks (CNNs) successful. Particularly important is the ability to incorporate domain knowledge of invariances, e.g., translational invariance of images. Kernels based on the \textit{maximum} similarity over a group of transformations are not generally positive definite. Perhaps it is for this reason that they have not been studied theoretically. We address this lacuna and show that positive definiteness indeed holds \textit{with high probability} for kernels based on the maximum similarity in the small training sample set regime of interest, and that they do yield the best results in that regime. We also show how additional properties such as their ability to incorporate local features at multiple spatial scales, e.g., as done in CNNs through max pooling, and to provide the benefits of composition through the architecture of multiple layers, can also be embedded into SVMs. We verify through experiments on widely available image sets that the resulting SVMs do provide superior accuracy in comparison to well-established deep neural network benchmarks for small sample sizes.

cs.LG↗

Anchor-Changing Regularized Natural Policy Gradient for Multi-Objective Reinforcement Learning

We study policy optimization for Markov decision processes (MDPs) with multiple reward value functions, which are to be jointly optimized according to given criteria such as proportional fairness (smooth concave scalarization), hard constraints (constrained MDP), and max-min trade-off. We propose an Anchor-changing Regularized Natural Policy Gradient (ARNPG) framework, which can systematically incorporate ideas from well-performing first-order methods into the design of policy optimization algorithms for multi-objective MDP problems. Theoretically, the designed algorithms based on the ARNPG framework achieve $\tilde{O}(1/T)$ global convergence with exact gradients. Empirically, the ARNPG-guided algorithms also demonstrate superior performance compared to some existing policy gradient-based approaches in both exact gradients and sample-based scenarios.

cs.LG↗

Terra: Blockage Resilience in Outdoor mmWave Networks

We address the problem of pedestrian blockage in outdoor mm-Wave networks, which can disrupt Line of Sight (LoS) communication and result in an outage. This necessitates either reacquisition as a new user that can take up to 1.28 sec in 5G New Radio, disrupting high-performance applications, or handover to a different base station (BS), which however requires very costly dense deployments to ensure multiple base stations are always visible to a mobile outdoors. We have found that there typically exists a strong ground reflection from concrete and gravel surfaces, within 4-6 dB of received signal strength (RSS) of LoS paths. The mobile can switch to such a Non-Line Sight (NLoS) beam to sustain the link for use as a control channel during a blockage event. This allows the mobile to maintain time-synchronization with the base station, allowing it to revert to the LoS path when the temporary blockage disappears. We present a protocol, Terra, to quickly discover, cache, and employ ground reflections. It can be used in most outdoor built environments since pedestrian blockages typically last only a few hundred milliseconds.

eess.SY↗

Policy Optimization for Constrained MDPs with Provable Fast Global Convergence

We address the problem of finding the optimal policy of a constrained Markov decision process (CMDP) using a gradient descent-based algorithm. Previous results have shown that a primal-dual approach can achieve an $\mathcal{O}(1/\sqrt{T})$ global convergence rate for both the optimality gap and the constraint violation. We propose a new algorithm called policy mirror descent-primal dual (PMD-PD) algorithm that can provably achieve a faster $\mathcal{O}(\log(T)/T)$ convergence rate for both the optimality gap and the constraint violation. For the primal (policy) update, the PMD-PD algorithm utilizes a modified value function and performs natural policy gradient steps, which is equivalent to a mirror descent step with appropriate regularization. For the dual update, the PMD-PD algorithm uses modified Lagrange multipliers to ensure a faster convergence rate. We also present two extensions of this approach to the settings with zero constraint violation and sample-based estimation. Experimental results demonstrate the faster convergence rate and the better performance of the PMD-PD algorithm compared with existing policy gradient-based algorithms.

cs.LG↗

Overcoming Pedestrian Blockage in mm-Wave Bands using Ground Reflections

mm-Wave communication employs directional beams to overcome high path loss. High data rate communication is typically along line-of-sight (LoS). In outdoor environments, such communication is susceptible to temporary blockage by pedestrians interposed between the transmitter and receiver. It results in outages in which the user is lost, and has to be reacquired as a new user, severely disrupting interactive and high throughput applications. It has been presumed that the solution is to have a densely deployed set of base stations that will allow the mobile to perform a handover to a different non-blocked base station every time a current base station is blocked. This is however a very costly solution for outdoor environments. Through extensive experiments we show that it is possible to exploit a strong ground reflection with a received signal strength (RSS) about 4dB less than the LoS path in outdoor built environments with concrete or gravel surfaces, for beams that are narrow in azimuth but wide in zenith. While such reflected paths cannot support the high data rates of LoS paths, they can support control channel communication, and, importantly, sustain time synchronization between the mobile and the base station. This allows a mobile to quickly recover to the LoS path upon the cessation of the temporary blockage, which typically lasts a few hundred milliseconds. We present a simple in-band protocol that quickly discovers ground reflected radiation and uses it to recover the LoS link when the temporary blockage disappears.

cs.IT↗

BeamSurfer: Minimalist Beam Management of Mobile mm-Wave Devices

Management of narrow directional beams is critical for mm-wave communication systems. Translational or rotational motion of the user can cause misalignment of transmit and receive beams with the base station losing track of the mobile. Reacquiring the user can take about one second in 5G NewRadio systems and significantly impair performance of applications, besides being energy intensive. It is therefore important to manage beams to continually maintain high received signal strength and prevent outages. It is also important to be able to recover from sudden but transient blockage caused by a handor face interposed in the Line-of-Sight (LoS) path. This work presents a beam management protocol called BeamSurfer that is targeted to the use case of users walking indoors near a base station. It is designed to be minimalistic, employing only in-band information, and not requiring knowledge such as location or orientation of the mobile device or any additional sensors. Evaluations under pedestrian mobility show that96%of the time it employs beams that are within3dB of what an omniscient Oracle would have chosen as the pair of optimal transmit-receive LoS beams. It also recovers from transient LoS blockage by continuing to maintain control packet communication over a pre-selected reflected path that preserves time-synchronization with the base station, which allows it to rapidly recover to the LoS link as soon as it is unblocked

eess.SY↗

Silent Tracker: In-band Beam Management for Soft Handover for mm-Wave Networks

In mm-wave networks, cell sizes are small due to high path and penetration losses. Mobiles need to frequently switch softly from one cell to another to preserve network connections and context. Each soft handover involves the mobile performing directional neighbor cell search, tracking cell beam, completing cell access request, and finally, context switching. The mobile must independently discover cell beams, derive timing information, and maintain beam alignment throughout the process to avoid packet loss and hard handover. We propose Silent tracker which enables a mobile to reliably manage handover events by maintaining an aligned beam until the successful handover completion. It is entirely in-band beam mechanism that does not need any side information. Experimental evaluations show that Silent Tracker maintains the mobile's receive beam aligned to the potential target base station's transmit beam till the successful conclusion of handover in three mobility scenarios: human walk, device rotation, and 20 mph vehicular speed.

eess.SY↗

An Efficient Network Solver for Dynamic Simulation of Power Systems Based on Hierarchical Inverse Computation and Modification

In power system dynamic simulation, up to 90% of the computational time is devoted to solve the network equations, i.e., a set of linear equations. Traditional approaches are based on sparse LU factorization, which is inherently sequential. In this paper, an inverse-based network solution is proposed by a hierarchical method for computing and store the approximate inverse of the conductance matrix in electromagnetic transient (EMT) simulations. The proposed method can also efficiently update the inverse by modifying only local sub-matrices to reflect changes in the network, e.g., loss of a line. Experiments on a series of simplified 179-bus Western Interconnection demonstrate the advantages of the proposed methods.

eess.SY↗

Reward Biased Maximum Likelihood Estimation for Reinforcement Learning

The Reward-Biased Maximum Likelihood Estimate (RBMLE) for adaptive control of Markov chains was proposed to overcome the central obstacle of what is variously called the fundamental "closed-identifiability problem" of adaptive control, the "dual control problem", or, contemporaneously, the "exploration vs. exploitation problem". It exploited the key observation that since the maximum likelihood parameter estimator can asymptotically identify the closed-transition probabilities under a certainty equivalent approach, the limiting parameter estimates must necessarily have an optimal reward that is less than the optimal reward attainable for the true but unknown system. Hence it proposed a counteracting reverse bias in favor of parameters with larger optimal rewards, providing a solution to the fundamental problem alluded to above. It thereby proposed an optimistic approach of favoring parameters with larger optimal rewards, now known as "optimism in the face of uncertainty". The RBMLE approach has been proved to be long-term average reward optimal in a variety of contexts. However, modern attention is focused on the much finer notion of "regret", or finite-time performance. Recent analysis of RBMLE for multi-armed stochastic bandits and linear contextual bandits has shown that it not only has state-of-the-art regret, but it also exhibits empirical performance comparable to or better than the best current contenders, and leads to strikingly simple index policies. Motivated by this, we examine the finite-time performance of RBMLE for reinforcement learning tasks that involve the general problem of optimal control of unknown Markov Decision Processes. We show that it has a regret of $\mathcal{O}( \log T)$ over a time horizon of $T$ steps, similar to state-of-the-art algorithms. Simulation studies show that RBMLE outperforms other algorithms such as UCRL2 and Thompson Sampling.

cs.LG↗

UNBLOCK: Low Complexity Transient Blockage Recovery for Mobile mm-Wave Devices

Directional radio beams are used in the mm-Wave band to combat the high path loss. The mm-Wave band also suffers from high penetration losses from drywall, wood, glass, concrete, etc., and also the human body. Hence, as a mobile user moves, the Line of Sight (LoS) path between the mobile and the Base Station (BS) can be blocked by objects interposed in the path, causing loss of the link. A mobile with a lost link will need to be re-acquired as a new user by initial access, a process that can take up to a second, causing disruptions to applications. UNBLOCK is a protocol that allows a mobile to recover from transient blockages, such as those caused by a human hand or another human walking into the line of path or other temporary occlusions by objects, which typically disappear within the order of $100$ ms, without having to go through re-acquisition. UNBLOCK is based on extensive experimentation in office type environments which has shown that while a LoS path is blocked, there typically exists a Non-LoS path, i.e., a reflected path through scatterers, with a loss within about $10$ dB of the LoS path. UNBLOCK proactively keeps such a NLoS path in reserve, to be used when blockage happens, typically without any warning. UNBLOCK uses this NLoS path to maintain time-synchronization with the BS until the blockage disappears, as well as to search for a better NLoS path if available. When the transient blockage disappears, it reestablishes LoS communication at the epochs that have been scheduled by the BS for communication with the mobile.

cs.NI↗

Detection of Cyber Attacks in Renewable-rich Microgrids Using Dynamic Watermarking

This paper presents the first demonstration of using an active mechanism to defend renewable-rich microgrids against cyber attacks. Cyber vulnerability of the renewable-rich microgrids is identified. The defense mechanism based on dynamic watermarking is proposed for detecting cyber anomalies in microgrids. The proposed mechanism is easily implementable and it has theoretically provable performance in term of detecting cyber attacks. The effectiveness of the proposed mechanism is tested and validated in a renewable-rich microgrid.

eess.SY↗

Exploration Through Reward Biasing: Reward-Biased Maximum Likelihood Estimation for Stochastic Multi-Armed Bandits

Inspired by the Reward-Biased Maximum Likelihood Estimate method of adaptive control, we propose RBMLE -- a novel family of learning algorithms for stochastic multi-armed bandits (SMABs). For a broad range of SMABs including both the parametric Exponential Family as well as the non-parametric sub-Gaussian/Exponential family, we show that RBMLE yields an index policy. To choose the bias-growth rate $α(t)$ in RBMLE, we reveal the nontrivial interplay between $α(t)$ and the regret bound that generally applies in both the Exponential Family as well as the sub-Gaussian/Exponential family bandits. To quantify the finite-time performance, we prove that RBMLE attains order-optimality by adaptively estimating the unknown constants in the expression of $α(t)$ for Gaussian and sub-Gaussian bandits. Extensive experiments demonstrate that the proposed RBMLE achieves empirical regret performance competitive with the state-of-the-art methods, while being more computationally efficient and scalable in comparison to the best-performing ones among them.

cs.LG↗

Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits

Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit trade-off in linear bandits problems as well as generalized linear bandits problems. We develop novel index policies that we prove achieve order-optimality, and show that they achieve empirical performance competitive with the state-of-the-art benchmark methods in extensive experiments. The new policies achieve this with low computation time per pull for linear bandits, and thereby resulting in both favorable regret as well as computational efficiency.

cs.LG↗

Incentive Compatibility in Stochastic Dynamic Systems

While the classic Vickrey-Clarke-Groves mechanism ensures incentive compatibility for a static one-shot game, it does not appear to be feasible to construct a dominant truth-telling mechanism for agents that are stochastic dynamic systems. The contribution of this paper is to show that for a set of LQG agents a mechanism consisting of a sequence of layered payments over time decouples the intertemporal effect of current bids on future payoffs and ensures truth-telling of dynamic states by their agents, if system parameters are known and agents are rational. Additionally, it is shown that there is a "Scaled" VCG mechanism that simultaneously satisfies incentive compatibility, social efficiency, budget balance as well as individual rationality under a certain "market power balance" condition where no agent is too negligible or too dominant. A further desirable property is that the SVCG payments converge to the Lagrange payments, the payments corresponding to the true price in the absence of strategic considerations, as the number of agents in the market increases. For LQ but non-Gaussian agents the optimal social welfare over the class of linear control laws is achieved.

eess.SY↗

Learning in Networked Control Systems

We design adaptive controller (learning rule) for a networked control system (NCS) in which data packets containing control information are transmitted across a lossy wireless channel. We propose Upper Confidence Bounds for Networked Control Systems (UCB-NCS), a learning rule that maintains confidence intervals for the estimates of plant parameters $(A_{(\star)},B_{(\star)})$, and channel reliability $p_{(\star)}$, and utilizes the principle of optimism in the face of uncertainty while making control decisions. We provide non-asymptotic performance guarantees for UCB-NCS by analyzing its "regret", i.e., performance gap from the scenario when $(A_{(\star)},B_{(\star)},p_{(\star)})$ are known to the controller. We show that with a high probability the regret can be upper-bounded as $\tilde{O}\left(C\sqrt{T}\right)$\footnote{Here $\tilde{O}$ hides logarithmic factors.}, where $T$ is the operating time horizon of the system, and $C$ is a problem dependent constant.

cs.LG↗

A Synchrophasor Data-driven Method for Forced Oscillation Localization under Resonance Conditions

This paper proposes a data-driven algorithm of locating the source of forced oscillations and suggests the physical interpretation of the method. By leveraging the sparsity of the forced oscillation sources along with the low-rank nature of synchrophasor data, the problem of source localization under resonance conditions is cast as computing the sparse and low-rank components using Robust Principal Component Analysis (RPCA), which can be efficiently solved by the exact Augmented Lagrange Multiplier method. Based on this problem formulation, an efficient and practically implementable algorithm is proposed to pinpoint the forced oscillation source during real-time operation. Furthermore, we provide theoretical insights into the efficacy of the proposed approach by use of physical model-based analysis, in specific by establishing the fact that the rank of the resonance component matrix is at most 2. The effectiveness of the proposed method is validated in the IEEE 68-bus power system and the WECC 179-bus benchmark system.

eess.SP↗

Optimal Information Updating based on Value of Information

We address the problem of how to optimally schedule data packets over an unreliable channel in order to minimize the estimation error of a simple-to-implement remote linear estimator using a constant "Kalman'' gain to track the state of a Gauss Markov process. The remote estimator receives time-stamped data packets which contain noisy observations of the process. Additionally, they also contain the information about the "quality'' of the sensor\ source, i.e., the variance of the observation noise that was used to generate the packet. In order to minimize the estimation error, the scheduler needs to use both while prioritizing packet transmissions. It is shown that a simple index rule that calculates the value of information (VoI) of each packet, and then schedules the packet with the largest current value of VoI, is optimal. The VoI of a packet decreases with its age, and increases with the precision of the source. Thus, we conclude that, for constant filter gains, a policy which minimizes the age of information does not necessarily maximize the estimator performance.

cs.IT↗

The Strategic LQG System: A Dynamic Stochastic VCG Framework for Optimal Coordination

The classic Vickrey-Clarke-Groves (VCG) mechanism ensures incentive compatibility, i.e., that truth-telling of all agents is a dominant strategy, for a static one-shot game. However, in a dynamic environment that unfolds over time, the agents' intertemporal payoffs depend on the expected future controls and payments, and a direct extension of the VCG mechanism is not sufficient to guarantee incentive compatibility. In fact, it does not appear to be feasible to construct mechanisms that ensure the dominance of dynamic truth-telling for agents comprised of general stochastic dynamic systems. The contribution of this paper is to show that such a dynamic stochastic extension does exist for the special case of Linear-Quadratic-Gaussian (LQG) agents with a careful construction of a sequence of layered payments over time. For a set of LQG agents, we propose a modified layered version of the VCG mechanism for payments that decouples the intertemporal effect of current bids on future payoffs, and prove that truth-telling of dynamic states forms a dominant strategy if system parameters are known and agents are rational. An important example of a problem needing such optimal dynamic coordination of stochastic agents arises in power systems where an Independent System Operator (ISO) has to ensure balance of generation and consumption at all time instants, while ensuring social optimality. The challenge is to determine a bidding scheme between all agents and the ISO that maximizes social welfare, while taking into account the stochastic dynamic models of agents, since renewable energy resources such as solar/wind are stochastic and dynamic in nature, as are consumptions by loads which are influenced by factors such as local temperatures and thermal inertias of facilities.

eess.SY↗