SearcharxivSearch

arXiv subjects

Arghyadip Roy

Publications and source records attributed to Arghyadip Roy.

8 recordsLinked to original sources

Adaptive KL-UCB based Bandit Algorithms for Markovian and i.i.d. Settings

In the regret-based formulation of Multi-armed Bandit (MAB) problems, except in rare instances, much of the literature focuses on arms with i.i.d. rewards. In this paper, we consider the problem of obtaining regret guarantees for MAB problems in which the rewards of each arm form a Markov chain which may not belong to a single parameter exponential family. To achieve a logarithmic regret in such problems is not difficult: a variation of standard Kullback-Leibler Upper Confidence Bound (KL-UCB) does the job. However, the constants obtained from such an analysis are poor for the following reason: i.i.d. rewards are a special case of Markov rewards and it is difficult to design an algorithm that works well independent of whether the underlying model is truly Markovian or i.i.d. To overcome this issue, we introduce a novel algorithm that identifies whether the rewards from each arm are truly Markovian or i.i.d. using a total variation distance-based test. Our algorithm then switches from using a standard KL-UCB to a specialized version of KL-UCB when it determines that the arm reward is Markovian, thus resulting in low regrets for both i.i.d. and Markovian settings.

cs.LG

A Policy Gradient Algorithm for the Risk-Sensitive Exponential Cost MDP

We study the risk-sensitive exponential cost MDP formulation and develop a trajectory-based gradient algorithm to find the stationary point of the cost associated with a set of parameterized policies. We derive a formula that can be used to compute the policy gradient from (state, action, cost) information collected from sample paths of the MDP for each fixed parameterized policy. Unlike the traditional average-cost problem, standard stochastic approximation theory cannot be used to exploit this formula. To address the issue, we introduce a truncated and smooth version of the risk-sensitive cost and show that this new cost criterion can be used to approximate the risk-sensitive cost and its gradient uniformly under some mild assumptions. We then develop a trajectory-based gradient algorithm to minimize the smooth truncated estimation of the risk-sensitive cost and derive conditions under which a sequence of truncations can be used to solve the original, untruncated cost problem.

eess.SY

Online Reinforcement Learning of Optimal Threshold Policies for Markov Decision Processes

To overcome the curses of dimensionality and modeling of Dynamic Programming (DP) methods to solve Markov Decision Process (MDP) problems, Reinforcement Learning (RL) methods are adopted in practice. Contrary to traditional RL algorithms which do not consider the structural properties of the optimal policy, we propose a structure-aware learning algorithm to exploit the ordered multi-threshold structure of the optimal policy, if any. We prove the asymptotic convergence of the proposed algorithm to the optimal policy. Due to the reduction in the policy space, the proposed algorithm provides remarkable improvements in storage and computational complexities over classical RL algorithms. Simulation results establish that the proposed algorithm converges faster than other RL algorithms.

cs.LG

Control and Management of Multiple RATs in Wireless Networks: An SDN Approach

Telecom operators are using a variety of Radio Access Technologies (RATs) for providing services to mobile subscribers. This development has emphasized the requirement for unified control and management of diverse RATs. Although multiple RATs co-exist within today's cellular networks, each RAT is controlled by a set of different entities. This may lead to suboptimal utilization of the overall network resources. In this article, we review various architectures for multi-RAT control proposed by both industry and academia. We also propose a novel SDN based network architecture for end-to-end control and management of diverse RATs. The architecture is scalable and provides a framework for improved network performance over the present day architecture and proposals in existing literature. Our architecture also provides a framework for deployment of applications in a RAT agnostic fashion. It facilitates network slicing and enables the provision of Quality of Service (QoS) guarantees to the end user. We have also developed an evaluation platform based on ns-3 to evaluate the performance offered by the architecture. Experimental results obtained using the platform demonstrate the benefits provided by our architecture.

eess.SP

A Structure-aware Online Learning Algorithm for Markov Decision Processes

To overcome the curse of dimensionality and curse of modeling in Dynamic Programming (DP) methods for solving classical Markov Decision Process (MDP) problems, Reinforcement Learning (RL) algorithms are popular. In this paper, we consider an infinite-horizon average reward MDP problem and prove the optimality of the threshold policy under certain conditions. Traditional RL techniques do not exploit the threshold nature of optimal policy while learning. In this paper, we propose a new RL algorithm which utilizes the known threshold structure of the optimal policy while learning by reducing the feasible policy space. We establish that the proposed algorithm converges to the optimal policy. It provides a significant improvement in convergence speed and computational and storage complexity over traditional RL algorithms. The proposed technique can be applied to a wide variety of optimization problems that include energy efficient data transmission and management of queues. We exhibit the improvement in convergence speed of the proposed algorithm over other RL algorithms through simulations.

cs.LG

An Energy-Aware WLAN Discovery Scheme for LTE HetNet

Recently, there has been significant interest in the integration and co-existence of Third Generation Partnership Project (3GPP) Long Term Evolution (LTE) with other Radio Access Technologies, like IEEE 802.11 Wireless Local Area Networks (WLANs). Although, the inter-working of IEEE 802.11 WLANs with 3GPP LTE has indicated enhanced network performance in the context of capacity and load balancing, the WLAN discovery scheme implemented in most of the commercially available smartphones is very inefficient and results in high battery drainage. In this paper, we have proposed an energy efficient WLAN discovery scheme for 3GPP LTE and IEEE 802.11 WLAN inter-working scenario. User Equipment (UE), in the proposed scheme, uses 3GPP network assistance along with the results of past channel scans, to optimally select the next channels to scan. Further, we have also developed an algorithm to accurately estimate the UE's mobility state, using 3GPP network signal strength patterns. We have implemented various discovery schemes in Android framework, to evaluate the performance of our proposed scheme against other solutions in the literature. Since, Android does not support selective scanning mode, we have implemented modules in Android to enable selective scanning. Further, we have also used simulation studies and justified the results using power consumption modeling. The results from the field experiments and simulations have shown high power savings using the proposed scanning scheme without any discovery performance deterioration.

cs.NI

Optimal Traffic Splitting Policy in LTE-based Heterogeneous Network

Dual Connectivity (DC) is a technique proposed to address the problem of increased handovers in heterogeneous networks. In DC, a foreground User Equipment (UE) with multiple transceivers has a possibility to connect to a Macro eNodeB (MeNB) and a Small cell eNodeB (SeNB) simultaneously. In downlink split bearer architecture of DC, a data radio bearer at MeNB gets divided into two; one part is forwarded to the SeNB through a non-ideal backhaul link to the UE, and the other part is forwarded by the MeNB. This may lead to an increase in the total delay at the UE since different packets corresponding to a single transmission may incur varying amounts of delays in the two different paths. Since the resources in the MeNB are shared by background legacy users and foreground users, DC may increase the blocking probability of background users. Moreover, single connectivity to the small cell may increase the blocking probability of foreground users. Therefore, we target to minimize the average delay of the system subject to a constraint on the blocking probability of background and foreground users. The optimal policy is computed and observed to contain a threshold structure. The variation of average system delay is studied for changes in different system parameters.

cs.NI

Performance Evaluation of Optimal Radio Access Technology Selection Algorithms for LTE-WiFi Network

A Heterogeneous Network (HetNet) comprises of multiple Radio Access Technologies (RATs) allowing a user to associate with a specific RAT and steer to other RATs in a seamless manner. To cope up with the unprecedented growth of data traffic, mobile data can be offloaded to Wireless Fidelity (WiFi) in a Long Term Evolution (LTE) based HetNet. In this paper, an optimal RAT selection problem is considered to maximize the total system throughput in an LTE-WiFi system with offload capability. Another formulation is also developed where maximizing the total system throughput is subject to a constraint on the voice user blocking probability. It is proved that the optimal policies for the association and offloading of voice/data users contain threshold structures. Based on the threshold structures, we propose algorithms for the association and offloading of users in LTE-WiFi HetNet. Simulation results are presented to demonstrate the voice user blocking probability and the total system throughput performance of the proposed algorithms in comparison to another benchmark algorithm.

cs.NI