Searcharxiv⌕ Search

arXiv subjects

Pingyi Fan

Publications and source records attributed to Pingyi Fan.

At least 109 records · Page 6Linked to original sources

Storage Space Allocation Strategy for Digital Data with Message Importance

This paper mainly focuses on the problem of lossy compression storage from the perspective of message importance when the reconstructed data pursues the least distortion within limited total storage size. For this purpose, we transform this problem to an optimization by means of the importance-weighted reconstruction error in data reconstruction. Based on it, this paper puts forward an optimal allocation strategy in the storage of digital data by a kind of restrictive water-filling. That is, it is a high efficient adaptive compression strategy since it can make rational use of all the storage space. It also characterizes the trade-off between the relative weighted reconstruction error and the available storage size. Furthermore, this paper also presents that both the users' preferences and the special characteristic of data distribution can trigger the small-probability event scenarios where only a fraction of data can cover the vast majority of users' interests. Whether it is for one of the reasons above, the data with highly clustered message importance is beneficial to compression storage. In contrast, the data with uniform information distribution is incompressible, which is consistent with that in information theory.

cs.IT↗

Energy Harvesting Powered Sensing in IoT: Timeliness Versus Distortion

We consider an Internet-of-Things (IoT) system in which an energy harvesting powered sensor node monitors the phenomenon of interest and transmits its observations to a remote monitor over a Gaussian channel. We measure the timeliness of the signals recovered by the monitor using age of information (AoI), which could be reduced by transmitting more observations to the monitor. We evaluate the corresponding distortion with the mean-squared error (MSE) metric, which would be reduced if a larger transmit power and a larger source coding rate were used. Since the energy harvested by the sensor node is random and limited, however, the timeliness and the distortion of the received signals cannot be optimized at the same time. Thus, we shall investigate the timeliness-distortion trade-off of the system by minimizing the average weighted-sum AoI and distortion over all possible transmit powers and transmission intervals. First, we explicitly present the optimal transmit powers for the performance limit achieving save-and-transmit policy and the easy-implementing fixed power transmission policy. Second, we propose a backward water-filling based offline power allocation algorithm and a genetic based offline algorithm to jointly optimize the transmission interval and transmit power. Third, we formulate the online power control as an Markov Decision Process (MDP) and solve the problem with an iterative algorithm, which closely approach the trade-off limit of the system. Also, we show that the optimal transmit power is a monotonic and bi-valued function of current AoI and distortion. Finally, we present our results via numerical simulations and extend results on the save-and-transmit policy to fading sensing systems.

cs.IT↗

Age-upon-Decisions Minimizing Scheduling in Internet of Things: To be Random or to be Deterministic?

We consider an Internet of Things (IoT) system in which a sensor delivers updates to a monitor with exponential service time and first-come-first-served (FCFS) discipline. We investigate the freshness of the received updates and propose a new metric termed as \textit{Age upon Decisions (AuD)}, which is defined as the time elapsed from the generation of each update to the epoch it is used to make decisions (e.g., estimations, inferences, controls). Within this framework, we aim at improving the freshness of updates at decision epochs by scheduling the update arrival process and the decision making process. Theoretical results show that 1) when the decisions are made according to a Poisson process, the average AuD is independent of decision rate and would be minimized if the arrival process is periodic (i.e., deterministic); 2) when both the decision process and the arrive process are periodic, the average AuD is larger than, but decreases with decision rate to, the average AuD of the corresponding system with Poisson decisions (i.e., random); 3) when both the decision process and the arrive process are periodic, the average AuD can be further decreased by optimally controlling the offset between the two processes. For practical IoT systems, therefore, it is suggested to employ periodic arrival processes and random decision processes. Nevertheless, making periodical updates and decisions with properly controlled offset also is a promising solution if the timing information of the two processes can be accessed by the monitor.

cs.IT↗

Importance of Small Probability Events in Big Data: Information Measures, Applications, and Challenges

In many applications (e.g., anomaly detection and security systems) of smart cities, rare events dominate the importance of the total information of big data collected by Internet of Things (IoTs). That is, it is pretty crucial to explore the valuable information associated with the rare events involved in minority subsets of the voluminous amounts of data. To do so, how to effectively measure the information with importance of the small probability events from the perspective of information theory is a fundamental question. This paper first makes a survey of some theories and models with respect to importance measures and investigates the relationship between subjective or semantic importance and rare events in big data. Moreover, some applications for message processing and data analysis are discussed in the viewpoint of information measures. In addition, based on rare events detection, some open challenges related to information measures, such as smart cities, autonomous driving, and anomaly detection in IoTs, are introduced which can be considered as future research directions.

cs.IT↗

Towards Big data processing in IoT: Path Planning and Resource Management of UAV Base Stations in Mobile-Edge Computing System

Heavy data load and wide cover range have always been crucial problems for online data processing in internet of things (IoT). Recently, mobile-edge computing (MEC) and unmanned aerial vehicle base stations (UAV-BSs) have emerged as promising techniques in IoT. In this paper, we propose a three-layer online data processing network based on MEC technique. On the bottom layer, raw data are generated by widely distributed sensors, which reflects local information. Upon them, unmanned aerial vehicle base stations (UAV-BSs) are deployed as moving MEC servers, which collect data and conduct initial steps of data processing. On top of them, a center cloud receives processed results and conducts further evaluation. As this is an online data processing system, the edge nodes should stabilize delay to ensure data freshness. Furthermore, limited onboard energy poses constraints to edge processing capability. To smartly manage network resources for saving energy and stabilizing delay, we develop an online determination policy based on Lyapunov Optimization. In cases of low data rate, it tends to reduce edge processor frequency for saving energy. In the presence of high data rate, it will smartly allocate bandwidth for edge data offloading. Meanwhile, hovering UAV-BSs bring a large and flexible service coverage, which results in the problem of effective path planning. In this paper, we apply deep reinforcement learning and develop an online path planning algorithm. Taking observations of around environment as input, a CNN network is trained to predict the reward of each action. By simulations, we validate its effectiveness in enhancing service coverage. The result will contribute to big data processing in future IoT.

cs.NI↗

Towards Big data processing in IoT: network management for online edge data processing

Heavy data load and wide cover range have always been crucial problems for internet of things (IoT). However, in mobile-edge computing (MEC) network, the huge data can be partly processed at the edge. In this paper, a MEC-based big data analysis network is discussed. The raw data generated by distributed network terminals are collected and processed by edge servers. The edge servers split out a large sum of redundant data and transmit extracted information to the center cloud for further analysis. However, for consideration of limited edge computation ability, part of the raw data in huge data sources may be directly transmitted to the cloud. To manage limited resources online, we propose an algorithm based on Lyapunov optimization to jointly optimize the policy of edge processor frequency, transmission power and bandwidth allocation. The algorithm aims at stabilizing data processing delay and saving energy without knowing probability distributions of data sources. The proposed network management algorithm may contribute to big data processing in future IoT.

cs.NI↗

Machine Learning Based Prediction and Classification of Computational Jobs in Cloud Computing Centers

With the rapid growth of the data volume and the fast increasing of the computational model complexity in the scenario of cloud computing, it becomes an important topic that how to handle users' requests by scheduling computational jobs and assigning the resources in data center. In order to have a better perception of the computing jobs and their requests of resources, we analyze its characteristics and focus on the prediction and classification of the computing jobs with some machine learning approaches. Specifically, we apply LSTM neural network to predict the arrival of the jobs and the aggregated requests for computing resources. Then we evaluate it on Google Cluster dataset and it shows that the accuracy has been improved compared to the current existing methods. Additionally, to have a better understanding of the computing jobs, we use an unsupervised hierarchical clustering algorithm, BIRCH, to make classification and get some interpretability of our results in the computing centers.

cs.LG↗

Fog-Assisted Multi-User SWIPT Networks: Local Computing or Offloading

This paper investigates a fog computing-assisted multi-user simultaneous wireless information and power transfer (SWIPT) network, where multiple sensors with power splitting (PS) receiver architectures receive information and harvest energy from a hybrid access point (HAP), and then process the received data by using local computing mode or fog offloading mode. For such a system, an optimization problem is formulated to minimize the sensors' required energy while guaranteeing their required information transmissions and processing rates by jointly optimizing the multi-user scheduling, the time assignment, the sensors' transmit powers and the PS ratios. Since the problem is a mixed integer programming (MIP) problem and cannot be solved with existing solution methods, we solve it by applying problem decomposition, variable substitutions and theoretical analysis. For a scheduled sensor, the closed-form and semi-closedform solutions to achieve its minimal required energy are derived, and then an efficient multi-user scheduling scheme is presented, which can achieve the suboptimal user scheduling with low computational complexity. Numerical results demonstrate our obtained theoretical results, which show that for each sensor, when it is located close to the HAP or the fog server (FS), the fog offloading mode is the better choice; otherwise, the local computing mode should be selected. The system performances in a frame-by-frame manner are also simulated, which show that using the energy stored in the batteries and that harvested from the signals transmitted by previous scheduled sensors can further decrease the total required energy of the sensors.

cs.IT↗

Matching Users' Preference Under Target Revenue Constraints in Optimal Data Recommendation Systems

This paper focuses on the problem of finding a particular data recommendation strategy based on the user preferences and a system expected revenue. To this end, we formulate this problem as an optimization by designing the recommendation mechanism as close to the user behavior as possible with a certain revenue constraint. In fact, the optimal recommendation distribution is the one that is the closest to the utility distribution in the sense of relative entropy and satisfies expected revenue. We show that the optimal recommendation distribution follows the same form as the message importance measure (MIM) if the target revenue is reasonable, i.e., neither too small nor too large. Therefore, the optimal recommendation distribution can be regarded as the normalized MIM, where the parameter, called importance coefficient, presents the concern of the system and switches the attention of the system over data sets with different occurring probability. By adjusting the importance coefficient, our MIM based framework of data recommendation can then be applied to system with various system requirements and data distributions.Therefore,the obtained results illustrate the physical meaning of MIM from the data recommendation perspective and validate the rationality of MIM in one aspect.

cs.IT↗

Optimal Online Transmission Policy in Wireless Powered Networks with Urgency-aware Age of Information

This paper investigates the age of information (AoI) for a radio frequency (RF) energy harvesting (EH) enabled network, where a sensor first scavenges energy from a wireless power station and then transmits the collected status update to a sink node. To capture the thirst for the fresh update becoming more and more urgent as time elapsing, urgency-aware AoI (U-AoI) is defined, which increases exponentially with time between two received updates. Due to EH, some waiting time is required at the sensor before transmitting the status update. To find the optimal transmission policy, an optimization problem is formulated to minimize the long-term average U-AoI under constraint of energy causality. As the problem is non-convex and with no known solution, a two-layer algorithm is presented to solve it, where the outer loop is designed based on Dinklebach's method, and in the inner loop, a semi-closed-form expression of the optimal waiting time policy is derived based on Karush-Kuhn-Tucker (KKT) optimality conditions. Numerical results shows that our proposed optimal transmission policy outperforms the the zero time waiting policy and equal time waiting policy in terms of long-term average U-AoI, especially when the networks are non-congested. It is also observed that in order to achieve the lower U-AoI, the sensor should transmit the next update without waiting when the network is congested while should wait a moment before transmitting the next update when the network is non-congested. Additionally, it also shows that the system U-AoI first decreases and then keep unchanged with the increments of EH circuit's saturation level and the energy outage probability.

cs.IT↗

Optimal Design of SWIPT-Aware Fog Computing Networks

This paper studies a simultaneous wireless information and power transfer (SWIPT)-aware fog computing network, where a multiple antenna fog function integrated hybrid access point (F-HAP) transfers information and energy to multiple heterogeneous single-antenna sensors and also helps some of them fulfill computing tasks. By jointly optimizing energy and information beamforming designs at the F-HAP, the bandwidth allocation and the computation offloading distribution, an optimization problem is formulated to minimize the required energy under communication and computation requirements, as well as energy harvesting constraints. Two optimal designs, i.e., fixed offloading time (FOT) and optimized offloading time (OOT) designs, are proposed. As both designs get involved in solving non-convex problems, there are no known solutions to them. Therefore, for the FOT design, the semidefinite relaxation (SDR) is adopted to solve it. It is theoretically proved that the rank-one constraints are always satisfied, so the global optimal solution is guaranteed. For the OOT design, since its non-convexity is hard to deal with, a penalty dual decomposition (PDD)-based algorithm is proposed, which is able to achieve a suboptimal solution. The computational complexity for two designs are analyzed. Numerical results show that the partial offloading mode is superior to binary benchmark modes. It is also shown that if the system is with strong enough computing capability, the OOT design is suggested to achieve lower required energy; Otherwise, the FOT design is preferred to achieve a relatively low computation complexity.

cs.CY↗

Information Measure Similarity Theory: Message Importance Measure via Shannon Entropy

Rare events attract more attention and interests in many scenarios of big data such as anomaly detection and security systems. To characterize the rare events importance from probabilistic perspective, the message importance measure (MIM) is proposed as a kind of semantics analysis tool. Similar to Shannon entropy, the MIM has its special functional on information processing, in which the parameter $\varpi$ of MIM plays a vital role. Actually, the parameter $\varpi$ dominates the properties of MIM, based on which the MIM has three work regions where the corresponding parameters satisfy $ 0 \le \varpi \le 2/\max\{p(x_i)\}$, $\varpi > 2/\max\{p(x_i)\}$ and $\varpi < 0$ respectively. Furthermore, in the case $ 0 \le \varpi \le 2/\max\{p(x_i)\}$, there are some similarity between the MIM and Shannon entropy in the information compression and transmission, which provide a new viewpoint for information theory. This paper first constructs a system model with message importance measure and proposes the message importance loss to enrich the information processing strategies. Moreover, we propose the message importance loss capacity to measure the information importance harvest in a transmission. Furthermore, the message importance distortion function is presented to give an upper bound of information compression based on message importance measure. Additionally, the bitrate transmission constrained by the message importance loss is investigated to broaden the scope for Shannon information theory.

cs.IT↗

State Variation Mining: On Information Divergence with Message Importance in Big Data

Information transfer which reveals the state variation of variables usually plays a vital role in big data analytics and processing. In fact, the measures for information transfer could reflect the system change by use of the variable distributions, similar to KL divergence and Renyi divergence. Furthermore, in terms of the information transfer in big data, small probability events usually dominate the importance of the total message to some degree. Therefore, it is significant to design an information transfer measure based on the message importance which emphasizes the small probability events. In this paper, we propose a message importance transfer measure (MITM) and investigate its characteristics and applications on three aspects. First, the message importance transfer capacity based on MITM is presented to offer an upper bound for the information transfer process with disturbance. Then, we extend the MITM to the continuous case and discuss the robustness by using it to measuring information distance. Finally, we utilize the MITM to guide the queue length selection in the caching operation of mobile edge computing.

cs.IT↗

Focusing on a Probability Element: Parameter Selection of Message Importance Measure in Big Data

Message importance measure (MIM) is applicable to characterize the importance of information in the scenario of big data, similar to entropy in information theory. In fact, MIM with a variable parameter can make an effect on the characterization of distribution. Furthermore, by choosing an appropriate parameter of MIM, it is possible to emphasize the message importance of a certain probability element in a distribution. Therefore, parametric MIM can play a vital role in anomaly detection of big data by focusing on probability of an anomalous event. In this paper, we propose a parameter selection method of MIM focusing on a probability element and then present its major properties. In addition, we discuss the parameter selection with prior probability, and investigate the availability in a statistical processing model of big data for anomaly detection problem.

cs.IT↗

Age of Information Upon Decisions

We consider an M/M/1 update-and-decide system where Poisson distributed decisions are made based on the received updates. We propose to characterize the freshness of the received updates at decision epochs with Age upon Decisions (AuD). Under the first-come-first-served policy (FCFS), the closed form average AuD is derived. We show that the average AuD of the system is determined by the arrival rate and the service rate, and is independent of the decision rate. Thus, merely increasing the decision rate does not improve the timeliness of decisions. Nevertheless, increasing the arrival rate and the service rate simultaneously can decrease the average AuD efficiently.

cs.IT↗

Minor probability events detection in big data: An integrated approach with Bayesian testing and MIM

The minor probability events detection is a crucial problem in Big data. Such events tend to include rarely occurring phenomenons which should be detected and monitored carefully. Given the prior probabilities of separate events and the conditional distributions of observations on the events, the Bayesian detection can be applied to estimate events behind the observations. It has been proved that Bayesian detection has the smallest overall testing error in average sense. However, when detecting an event with very small prior probability, the conditional Bayesian detection would result in high miss testing rate. To overcome such a problem, a modified detection approach is proposed based on Bayesian detection and message importance measure, which can reduce miss testing rate in conditions of detecting events with minor probability. The result can help to dig minor probability events in big data.

eess.SP↗

A Switch to the Concern of User: Importance Coefficient in Utility Distribution and Message Importance Measure

This paper mainly focuses on the utilization frequency in receiving end of communication systems, which shows the inclination of the user about different symbols. When the average number of use is limited, a specific utility distribution is proposed on the best effort in term of fairness, which is also the closest one to occurring probability in the relative entropy. Similar to a switch, its parameter can be selected to make it satisfy different users' requirements: negative parameter means the user focus on high-probability events and positive parameter means the user is interested in small-probability events. In fact, the utility distribution is a measure of message importance in essence. It illustrates the meaning of message importance measure (MIM), and extend it to the general case by selecting the parameter. Numerical results show that this utility distribution characterizes the message importance like MIM and its parameter determines the concern of users.

cs.IT↗

Uplink Age of Information of Unilaterally Powered Two-way Data Exchanging Systems

We consider a two-way data exchanging system where a master node transfers energy and data packets to a slave node alternatively. The slave node harvests the transferred energy and performs information transmission as long as it has sufficient energy for current block, i.e., according to the best-effort policy. We examine the freshness of the received packets at the master node in terms of age of information (AoI), which is defined as the time elapsed after the generation of the latest received packet. We derive average uplink AoI and uplink data rate as functions of downlink data rate in closed form. The obtained results illustrate the performance limit of the unilaterally powered two-way data exchanging system in terms of timeliness and efficiency. The results also specify the achievable tradeoff between the data rates of the two-way data exchanging system.

cs.IT↗