SearcharxivSearch

arXiv subjects

Xueying Guo

Publications and source records attributed to Xueying Guo.

At least 19 recordsLinked to original sources

Opportunistic Episodic Reinforcement Learning

In this paper, we propose and study opportunistic reinforcement learning - a new variant of reinforcement learning problems where the regret of selecting a suboptimal action varies under an external environmental condition known as the variation factor. When the variation factor is low, so is the regret of selecting a suboptimal action and vice versa. Our intuition is to exploit more when the variation factor is high, and explore more when the variation factor is low. We demonstrate the benefit of this novel framework for finite-horizon episodic MDPs by designing and evaluating OppUCRL2 and OppPSRL algorithms. Our algorithms dynamically balance the exploration-exploitation trade-off for reinforcement learning by introducing variation factor-dependent optimism to guide exploration. We establish an $\tilde{O}(HS \sqrt{AT})$ regret bound for the OppUCRL2 algorithm and show through simulations that both OppUCRL2 and OppPSRL algorithm outperform their original corresponding algorithms.

cs.LG

ETA Prediction with Graph Neural Networks in Google Maps

Travel-time prediction constitutes a task of high importance in transportation networks, with web mapping services like Google Maps regularly serving vast quantities of travel time queries from users and enterprises alike. Further, such a task requires accounting for complex spatiotemporal interactions (modelling both the topological properties of the road network and anticipating events -- such as rush hours -- that may occur in the future). Hence, it is an ideal target for graph representation learning at scale. Here we present a graph neural network estimator for estimated time of arrival (ETA) which we have deployed in production at Google Maps. While our main architecture consists of standard GNN building blocks, we further detail the usage of training schedule methods such as MetaGradients in order to make our model robust and production-ready. We also provide prescriptive studies: ablating on various architectural decisions and training regimes, and qualitative analyses on real-world situations where our model provides a competitive edge. Our GNN proved powerful when deployed, significantly reducing negative ETA outcomes in several regions compared to the previous production baseline (40+% in cities like Sydney).

cs.LG

A Giant Planet Candidate Transiting a White Dwarf

Astronomers have discovered thousands of planets outside the solar system, most of which orbit stars that will eventually evolve into red giants and then into white dwarfs. During the red giant phase, any close-orbiting planets will be engulfed by the star, but more distant planets can survive this phase and remain in orbit around the white dwarf. Some white dwarfs show evidence for rocky material floating in their atmospheres, in warm debris disks, or orbiting very closely, which has been interpreted as the debris of rocky planets that were scattered inward and tidally disrupted. Recently, the discovery of a gaseous debris disk with a composition similar to ice giant planets demonstrated that massive planets might also find their way into tight orbits around white dwarfs, but it is unclear whether the planets can survive the journey. So far, the detection of intact planets in close orbits around white dwarfs has remained elusive. Here, we report the discovery of a giant planet candidate transiting the white dwarf WD 1856+534 (TIC 267574918) every 1.4 days. The planet candidate is roughly the same size as Jupiter and is no more than 14 times as massive (with 95% confidence). Other cases of white dwarfs with close brown dwarf or stellar companions are explained as the consequence of common-envelope evolution, wherein the original orbit is enveloped during the red-giant phase and shrinks due to friction. In this case, though, the low mass and relatively long orbital period of the planet candidate make common-envelope evolution less likely. Instead, the WD 1856+534 system seems to demonstrate that giant planets can be scattered into tight orbits without being tidally disrupted, and motivates searches for smaller transiting planets around white dwarfs.

astro-ph.EP

Updated Parameters and a New Transmission Spectrum of HD 97658b

Recent years have seen increasing interest in the characterization of sub-Neptune sized planets because of their prevalence in the Galaxy, contrasted with their absence in our solar system. HD 97658 is one of the brightest stars hosting a planet of this kind, and we present the transmission spectrum of this planet by combining four HST transits, twelve Spitzer/IRAC transits, and eight MOST transits of this system. Our transmission spectrum has higher signal to noise ratio than that from previous works, and the result suggests that the slight increase in transit depth from wavelength 1.1 to 1.7 microns reported in previous works on the transmission spectrum of this planet is likely systematic. Nonetheless, our atmospheric modeling results are not conclusive as no model provides an excellent match to our data. Nonetheless we find that atmospheres with high C/O ratios (C/O >~ 0.8) and metallicities of >~ 100x solar metallicity are favored. We combine the mid-transit times from all the new Spitzer and MOST observations and obtain an updated orbital period of P=9.489295 +/- 0.000005 d, with a best-fit transit time center at T_0 = 2456361.80690 +/- 0.00038 (BJD). No transit timing variations are found in this system. We also present new measurements of the stellar rotation period (34 +/- 2 d) and stellar activity cycle (9.6 yr) of the host star HD 97658. Finally, we calculate and rank the Transmission Spectroscopy Metric of all confirmed planets cooler than 1000 K and with sizes between 1 and 4 R_Earth. We find that at least a third of small planets cooler than 1000 K can be well characterized using JWST, and of those, HD 97658b is ranked fifth, meaning it remains a high-priority target for atmospheric characterization.

astro-ph.EP

Zodiacal Exoplanets in Time (ZEIT) IX: a flat transmission spectrum and a highly eccentric orbit for the young Neptune K2-25b as revealed by Spitzer

Transiting planets in nearby young clusters offer the opportunity to study the atmospheres and dynamics of planets during their formative years. To this end, we focused on K2-25b -- a close-in ($P$=3.48 days), Neptune-sized exoplanet orbiting a M4.5 dwarf in the 650Myr Hyades cluster. We combined photometric observations of K2-25 covering a total of 44 transits and spanning >2 yr, drawn from a mix of space-based telescopes (Spitzer Space Telescope and K2) and ground-based facilities (Las Cumbres Observatory Global Telescope network and MEarth). The transit photometry spanned 0.6--4.5$μ$m, which enabled our study of K2-25b's transmission spectrum. We combined and fit each dataset at a common wavelength within a Markov Chain Monte Carlo framework, yielding consistent planet parameters. The resulting transit depths ruled out a solar-composition atmosphere for K2-25b for the range of expected planetary masses and equilibrium temperature at a $>4σ$ confidence level, and are consistent with a flat transmission spectrum. Mass constraints and transit observations at a finer grid of wavelengths (e.g., from the Hubble Space Telescope) are needed to make more definitive statements about the presence of clouds or an atmosphere of high mean molecular weight. Our precise measurements of K2-25b's transit duration also enabled new constraints on the eccentricity of K2-25's orbit. We find K2-25b's orbit to be eccentric ($e>0.20$) for all reasonable stellar densities and independent of the observation wavelength or instrument. The high eccentricity is suggestive of a complex dynamical history and motivates future searches for additional planets or stellar companions.

astro-ph.EP

Absence of a thick atmosphere on the terrestrial exoplanet LHS 3844b

Most known terrestrial planets orbit small stars with radii less than 60% that of the Sun. Theoretical models predict that these planets are more vulnerable to atmospheric loss than their counterparts orbiting Sun-like stars. To determine whether a thick atmosphere has survived on a small planet, one approach is to search for signatures of atmospheric heat redistribution in its thermal phase curve. Previous phase curve observations of the super-Earth 55 Cancri e (1.9 Earth radii) showed that its peak brightness is offset from the substellar point $-$ possibly indicative of atmospheric circulation. Here we report a phase curve measurement for the smaller, cooler planet LHS 3844b, a 1.3 Earth radius world in an 11-hour orbit around a small, nearby star. The observed phase variation is symmetric and has a large amplitude, implying a dayside brightness temperature of $1040\pm40$ kelvin and a nightside temperature consistent with zero kelvin (at one standard deviation). Thick atmospheres with surface pressures above 10 bar are ruled out by the data (at three standard deviations), and less-massive atmospheres are unstable to erosion by stellar wind. The data are well fitted by a bare rock model with a low Bond albedo (lower than 0.2 at two standard deviations). These results support theoretical predictions that hot terrestrial planets orbiting small stars may not retain substantial atmospheres.

astro-ph.EP

Kernel-based Multi-Task Contextual Bandits in Cellular Network Configuration

Cellular network configuration plays a critical role in network performance. In current practice, network configuration depends heavily on field experience of engineers and often remains static for a long period of time. This practice is far from optimal. To address this limitation, online-learning-based approaches have great potentials to automate and optimize network configuration. Learning-based approaches face the challenges of learning a highly complex function for each base station and balancing the fundamental exploration-exploitation tradeoff while minimizing the exploration cost. Fortunately, in cellular networks, base stations (BSs) often have similarities even though they are not identical. To leverage such similarities, we propose kernel-based multi-BS contextual bandit algorithm based on multi-task learning. In the algorithm, we leverage the similarity among different BSs defined by conditional kernel embedding. We present theoretical analysis of the proposed algorithm in terms of regret and multi-task-learning efficiency. We evaluate the effectiveness of our algorithm based on a simulator built by real traces.

cs.LG

AdaLinUCB: Opportunistic Learning for Contextual Bandits

In this paper, we propose and study opportunistic contextual bandits - a special case of contextual bandits where the exploration cost varies under different environmental conditions, such as network load or return variation in recommendations. When the exploration cost is low, so is the actual regret of pulling a sub-optimal arm (e.g., trying a suboptimal recommendation). Therefore, intuitively, we could explore more when the exploration cost is relatively low and exploit more when the exploration cost is relatively high. Inspired by this intuition, for opportunistic contextual bandits with Linear payoffs, we propose an Adaptive Upper-Confidence-Bound algorithm (AdaLinUCB) to adaptively balance the exploration-exploitation trade-off for opportunistic learning. We prove that AdaLinUCB achieves O((log T)^2) problem-dependent regret upper bound, which has a smaller coefficient than that of the traditional LinUCB algorithm. Moreover, based on both synthetic and real-world dataset, we show that AdaLinUCB significantly outperforms other contextual bandit algorithms, under large exploration cost fluctuations.

cs.LG

Adaptive Learning-Based Task Offloading for Vehicular Edge Computing Systems

The vehicular edge computing (VEC) system integrates the computing resources of vehicles, and provides computing services for other vehicles and pedestrians with task offloading. However, the vehicular task offloading environment is dynamic and uncertain, with fast varying network topologies, wireless channel states and computing workloads. These uncertainties bring extra challenges to task offloading. In this work, we consider the task offloading among vehicles, and propose a solution that enables vehicles to learn the offloading delay performance of their neighboring vehicles while offloading computation tasks. We design an adaptive learning-based task offloading (ALTO) algorithm based on the multi-armed bandit (MAB) theory, in order to minimize the average offloading delay. ALTO works in a distributed manner without requiring frequent state exchange, and is augmented with input-awareness and occurrence-awareness to adapt to the dynamic environment. The proposed algorithm is proved to have a sublinear learning regret. Extensive simulations are carried out under both synthetic scenario and realistic highway scenario, and results illustrate that the proposed algorithm achieves low delay performance, and decreases the average delay up to 30% compared with the existing upper confidence bound based learning algorithm.

cs.IT

Adaptive Exploration-Exploitation Tradeoff for Opportunistic Bandits

In this paper, we propose and study opportunistic bandits - a new variant of bandits where the regret of pulling a suboptimal arm varies under different environmental conditions, such as network load or produce price. When the load/price is low, so is the cost/regret of pulling a suboptimal arm (e.g., trying a suboptimal network configuration). Therefore, intuitively, we could explore more when the load/price is low and exploit more when the load/price is high. Inspired by this intuition, we propose an Adaptive Upper-Confidence-Bound (AdaUCB) algorithm to adaptively balance the exploration-exploitation tradeoff for opportunistic bandits. We prove that AdaUCB achieves $O(\log T)$ regret with a smaller coefficient than the traditional UCB algorithm. Furthermore, AdaUCB achieves $O(1)$ regret with respect to $T$ if the exploration cost is zero when the load level is below a certain threshold. Last, based on both synthetic data and real-world traces, experimental results show that AdaUCB significantly outperforms other bandit algorithms, such as UCB and TS (Thompson Sampling), under large load/price fluctuations.

cs.LG

Task Replication for Vehicular Edge Computing: A Combinatorial Multi-Armed Bandit based Approach

In vehicular edge computing (VEC) system, some vehicles with surplus computing resources can provide computation task offloading opportunities for other vehicles or pedestrians. However, vehicular network is highly dynamic, with fast varying channel states and computation loads. These dynamics are difficult to model or to predict, but they have major impact on the quality of service (QoS) of task offloading, including delay performance and service reliability. Meanwhile, the computing resources in VEC are often redundant due to the high density of vehicles. To improve the QoS of VEC and exploit the abundant computing resources on vehicles, we propose a learning-based task replication algorithm (LTRA) based on combinatorial multi-armed bandit (CMAB) theory, in order to minimize the average offloading delay. LTRA enables multiple vehicles to process the replicas of the same task simultaneously, and vehicles that require computing services can learn the delay performance of other vehicles while offloading tasks. We take the occurrence time of vehicles into consideration, and redesign the utility function of existing CMAB algorithm, so that LTRA can adapt to the time varying network topology of VEC. We use a realistic highway scenario to evaluate the delay performance and service reliability of LTRA through simulations, and show that compared with single task offloading, LTRA can improve the task completion ratio with deadline 0.6s from 80% to 98%.

cs.IT

Learning-Based Task Offloading for Vehicular Cloud Computing Systems

Vehicular cloud computing (VCC) is proposed to effectively utilize and share the computing and storage resources on vehicles. However, due to the mobility of vehicles, the network topology, the wireless channel states and the available computing resources vary rapidly and are difficult to predict. In this work, we develop a learning-based task offloading framework using the multi-armed bandit (MAB) theory, which enables vehicles to learn the potential task offloading performance of its neighboring vehicles with excessive computing resources, namely service vehicles (SeVs), and minimizes the average offloading delay. We propose an adaptive volatile upper confidence bound (AVUCB) algorithm and augment it with load-awareness and occurrence-awareness, by redesigning the utility function of the classic MAB algorithms. The proposed AVUCB algorithm can effectively adapt to the dynamic vehicular environment, balance the tradeoff between exploration and exploitation in the learning process, and converge fast to the optimal SeV with theoretical performance guarantee. Simulations under both synthetic scenario and a realistic highway scenario are carried out, showing that the proposed algorithm achieves close-to-optimal delay performance.

cs.IT

Temperate super-Earths/mini-Neptunes around M/K dwarfs Consist of 2 Populations Distinguished by Their Atmospheres

Studies of the atmospheres of hot Jupiters reveal a diversity of atmospheric composition and haze properties. Similar studies on individual smaller, temperate planets are rare due to the inherent difficulty of the observations and also to the average faintness of their host stars. To investigate their ensemble atmospheric properties, we construct a sample of 28 similar planets, all possess equilibrium temperature within 300-500K, have similar size (1-3 R_e), and orbit early M dwarfs and late K dwarfs with effective temperatures within a few hundred Kelvin of one another. In addition, NASA's Kepler/K2 and Spitzer missions gathered transit observations of each planet, producing an uniform transit data set both in wavelength and coarse planetary type. With the transits measured in Kepler's broad optical bandpass and Spitzer's 4.5 micron wavelength bandpass, we measure the transmission spectral slope, alpha, for the entire sample. While this measurement is too uncertain in nearly all cases to infer the properties of any individual planet, the distribution of alpha among several dozen similar planets encodes a key trend. We find that the distribution of alpha is not well-described by a single Gaussian distribution. Rather, a ratio of the Bayesian evidences between the likeliest 1-component and 2-component Gaussian models favors the latter by a ratio of 100:1. One Gaussian is centered around an average alpha=-1.3, indicating hazy/cloudy atmospheres or bare cores with atmosphere evaporated. A smaller but significant second population (20+\-10% of all) is necessary to model significantly higher alpha values, which indicate atmospheres with potentially detectable molecular features. We conclude that the atmospheres of small and temperate planets are far from uniformly flat, and that a subset are particularly favorable for follow-up observation from space-based platforms like HST and JWST.

astro-ph.EP

Task Replication for Deadline-Constrained Vehicular Cloud Computing: Optimal Policy, Performance Analysis and Implications on Road Traffic

In vehicular cloud computing (VCC) systems, the computational resources of moving vehicles are exploited and managed by infrastructures, e.g., roadside units, to provide computational services. The offloading of computational tasks and collection of results rely on successful transmissions between vehicles and infrastructures during encounters. In this paper, we investigate how to provide timely computational services in VCC systems. In particular, we seek to minimize the deadline violation probability given a set of tasks to be executed in vehicular clouds. Due to the uncertainty of vehicle movements, the task replication methodology is leveraged which allows one task to be executed by several vehicles, and thus trading computational resources for delay reduction. The optimal task replication policy is of key interest. We first formulate the problem as a finite-horizon sampled-time Markov decision problem and obtain the optimal policy by value iterations. To conquer the complexity issue, we propose the balanced-task-assignment (BETA) policy which is proved optimal and has a clear structure: it always assigns the task with the minimum number of replicas. Moreover, a tight closed-form performance upper bound for the BETA policy is derived, which indicates that the deadline violation probability follows the Rayleigh distribution approximately. Applying the vehicle speed-density relationship in the traffic flow theory, we find that vehicle mobility benefits VCC systems more compared with road traffic systems, by showing that the optimum vehicle speed to minimize the deadline violation probability is larger than the critical vehicle speed in traffic theory which maximizes traffic flow efficiency.

cs.IT

Risk-Sensitive Optimal Control of Queues

We consider the problem of designing risk-sensitive optimal control policies for scheduling packet transmissions in a stochastic wireless network. A single client is connected to an access point (AP) through a wireless channel. Packet transmission incurs a cost $C$, while packet delivery yields a reward of $R$ units. The client maintains a finite buffer of size $B$, and a penalty of $L$ units is imposed upon packet loss which occurs due to finite queueing buffer. We show that the risk-sensitive optimal control policy for such a simple set-up is of threshold type, i.e., it is optimal to carry out packet transmissions only when $Q(t)$, i.e., the queue length at time $t$ exceeds a certain threshold $τ$. It is also shown that the value of threshold $τ$ increases upon increasing the cost per unit packet transmission $C$. Furthermore, it is also shown that a threshold policy with threshold equal to $τ$ is optimal for a set of problems in which cost $C$ lies within an interval $[C_l,C_u]$. Equations that need to be solved in order to obtain $C_l,C_u$ are also provided.

eess.SY

The metallicity distribution and hot Jupiter rate of the Kepler field: Hectochelle High-resolution spectroscopy for 776 Kepler target stars

The occurrence rate of hot Jupiters from the Kepler transit survey is roughly half that of radial velocity surveys targeting solar neighborhood stars. One hypothesis to explain this difference is that the two surveys target stars with different stellar metallicity distributions. To test this hypothesis, we measure the metallicity distribution of the Kepler targets using the Hectochelle multi-fiber, high-resolution spectrograph. Limiting our spectroscopic analysis to 610 dwarf stars in our sample with log(g)>3.5, we measure a metallicity distribution characterized by a mean of [M/H]_{mean} = -0.045 +/- 0.00, in agreement with previous studies of the Kepler field target stars. In comparison, the metallicity distribution of the California Planet Search radial velocity sample has a mean of [M/H]_{CPS, mean} = -0.005 +/- 0.006, and the samples come from different parent populations according to a Kolmogorov-Smirnov test. We refit the exponential relation between the fraction of stars hosting a close-in giant planet and the host star metallicity using a sample of dwarf stars from the California Planet Search with updated metallicities. The best-fit relation tells us that the difference in metallicity between the two samples is insufficient to explain the discrepant Hot Jupiter occurrence rates; the metallicity difference would need to be $\simeq$0.2-0.3 dex for perfect agreement. We also show that (sub)giant contamination in the Kepler sample cannot reconcile the two occurrence calculations. We conclude that other factors, such as binary contamination and imperfect stellar properties, must also be at play.

astro-ph.SR

A Cooperative Scheduling Scheme of Local Cloud and Internet Cloud for Delay-Aware Mobile Cloud Computing

With the proliferation of mobile applications, Mobile Cloud Computing (MCC) has been proposed to help mobile devices save energy and improve computation performance. To further improve the quality of service (QoS) of MCC, cloud servers can be deployed locally so that the latency is decreased. However, the computational resource of the local cloud is generally limited. In this paper, we design a threshold-based policy to improve the QoS of MCC by cooperation of the local cloud and Internet cloud resources, which takes the advantages of low latency of the local cloud and abundant computational resources of the Internet cloud simultaneously. This policy also applies a priority queue in terms of delay requirements of applications. The optimal thresholds depending on the traffic load is obtained via a proposed algorithm. Numerical results show that the QoS can be greatly enhanced with the assistance of Internet cloud when the local cloud is overloaded. Better QoS is achieved if the local cloud order tasks according to their delay requirements, where delay-sensitive applications are executed ahead of delay-tolerant applications. Moreover, the optimal thresholds of the policy have a sound impact on the QoS of the system.

cs.NI

A High Reliability Asymptotic Approach for Packet Inter-Delivery Time Optimization in Cyber-Physical Systems

In cyber-physical systems such as automobiles, measurement data from sensor nodes should be delivered to other consumer nodes such as actuators in a regular fashion. But, in practical systems over unreliable media such as wireless, it is a significant challenge to guarantee small enough inter-delivery times for different clients with heterogeneous channel conditions and inter-delivery requirements. In this paper, we design scheduling policies aiming at satisfying the inter-delivery requirements of such clients. We formulate the problem as a risk-sensitive Markov Decision Process (MDP). Although the resulting problem involves an infinite state space, we first prove that there is an equivalent MDP involving only a finite number of states. Then we prove the existence of a stationary optimal policy and establish an algorithm to compute it in a finite number of steps. However, the bane of this and many similar problems is the resulting complexity, and, in an attempt to make fundamental progress, we further propose a new high reliability asymptotic approach. In essence, this approach considers the scenario when the channel failure probabilities for different clients are of the same order, and asymptotically approach zero. We thus proceed to determine the asymptotically optimal policy: in a two-client scenario, we show that the asymptotically optimal policy is a "modified least time-to-go" policy, which is intuitively appealing and easily implementable; in the general multi-client scenario, we are led to an SN policy, and we develop an algorithm of low computational complexity to obtain it. Simulation results show that the resulting policies perform well even in the pre-asymptotic regime with moderate failure probabilities.

cs.NI