SearcharxivSearch

arXiv subjects

Vincent Leon

Publications and source records attributed to Vincent Leon.

10 recordsLinked to original sources

Certifying Concavity and Monotonicity in Games via Sum-of-Squares Hierarchies

Concavity and its refinements underpin tractability in multiplayer games, where players independently choose actions to maximize their own payoffs which depend on other players' actions. In concave games, where players' strategy sets are compact and convex, and their payoffs are concave in their own actions, strong guarantees follow: Nash equilibria always exist and decentralized algorithms converge to equilibria. If the game is furthermore monotone, an even stronger guarantee holds: Nash equilibria are unique under strictness assumptions. Unfortunately, we show that certifying concavity or monotonicity is NP-hard, already for games where utilities are multivariate polynomials and compact, convex basic semialgebraic strategy sets -- an expressive class that captures extensive-form games with imperfect recall. On the positive side, we develop two hierarchies of sum-of-squares programs that certify concavity and monotonicity of a given game, and each level of the hierarchies can be solved in polynomial time. We show that almost all concave/monotone games are certified at some finite level of the hierarchies. Subsequently, we introduce SOS-concave/monotone games, which globally approximate concave/monotone games, and show that for any given game we can compute the closest SOS-concave/monotone game in polynomial time. Finally, we apply our techniques to canonical examples of imperfect recall extensive-form games.

cs.GT

Online Learning for Dynamic Vickrey-Clarke-Groves Mechanism in Unknown Environments

We consider the problem of online dynamic mechanism design for sequential auctions in unknown environments, where the underlying market and, thus, the bidders' values vary over time as interactions between the seller and the bidders progress. We model the sequential auctions as an infinite-horizon average-reward Markov decision process (MDP). In each round, the seller determines an allocation and sets a payment for each bidder, while each bidder receives a private reward and submits a sealed bid to the seller. The state, which represents the underlying market, evolves according to an unknown transition kernel and the seller's allocation policy without episodic resets. We first extend the Vickrey-Clarke-Groves (VCG) mechanism to sequential auctions, thereby obtaining a dynamic counterpart that preserves the desired properties: efficiency, truthfulness, and individual rationality. We then focus on the online setting and develop a reinforcement learning algorithm for the seller to learn the underlying MDP and implement a mechanism that closely resembles the dynamic VCG mechanism. We show that the learned mechanism approximately satisfies efficiency, truthfulness, and individual rationality and achieves guaranteed performance in terms of various notions of regret.

cs.GT

Online Reinforcement Learning in Markov Decision Process Using Linear Programming

We consider online reinforcement learning in episodic Markov decision process (MDP) with unknown transition function and stochastic rewards drawn from some fixed but unknown distribution. The learner aims to learn the optimal policy and minimize their regret over a finite time horizon through interacting with the environment. We devise a simple and efficient model-based algorithm that achieves $\widetilde{O}(LX\sqrt{TA})$ regret with high probability, where $L$ is the episode length, $T$ is the number of episodes, and $X$ and $A$ are the cardinalities of the state space and the action space, respectively. The proposed algorithm, which is based on the concept of ``optimism in the face of uncertainty", maintains confidence sets of transition and reward functions and uses occupancy measures to connect the online MDP with linear programming. It achieves a tighter regret bound compared to the existing works that use a similar confidence set framework and improves computational effort compared to those that use a different framework but with a slightly tighter regret bound.

cs.LG

Limited-Trust in Diffusion of Competing Alternatives over Social Networks

We consider the diffusion of two alternatives in social networks using a game-theoretic approach. Each individual plays a coordination game with its neighbors repeatedly and decides which to adopt. As products are used in conjunction with others and through repeated interactions, individuals are more interested in their long-term benefits and tend to show trust to others to maximize their long-term utility by choosing a suboptimal option with respect to instantaneous payoff. To capture such trust behavior, we deploy limited-trust equilibrium (LTE) in diffusion process. We analyze the convergence of emerging dynamics to equilibrium points using mean-field approximation and study the equilibrium state and the convergence rate of diffusion using absorption probability and expected absorption time of a reduced-size absorbing Markov chain. We also show that the diffusion model on LTE under the best-response strategy can be converted to the well-known linear threshold model. Simulation results show that when agents behave trustworthy, their long-term utility will increase significantly compared to the case when they are solely self-interested. Moreover, the Markov chain analysis provides a good estimate of convergence properties over random networks.

cs.SI

Online Learning in Budget-Constrained Dynamic Colonel Blotto Games

In this paper, we study the strategic allocation of limited resources using a Colonel Blotto game (CBG) under a dynamic setting and analyze the problem using an online learning approach. In this model, one of the players is a learner who has limited troops to allocate over a finite time horizon, and the other player is an adversary. In each round, the learner plays a one-shot Colonel Blotto game with the adversary and strategically determines the allocation of troops among battlefields based on past observations. The adversary chooses its allocation action randomly from some fixed distribution that is unknown to the learner. The learner's objective is to minimize its regret, which is the difference between the cumulative reward of the best mixed strategy and the realized cumulative reward by following a learning algorithm while not violating the budget constraint. The learning in dynamic CBG is analyzed under the framework of combinatorial bandits and bandits with knapsacks. We first convert the budget-constrained dynamic CBG to a path planning problem on a directed graph. We then devise an efficient algorithm that combines a special combinatorial bandit algorithm for path planning problem and a bandits with knapsack algorithm to cope with the budget constraint. The theoretical analysis shows that the learner's regret is bounded by a term sublinear in time horizon and polynomial in other parameters. Finally, we justify our theoretical results by carrying out simulations for various scenarios.

cs.LG

Toward Optimal Adversarial Policies in the Multiplicative Learning System with a Malicious Expert

We consider a learning system based on the conventional multiplicative weight (MW) rule that combines experts' advice to predict a sequence of true outcomes. It is assumed that one of the experts is malicious and aims to impose the maximum loss on the system. The loss of the system is naturally defined to be the aggregate absolute difference between the sequence of predicted outcomes and the true outcomes. We consider this problem under both offline and online settings. In the offline setting where the malicious expert must choose its entire sequence of decisions a priori, we show somewhat surprisingly that a simple greedy policy of always reporting false prediction is asymptotically optimal with an approximation ratio of $1+O(\sqrt{\frac{\ln N}{N}})$, where $N$ is the total number of prediction stages. In particular, we describe a policy that closely resembles the structure of the optimal offline policy. For the online setting where the malicious expert can adaptively make its decisions, we show that the optimal online policy can be efficiently computed by solving a dynamic program in $O(N^3)$. Our results provide a new direction for vulnerability assessment of commonly used learning algorithms to adversarial attacks where the threat is an integral part of the system.

cs.LG

Spectroscopic study of double-walled carbon nanotubes functionalization for preparation of carbon nanotube / epoxy composites

A spectroscopic study of the amino functionalization of double-walled carbon nanotube (DWCNT) is performed. Original experimental investigations by near edge X-ray absorption fine structure spectroscopy at the C and O K-edges allow one to follow the efficiency of the chemistry during the different steps of covalent functionalization. Combined with Raman spectroscopy, the characterization gives a direct evidence of the grafting of amino-terminated molecules on the structural defects of the DWCNT external wall, whereas the internal wall does not undergo any change. Structural and mechanical investigation of the amino functionalized DWCNT / epoxy composites show coupling between epoxy molecules and the DWCNTs. Functionalization improves the interface between amino-functionalized DWCNT and the epoxy molecules. The electrical transport measurements indicate a percolating network formed only by inner metallic tubes of the DWCNTs. The activation energy of the barriers between connected metallic tubes is determined around 20 meV.

cond-mat.mtrl-sci

Effect of Nanoscale Confinement on the β-αPhase Transition in Ag2Se

The confinement of silver selenide was investigated using mesoporous silica. Results from x-ray diffraction and electron microscopy show that the confined material still exhibits a βto αtransition similar to the one that takes place in the bulk crystalline state but with a transition temperature that depends significantly on the confinement conditions. Decreasing the pore size leads to an increase of the transition temperature, opposite to the behavior of the melting point observed in several metallic and organic materials. In the free particles, on the other hand, no size dependence is observed with particle sizes down to 4 nm.

cond-mat.mtrl-sci

Nature of the bound states of molecular hydrogen in carbon nanohorns

The effects of confining molecular hydrogen within carbon nanohorns are studied via high-resolution quasielastic and inelastic neutron spectroscopies. Both sets of data are remarkably different from those obtained in bulk samples in the liquid and crystalline states. At temperatures where bulk hydrogen is liquid, the spectra of the confined sample show an elastic component indicating a significant proportion of immobile molecules as well as distinctly narrower quasielastic line widths and a strong distortion of the line shape of the para - ortho rotational transition. The results show that hydrogen interacts far more strongly with such carbonous structures than it does to carbon nanotubes, suggesting that nanohorns and related nanostructures may offer significantly better prospects as lightweight media for hydrogen storage applications.

cond-mat.mtrl-sci

Collective excitations in liquid D2 confined within the mesoscopic pores of a MCM-41 molecular sieve

We present a comparative study of the excitations in bulk and liquid D2 confined within the pores of MCM-41. The material (Mobile Crystalline Material-41) is a silicate obtained by means of a template that yields a partially crystalline structure composed by arrays of nonintersecting hexagonal channels of controlled width having walls made of amorphous SiO2. Its porosity was characterized by means of adsorption isotherms and found to be composed by a regular array of pores having a narrow distribution of sizes with a most probable value of 2.45 nm. The assessment of the precise location of the sample within the pores is carried out by means of pressure isotherms. The study was conducted at two pressures which correspond to pore fillings above the capillary condensation regime. Within the range of wave vectors where collective excitations can be followed up (0.3<Q<3.0 $Å$−1), we found confinement brings forward a large shortening of the excitation lifetimes that shifts the characteristic frequencies to higher energies. In addition, the coherent quasielastic scattering shows signatures of reduced diffusivity.

cond-mat.mtrl-sci