SearcharxivSearch

arXiv subjects

William Chang

Publications and source records attributed to William Chang.

At least 19 recordsLinked to original sources

The Price of Decentralization in Top-$K$ Arm Identification

Cooperative teams often need to agree on the best few options rather than simply accumulate reward, and they must do so while each member sees only a fragment of the team's collective experience. We study this as top-$K$ joint-arm identification in multi-agent multi-armed bandits: at every round $M$ agents simultaneously choose individual actions that compose a joint arm, and the team must ultimately return the $K$ joint arms of highest mean reward. The difficulty is that an agent may not observe the actions of others, their rewards, or either. We treat three observability regimes---(A) shared rewards with hidden actions, (B) observed actions with private rewards, and (C) full asymmetry---and design communication-free elimination algorithms (UCB-Intervals) that reconstruct implicit coordination from whatever signal each regime leaves intact: a shared arm ordering in (A), observable deviations in (B), and enlarged confidence radii under (C). We give matching analyses in both the fixed-budget and fixed-confidence objectives, then fold all three regimes into a single meta-guarantee indexed by a multiplicity $c$ and a consensus factor $\rho$. Our central result is quantitative rather than merely algorithmic: change-of-measure lower bounds show that shared-reward identification is optimal up to one universal logarithmic factor, and that the entire statistical price of removing communication is a multiplicative $\rho^2$ in sample complexity---a fixed $4\times$ penalty under full asymmetry. The resulting stopping time scales as $O\!\left(\sum_{\mathbf{a}} \frac{\log(A^M/\delta)}{\Delta_{\mathbf{a}}^2}\right)$ and the fixed-budget error as $\exp(-\Theta(T/H_1))$, with the dependence on the joint-action count $A^M$ shown to be unavoidable.

cs.LG

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards. Players cannot communicate during learning but may agree on a protocol a priori. For Problems A and B we propose \texttt{mQ-learning} and \texttt{mQ-learning-intervals}, achieving $\tilde{O}(\sqrt{H^4 S A_{\text{joint}}\, T})$ regret, where $H$ is the horizon, $S$ the state count, $T = KH$ the total steps, and $A_{\text{joint}} = \prod_{i=1}^M |\mathcal{A}_i|$ the joint action space across $M$ players. For Problem C we give \texttt{mEXC} and \texttt{mEXC-Bellman}, two-phase explore-then-commit algorithms with regret $\tilde{O}(H (S A_{\text{joint}})^{1/3} T^{2/3})$. Against the centralized joint-action benchmark, decentralized learning under information asymmetry matches the single-agent Q-learning rate of \cite{jin2018q} up to logarithmic factors. Because $A_{\text{joint}}$ grows exponentially in $M$, the bounds are most meaningful for small $M$ or small per-player action sets.

cs.LG

DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks

In this work, we extend the Dependent Click Model (DCM) Bandits to a multiplayer information-asymmetric setting, where multiple agents interact with a shared ranked list and may observe multiple clicks per session, introducing new challenges for selection strategies. We study asymmetry in (1) actions and (2) rewards, providing sublinear regret guarantees for three settings where at least one asymmetry is present. Establishing matching information-theoretic lower bounds for these settings is left as an open problem. We further show that for small termination probabilities, the termination ranking need not be known, improving on prior single-agent results. Experiments confirm that our algorithms perform well across asymmetric environments and highlight the critical role of feedback structure, specifically the distinction between full versus first-click feedback, in coordinating exploration and minimizing regret.

cs.LG

Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. We consider three information structures: (A)~unobserved actions with common rewards, (B)~observed actions with independent rewards, and (C)~unobserved actions with independent rewards. In each case we design and analyze an algorithm that estimates the Lipschitz constant, chooses a discretization of the joint action space, and applies a cooperative bandit method to the induced discrete problem. Players never communicate once learning starts, so the central difficulty is that they must reach the \emph{same} discretization from their own data. We prove regret guarantees showing that common rewards and observable actions each supply this agreement for free, and that in their absence agreement can still be bought, through a dithered quantization of the estimate, at no cost in the leading order of the regret.

cs.LG

Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. However, real-world applications often involve heavy-tailed reward distributions and decentralized, information-asymmetric interactions. We study multi-agent multi-armed bandits with heavy-tailed rewards under three information-asymmetry regimes: unobserved actions with common rewards, observed actions with independent rewards, and unobserved actions with independent rewards. We develop robust decentralized algorithms for each setting and derive regret guarantees that nearly match centralized heavy-tailed rates. Experiments on a Pareto-distributed reward environment validate our theoretical findings and illustrate the trade-offs between synchronization, coordination, and exploration across the three regimes.

cs.LG

Generalised higher order vectorial $\infty$-eigenvalue problems

We study the problem of minimising the $L^\infty$ norm of a function of the $k$-th derivative over a class of maps, subject to a constraint involving the $L^\infty$ norm of a function of the map and its lower-order derivatives, for any integer $k\geq 2$. We impose boundary conditions corresponding to the $k$-th order analogues of the classical ``clamped'' and ``hinged'' cases. By employing the method of $L^p$ approximations, we establish the existence of a special $L^\infty$ minimiser, which solves a divergence PDE system with measure coefficients as parameters. This system constitutes the counterpart of the Aronsson--Euler equations for the constrained variational problem under consideration. Furthermore, we establish a lower bound for the eigenvalue. The present work extends the second-order vectorial results of Clark and Katzourakis [Generalised second order vectorial $\infty$-eigenvalue problems, PRSE A, 1-21, 2024] to the general higher-order setting.

math.AP

Accelerating Low-Frequency Convergence for Limited-Angle DBT via Two-Channel Fidelity in PDHG

Reconstruction in limited-angle digital breast tomosynthesis (DBT) suffers from slow convergence of low spatial-frequency components when using weighted data-fidelity terms within primal-dual optimization. We introduce a two-channel fidelity strategy that decomposes the sinogram residual into complementary low-pass and high-pass bands using square-root Hanning (Hann^{1/2}) filter families, each driven by an independent \ell_2-ball constraint and dual update in the PDHG (Chambolle-Pock) algorithm with He-Yuan predictor-corrector relaxation. By assigning a larger dual step size and slightly looser tolerance to the low-frequency channel, the method delivers stronger per-iteration correction to the near-DC band without violating global PDHG stability. Experiments on a 2D digital breast phantom across multiple resolutions demonstrate that the two-channel approach yields 19%--61% RMSE improvement over the single-channel baseline, with larger gains at coarser discretizations where problem conditioning is more favorable, supporting more balanced spectral convergence in clinically realistic limited-angle regimes.

math.OC

Multiplayer Information Asymmetric Bandits in Metric Spaces

In recent years the information asymmetric Lipschitz bandits In this paper we studied the Lipschitz bandit problem applied to the multiplayer information asymmetric problem studied in \cite{chang2022online, chang2023optimal}. More specifically we consider information asymmetry in rewards, actions, or both. We adopt the CAB algorithm given in \cite{kleinberg2004nearly} which uses a fixed discretization to give regret bounds of the same order (in the dimension of the action) space in all 3 problem settings. We also adopt their zooming algorithm \cite{ kleinberg2008multi}which uses an adaptive discretization and apply it to information asymmetry in rewards and information asymmetry in actions.

cs.LG

Multiplayer Information Asymmetric Contextual Bandits

Single-player contextual bandits are a well-studied problem in reinforcement learning that has seen applications in various fields such as advertising, healthcare, and finance. In light of the recent work on \emph{information asymmetric} bandits \cite{chang2022online, chang2023online}, we propose a novel multiplayer information asymmetric contextual bandit framework where there are multiple players each with their own set of actions. At every round, they observe the same context vectors and simultaneously take an action from their own set of actions, giving rise to a joint action. However, upon taking this action the players are subjected to information asymmetry in (1) actions and/or (2) rewards. We designed an algorithm \texttt{LinUCB} by modifying the classical single-player algorithm \texttt{LinUCB} in \cite{chu2011contextual} to achieve the optimal regret $O(\sqrt{T})$ when only one kind of asymmetry is present. We then propose a novel algorithm \texttt{ETC} that is built on explore-then-commit principles to achieve the same optimal regret when both types of asymmetry are present.

cs.LG

Mixing on Generalized Associahedra

Eppstein and Frishberg recently proved that the mixing time for the simple random walk on the $1$-skeleton of the associahedron is $O(n^3\log^3 n)$. We obtain similar rapid mixing results for the simple random walks on the $1$-skeleta of the type-$B$ and type-$D$ associahedra. We adapt Eppstein and Frishberg's technique to obtain the same bound of $O(n^3\log^3 n)$ in type $B$ and a bound of $O(n^{13} \log^2 n)$ in type $D$; in the process, we establish an expansion bound that is tight up to logarithmic factors in type $B$.

math.CO

Finite-Time Frequentist Regret Bounds of Multi-Agent Thompson Sampling on Sparse Hypergraphs

We study the multi-agent multi-armed bandit (MAMAB) problem, where $m$ agents are factored into $\rho$ overlapping groups. Each group represents a hyperedge, forming a hypergraph over the agents. At each round of interaction, the learner pulls a joint arm (composed of individual arms for each agent) and receives a reward according to the hypergraph structure. Specifically, we assume there is a local reward for each hyperedge, and the reward of the joint arm is the sum of these local rewards. Previous work introduced the multi-agent Thompson sampling (MATS) algorithm \citep{verstraeten2020multiagent} and derived a Bayesian regret bound. However, it remains an open problem how to derive a frequentist regret bound for Thompson sampling in this multi-agent setting. To address these issues, we propose an efficient variant of MATS, the $\epsilon$-exploring Multi-Agent Thompson Sampling ($\epsilon$-MATS) algorithm, which performs MATS exploration with probability $\epsilon$ while adopts a greedy policy otherwise. We prove that $\epsilon$-MATS achieves a worst-case frequentist regret bound that is sublinear in both the time horizon and the local arm size. We also derive a lower bound for this setting, which implies our frequentist regret upper bound is optimal up to constant and logarithm terms, when the hypergraph is sufficiently sparse. Thorough experiments on standard MAMAB problems demonstrate the superior performance and the improved computational efficiency of $\epsilon$-MATS compared with existing algorithms in the same setting.

cs.LG

Optimal Cooperative Multiplayer Learning Bandits with Noisy Rewards and No Communication

We consider a cooperative multiplayer bandit learning problem where the players are only allowed to agree on a strategy beforehand, but cannot communicate during the learning process. In this problem, each player simultaneously selects an action. Based on the actions selected by all players, the team of players receives a reward. The actions of all the players are commonly observed. However, each player receives a noisy version of the reward which cannot be shared with other players. Since players receive potentially different rewards, there is an asymmetry in the information used to select their actions. In this paper, we provide an algorithm based on upper and lower confidence bounds that the players can use to select their optimal actions despite the asymmetry in the reward information. We show that this algorithm can achieve logarithmic $O(\frac{\log T}{\Delta_{\bm{a}}})$ (gap-dependent) regret as well as $O(\sqrt{T\log T})$ (gap-independent) regret. This is asymptotically optimal in $T$. We also show that it performs empirically better than the current state of the art algorithm for this environment.

cs.LG

Closure of Certain Matrix Varieties and Applications

We prove some results about closures of certain matrix varieties consisting of elements with the same centralizer dimension. This generalizes a result of Dixmier and has applications to topological generation of simple algebraic groups.

math.AG

No-Regret Online Reinforcement Learning with Adversarial Losses and Transitions

Existing online learning algorithms for adversarial Markov Decision Processes achieve ${O}(\sqrt{T})$ regret after $T$ rounds of interactions even if the loss functions are chosen arbitrarily by an adversary, with the caveat that the transition function has to be fixed. This is because it has been shown that adversarial transition functions make no-regret learning impossible. Despite such impossibility results, in this work, we develop algorithms that can handle both adversarial losses and adversarial transitions, with regret increasing smoothly in the degree of maliciousness of the adversary. More concretely, we first propose an algorithm that enjoys $\widetilde{{O}}(\sqrt{T} + C^{\textsf{P}})$ regret where $C^{\textsf{P}}$ measures how adversarial the transition functions are and can be at most ${O}(T)$. While this algorithm itself requires knowledge of $C^{\textsf{P}}$, we further develop a black-box reduction approach that removes this requirement. Moreover, we also show that further refinements of the algorithm not only maintains the same regret bound, but also simultaneously adapts to easier environments (where losses are generated in a certain stochastically constrained manner as in Jin et al. [2021]) and achieves $\widetilde{{O}}(U + \sqrt{UC^{\textsf{L}}} + C^{\textsf{P}})$ regret, where $U$ is some standard gap-dependent coefficient and $C^{\textsf{L}}$ is the amount of corruption on losses.

cs.LG

Bode Integral Limitation For Irrational Systems

Bode integrals of sensitivity and sensitivity-like functions along with complementary sensitivity and complementary sensitivity-like functions are conventionally used for describing performance limitations of a feedback control system. In this paper, we investigate the Bode integral and evaluate what happens when a fractional order Proportional-Integral-Derivative (PID) controller is used in a feedback control system. We extend our analysis to when fractal PID controllers are applied to irrational systems. We split this into two cases: when the sequence of infinitely many right half plane open-loop poles doesn't have any limit points and when it does have a limit point. In both cases, we prove that the structure of the Bode Integral is similar to the classical version under certain conditions of convergence. We also provide a sufficient condition for the controller to lower the Bode sensitivity integral.

eess.SY

Finite-time self-similar rupture in a generalized elastohydrodynamic lubrication model

Thin film rupture is a type of nonlinear instability that causes the solution to touch down to zero at finite time. We investigate the finite-time rupture behavior of a generalized elastohydrodynamic lubrication model. This model features the interplay between destabilizing disjoining pressure and stabilizing elastic bending pressure and surface tension. The governing equation is a sixth-order nonlinear degenerate parabolic partial differential equation parameterized by exponents in the mobility function and the disjoining pressure, respectively. Asymptotic self-similar finite-time rupture solutions governed by a sixth-order leading-order equation are analyzed. In the weak elasticity limit, transient self-similar dynamics governed by a fourth-order similarity equation are also identified.

math.AP

Stochastic Couplings and Bijections from the Symmetric Group to Itself

Inspired by the Stochastic processes described by the Feller Coupling and Chinese Restaurant Processes, we create four different bijections from words in the set $[1]\times [2] \times\cdot \times[n]$ to $S_n$. We then compose these maps with their inverse to obtain a toal of six bijections $S_n \to S_n$. Following that, we investigate the fixed points ($1$-cycle) and higher $k$-cycles of these maps. We characterized some of their properties completely as well as empirically showing the complexity of the higher $k$-cycle structures for these maps.

math.CO

Approximation Capabilities of Neural Networks using Morphological Perceptrons and Generalizations

Standard artificial neural networks (ANNs) use sum-product or multiply-accumulate node operations with a memoryless nonlinear activation. These neural networks are known to have universal function approximation capabilities. Previously proposed morphological perceptrons use max-sum, in place of sum-product, node processing and have promising properties for circuit implementations. In this paper we show that these max-sum ANNs do not have universal approximation capabilities. Furthermore, we consider proposed signed-max-sum and max-star-sum generalizations of morphological ANNs and show that these variants also do not have universal approximation capabilities. We contrast these variations to log-number system (LNS) implementations which also avoid multiplications, but do exhibit universal approximation capabilities.

cs.LG