Searcharxiv⌕ Search

arXiv subjects

Atsushi Iwasaki

Publications and source records attributed to Atsushi Iwasaki.

At least 19 recordsLinked to original sources

Asymmetric Perturbation in Solving Bilinear Saddle-Point Optimization

This paper proposes asymmetric perturbation, where only one player's payoff function is perturbed, for solving bilinear saddle-point optimization problems, commonly arising in minimax problems, game theory, and constrained optimization. Symmetric perturbation is known to require decreasing its strength to ensure convergence to a solution, i.e., an equilibrium in the original game, resulting in a slower rate. First, with asymmetric perturbation, we show that, for a sufficiently small perturbation strength, the equilibrium strategy of the asymmetrically perturbed game coincides with an equilibrium strategy of the original unperturbed game. Second, building on this coincidence, we construct a learning algorithm with a linear last-iterate convergence rate. Third, motivated by the fact that the coincidence relies on the perturbation strength being sufficiently small, we also provide a parameter-free variant, retaining the linear rate. Finally, we empirically demonstrate fast convergence toward equilibria in both normal-form and extensive-form games.

math.OC↗

On the Power of Perturbation under Sampling in Solving Extensive-Form Games

We investigate how perturbation does and does not improve the Follow-the-Regularized-Leader (FTRL) algorithm in solving imperfect-information extensive-form games under sampling, where payoffs are estimated from sampled trajectories. While optimistic algorithms are effective under full feedback, they often become unstable in the presence of sampling noise. Payoff perturbation offers a promising alternative for stabilizing learning and achieving \textit{last-iterate convergence}. We present a unified framework for \textit{Perturbed FTRL} algorithms and study two variants: PFTRL-KL (standard KL divergence) and PFTRL-RKL (Reverse KL divergence), the latter featuring an estimator with both unbiasedness and conditional zero variance. While PFTRL-KL generally achieves equivalent or better performance across benchmark games, PFTRL-RKL consistently outperforms it in Leduc poker, whose structure is more asymmetric than the other games in a sense. Given the modest advantage of PFTRL-RKL, we design the second experiment to isolate the effect of conditional zero variance, showing that the variance-reduction property of RKL improve last-iterate performance.

cs.GT↗

Evaluating the Efficiency of Regulation in Matching Markets with Distributional Disparities

Cap-based regulations are widely used to address distributional disparities in matching markets, but their efficiency relative to alternative instruments such as subsidies remains poorly understood. This paper develops a framework for evaluating policy interventions by incorporating regional constraints into a transferable utility matching model. We show that a policymaker with aggregate-level match data can implement a taxation policy that maximizes social welfare and outperforms any cap-based policy. Using newly collected data from the Japan Residency Matching Program, we estimate participant preferences and simulate counterfactual match outcomes under both cap-based and subsidy-based policies. The results reveal that the status quo cap-based regulation generates substantial efficiency losses, whereas small, targeted subsidies can achieve similar distributional goals with significantly higher social welfare.

cs.GT↗

Boosting Perturbed Gradient Ascent for Last-Iterate Convergence in Games

This paper presents a payoff perturbation technique, introducing a strong convexity to players' payoff functions in games. This technique is specifically designed for first-order methods to achieve last-iterate convergence in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. Although perturbation is known to facilitate the convergence of learning algorithms, the magnitude of perturbation requires careful adjustment to ensure last-iterate convergence. Previous studies have proposed a scheme in which the magnitude is determined by the distance from a periodically re-initialized anchoring or reference strategy. Building upon this, we propose Gradient Ascent with Boosting Payoff Perturbation, which incorporates a novel perturbation into the underlying payoff function, maintaining the periodically re-initializing anchoring strategy scheme. This innovation empowers us to provide faster last-iterate convergence rates against the existing payoff perturbed algorithms, even in the presence of additive noise.

cs.GT↗

Approximate State Abstraction for Markov Games

This paper introduces state abstraction for two-player zero-sum Markov games (TZMGs), where the payoffs for the two players are determined by the state representing the environment and their respective actions, with state transitions following Markov decision processes. For example, in games like soccer, the value of actions changes according to the state of play, and thus such games should be described as Markov games. In TZMGs, as the number of states increases, computing equilibria becomes more difficult. Therefore, we consider state abstraction, which reduces the number of states by treating multiple different states as a single state. There is a substantial body of research on finding optimal policies for Markov decision processes using state abstraction. However, in the multi-player setting, the game with state abstraction may yield different equilibrium solutions from those of the ground game. To evaluate the equilibrium solutions of the game with state abstraction, we derived bounds on the duality gap, which represents the distance from the equilibrium solutions of the ground game. Finally, we demonstrate our state abstraction with Markov Soccer, compute equilibrium policies, and examine the results.

cs.GT↗

Adaptively Perturbed Mirror Descent for Learning in Games

This paper proposes a payoff perturbation technique for the Mirror Descent (MD) algorithm in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. The optimistic family of learning algorithms, exemplified by optimistic MD, successfully achieves {\it last-iterate} convergence in scenarios devoid of noise, leading the dynamics to a Nash equilibrium. A recent re-emerging trend underscores the promise of the perturbation approach, where payoff functions are perturbed based on the distance from an anchoring, or {\it slingshot}, strategy. In response, we propose {\it Adaptively Perturbed MD} (APMD), which adjusts the magnitude of the perturbation by repeatedly updating the slingshot strategy at a predefined interval. This innovation empowers us to find a Nash equilibrium of the underlying game with guaranteed rates. Empirical demonstrations affirm that our algorithm exhibits significantly accelerated convergence.

cs.GT↗

Learning Fair Division from Bandit Feedback

This work addresses learning online fair division under uncertainty, where a central planner sequentially allocates items without precise knowledge of agents' values or utilities. Departing from conventional online algorithm, the planner here relies on noisy, estimated values obtained after allocating items. We introduce wrapper algorithms utilizing \textit{dual averaging}, enabling gradual learning of both the type distribution of arriving items and agents' values through bandit feedback. This approach enables the algorithms to asymptotically achieve optimal Nash social welfare in linear Fisher markets with agents having additive utilities. We establish regret bounds in Nash social welfare and empirically validate the superior performance of our proposed algorithms across synthetic and empirical datasets.

cs.LG↗

Last-Iterate Convergence with Full and Noisy Feedback in Two-Player Zero-Sum Games

This paper proposes Mutation-Driven Multiplicative Weights Update (M2WU) for learning an equilibrium in two-player zero-sum normal-form games and proves that it exhibits the last-iterate convergence property in both full and noisy feedback settings. In the former, players observe their exact gradient vectors of the utility functions. In the latter, they only observe the noisy gradient vectors. Even the celebrated Multiplicative Weights Update (MWU) and Optimistic MWU (OMWU) algorithms may not converge to a Nash equilibrium with noisy feedback. On the contrary, M2WU exhibits the last-iterate convergence to a stationary point near a Nash equilibrium in both feedback settings. We then prove that it converges to an exact Nash equilibrium by iteratively adapting the mutation term. We empirically confirm that M2WU outperforms MWU and OMWU in exploitability and convergence rates.

cs.GT↗

Mutation-Driven Follow the Regularized Leader for Last-Iterate Convergence in Zero-Sum Games

In this study, we consider a variant of the Follow the Regularized Leader (FTRL) dynamics in two-player zero-sum games. FTRL is guaranteed to converge to a Nash equilibrium when time-averaging the strategies, while a lot of variants suffer from the issue of limit cycling behavior, i.e., lack the last-iterate convergence guarantee. To this end, we propose mutant FTRL (M-FTRL), an algorithm that introduces mutation for the perturbation of action probabilities. We then investigate the continuous-time dynamics of M-FTRL and provide the strong convergence guarantees toward stationary points that approximate Nash equilibria under full-information feedback. Furthermore, our simulation demonstrates that M-FTRL can enjoy faster convergence rates than FTRL and optimistic FTRL under full-information feedback and surprisingly exhibits clear convergence under bandit feedback.

cs.GT↗

Anytime Capacity Expansion in Medical Residency Match by Monte Carlo Tree Search

This paper considers the capacity expansion problem in two-sided matchings, where the policymaker is allowed to allocate some extra seats as well as the standard seats. In medical residency match, each hospital accepts a limited number of doctors. Such capacity constraints are typically given in advance. However, such exogenous constraints can compromise the welfare of the doctors; some popular hospitals inevitably dismiss some of their favorite doctors. Meanwhile, it is often the case that the hospitals are also benefited to accept a few extra doctors. To tackle the problem, we propose an anytime method that the upper confidence tree searches the space of capacity expansions, each of which has a resident-optimal stable assignment that the deferred acceptance method finds. Constructing a good search tree representation significantly boosts the performance of the proposed method. Our simulation shows that the proposed method identifies an almost optimal capacity expansion with a significantly smaller computational budget than exact methods based on mixed-integer programming.

cs.GT↗

The reference distributions of Maurer's universal statistical test and its improved tests

Maurer's universal statistical test can widely detect non-randomness of given sequences. Coron proposed an improved test, and further Yamamoto and Liu proposed a new test based on Coron's test. These tests use normal distributions as their reference distributions, but the soundness has not been theoretically discussed so far. Additionally, Yamamoto and Liu's test uses an experimental value as the variance of its reference distribution. In this paper, we theoretically derive the variance of the reference distribution of Yamamoto and Liu's test and prove that the true reference distribution of Coron's test converges to a normal distribution in some sense. We can apply the proof to the other tests with small changes.

math.ST↗

Repeated Multimarket Contact with Private Monitoring: A Belief-Free Approach

This paper studies repeated games where two players play multiple duopolistic games simultaneously (multimarket contact). A key assumption is that each player receives a noisy and private signal about the other's actions (private monitoring or observation errors). There has been no game-theoretic support that multimarket contact facilitates collusion or not, in the sense that more collusive equilibria in terms of per-market profits exist than those under a benchmark case of one market. An equilibrium candidate under the benchmark case is belief-free strategies. We are the first to construct a non-trivial class of strategies that exhibits the effect of multimarket contact from the perspectives of simplicity and mild punishment. Strategies must be simple because firms in a cartel must coordinate each other with no communication. Punishment must be mild to an extent that it does not hurt even the minimum required profits in the cartel. We thus focus on two-state automaton strategies such that the players are cooperative in at least one market even when he or she punishes a traitor. Furthermore, we identify an additional condition (partial indifference), under which the collusive equilibrium yields the optimal payoff.

cs.GT↗

Near-Feasible Stable Matchings with Budget Constraints

We consider the matching with contracts framework of Hatfield and Milgrom when one side (a firm or hospital) can make monetary transfers (offer wages) to the other (a worker or doctor). In a standard model, monetary transfers are not restricted. However, we assume that each hospital has a fixed budget; that is, the total amount of wages allocated by each hospital to the doctors is constrained. With this constraint, stable matchings may fail to exist and checking for the existence is hard. To deal with the nonexistence, we focus on near-feasible matchings that can exceed each hospital budget by a certain amount, and We introduce a new concept of compatibility. We show that the compatibility condition is a sufficient condition for the existence of a near-feasible stable matching in the matching with contracts framework. Under a slight restriction on hospitals' preferences, we provide mechanisms that efficiently return a near-feasible stable matching with respect to the actual amount of wages allocated by each hospital. By sacrificing strategy-proofness, the best possible bound of budget excess is achieved.

cs.GT↗

Independent Randomness Tests based on the Orthogonalized Non-overlapping Template Matching Test

In general, randomness tests included in a test suite are not independent of each other. This renders it difficult to fix a rational criterion through the whole test suite with an explicit significance level. In this paper, we focus on the Non-overlapping Template Matching Test, which is a randomness test included in the NIST statistical test suite. The test uses a parameter called "template" and we can consider a test item for each template. We investigate dependency between two test items by deriving the joint probability density function of the two p-values and propose a transformation to make multi test items independent of each other.

math.ST↗

Approximately Stable Matchings with General Constraints

This paper focuses on two-sided matching where one side (a hospital or firm) is matched to the other side (a doctor or worker) so as to maximize a cardinal objective under general feasibility constraints. In a standard model, even though multiple doctors can be matched to a single hospital, a hospital has a responsive preference and a maximum quota. However, in practical applications, a hospital has some complicated cardinal preference and constraints. With such preferences (e.g., submodular) and constraints (e.g., knapsack or matroid intersection), stable matchings may fail to exist. This paper first determines the complexity of checking and computing stable matchings based on preference class and constraint class. Second, we establish a framework to analyze this problem on packing problems and the framework enables us to access the wealth of online packing algorithms so that we construct approximately stable algorithms as a variant of generalized deferred acceptance algorithm. We further provide some inapproximability results.

cs.GT↗

Deriving the Variance of the Discrete Fourier Transform Test Using Parseval's Theorem

The discrete Fourier transform test is a randomness test included in NIST SP800-22. However, the variance of the test statistic is smaller than expected and the theoretical value of the variance is not known. Hitherto, the mechanism explaining why the former variance is smaller than expected has been qualitatively explained based on Parseval's theorem. In this paper, we explore this quantitatively and derive the variance using Parseval's theorem under particular assumptions. Numerical experiments are then used to show that this derived variance is robust.

math.ST↗

Approximately Stable Matchings with Budget Constraints

This paper considers two-sided matching with budget constraints where one side (firm or hospital) can make monetary transfers (offer wages) to the other (worker or doctor). In a standard model, while multiple doctors can be matched to a single hospital, a hospital has a maximum quota: the number of doctors assigned to a hospital cannot exceed a certain limit. In our model, a hospital instead has a fixed budget: the total amount of wages allocated by each hospital to doctors is constrained. With budget constraints, stable matchings may fail to exist and checking for the existence is hard. To deal with the nonexistence of stable matchings, we extend the "matching with contracts" model of Hatfield and Milgrom, so that it handles approximately stable matchings where each of the hospitals' utilities after deviation can increase by factor up to a certain amount. We then propose two novel mechanisms that efficiently return such a stable matching that exactly satisfies the budget constraints. In particular, by sacrificing strategy-proofness, our first mechanism achieves the best possible bound. Furthermore, we find a special case such that a simple mechanism is strategy-proof for doctors, keeping the best possible bound of the general case.

cs.GT↗

Analysis of NIST SP800-22 focusing on randomness of each sequence

NIST SP800-22 is a randomness test set applied for a set of sequences. Although SP800-22 widely used, a rational criterion throughout all test items has not been shown. The main reason is that the dependency of test items has not been perfectly clear. In this paper, a certain scalar is computed for each sequence throughout all test items and make the histogram of the scalar. By comparing the histogram and the theoretical distribution under some assumptions, the dependency is visually shown. In addition, an algorithmic method to derive "minimum set" using the histogram is proposed.

stat.AP↗