SearcharxivSearch

arXiv · 2105.10099

Learning from zero: how to make consumption-saving decisions in a stochastic environment with an AI algorithm

Abstract

This exercise proposes a learning mechanism to model economic agent's decision-making process using an actor-critic structure in the literature of artificial intelligence. It is motivated by the psychology literature of learning through reinforcing good or bad decisions. In a model of an environment, to learn to make decisions, this AI agent needs to interact with its environment and make explorative actions. Each action in a given state brings a reward signal to the agent. These interactive experience is saved in the agent's memory, which is then used to update its subjective belief of the world. The agent's decision-making strategy is formed and adjusted based on this evolving subjective belief. This agent does not only take an action that it knows would bring a high reward, it also explores other possibilities. This is the process of taking explorative actions, and it ensures that the agent notices changes in its environment and adapt its subjective belief and decisions accordingly. Through a model of stochastic optimal growth, I illustrate that the economic agent under this proposed learning structure is adaptive to changes in an underlying stochastic process of the economy. AI agents can differ in their levels of exploration, which leads to different experience in the same environment. This reflects on to their different learning behaviours and welfare obtained. The chosen economic structure possesses the fundamental decision making problems of macroeconomic models, i.e., how to make consumption-saving decisions in a lifetime, and it can be generalised to other decision-making processes and economic models.

Explore related subjects

Keep this discovery

BibTeXRIS

Rui, Shi. 2021-05-21. Learning from zero: how to make consumption-saving decisions in a stochastic environment with an AI algorithm. https://arxiv.org/abs/2105.10099

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Social Preferences and Cooperation: Beliefs, Robustness, and the Limits of Altruism

We study a mechanism of cooperation in the Prisoner's Dilemma (PD). Incorporating social preferences as efficiency concerns into the PD game, we study how altruism translates into cooperation. Under complete information, cooperation requires the opponent's altruism to clear a threshold. We then introduce a subjective extension of Bayesian Nash equilibrium that relaxes the Common Prior Assumption, letting players hold heterogeneous, potentially misspecified beliefs about each other's altruistic type. Cooperation then depends on beliefs about altruism rather than altruism itself, and can be sustained even when opponents are, on average, only weakly altruistic. When fear of exploitation dominates the temptation to defect, beliefs about the opponent's cooperation become strategic complements, so a cooperative and an uncooperative equilibrium can coexist under identical payoffs and an identical, correctly specified prior. Using multiplier preferences, we then study how robust this belief-driven cooperation is to model misspecification. Cooperation is fragile: it survives only above a threshold level of confidence in one's own belief, and can unravel even when the belief itself correctly supports cooperation. As a formal extension, the same robust-control apparatus, applied to a player's action choice, nests Nash equilibrium, Bayesian Nash equilibrium, and logit Quantal Response Equilibrium as limiting cases. Cooperation depends less on how altruistic agents are than on what they believe about each other, and how confident they are that this belief is right.

econ.TH

Utility-Level-Dependent Ambiguity

Experimental evidence suggests that ambiguity-sensitive choice can vary systematically with the circumstances of a decision. This paper isolates one channel within a stable preference relation: ambiguity weighting may depend on the act's certainty-equivalent level. After the standard Anscombe-Aumann calibration of consequence utility, a set of behavioral axioms yields a unique continuous family of normalized monotone capacities $\{\nu_v\}_{v\in(0,1)}$. Each nonendpoint act is evaluated by the Choquet integral associated with the capacity at its own interior certainty-equivalent level, while nonendpoint acts on the same indifference surface share the same capacity. Binary event comparisons identify local event weights at each elicited level and trace their cross-level variation, providing tests of the fixed-capacity restriction. Local uncertainty aversion is equivalent to convexity of $\nu_v$ and yields an implicit multiple-priors representation with certainty-equivalent-indexed local cores. Certainty translation invariance holds if and only if the capacity is fixed across levels, recovering the maintained nondegenerate fixed-capacity Choquet expected utility benchmark; global mixture-betweenness yields implicit additive utility, and imposing both restrictions recovers full-support subjective expected utility. The capacity schedule is a reduced-form ambiguity weighting whose variation may reflect changes in ambiguity perception, ambiguity attitude, or both.

econ.TH

The Attention Cost of Stable Matching

In large markets, scarce attention limits partner evaluation and creates allocation loss, which stability magnifies. In an independent random market with average executable degree $d$, unmatched shares fall at rates $e^{-\sqrt d}$ under stability and $e^{-d}$ under maximum matching on the same graph. Changing consideration can make applications rejected in a provisional active-screen computation relevant again. Exact query-neutral implementation must retain allocation-relevant off-screen authorizations; otherwise, missing authorization must be reacquired. Limited-attention deferred acceptance (LA-DA) preserves valid authorizations and reengages eligible pairs. Conditional on exact next-best information and persistent execution rights, adaptive discovery saves a logarithmic factor in reached proposals relative to independent exposure. In an application to speed dating, bilateral reports let us compare stable and maximum matching on restricted graphs, separating missed opportunities from same-graph stability loss. In Chilean school choice, we document 9,502 applicants accepting higher-ranked or new placements through retained rankings.

econ.TH