Source-linked AI summary
The AI Economist: Improving Equality and Productivity with AI-Driven Tax Policies
Stephan Zheng, Alexander Trott, Sunil Srinivasa, Nikhil Naik, Melvin Gruesbeck, David C. Parkes, Richard Socher
TL;DR
The paper tackles the difficulty of designing and testing tax policies that balance equality with productivity. It trains adaptive social planners and economic agents through two-level deep reinforcement learning in economic simulations, finding improved trade-offs over the Saez framework and effective transfer to human-participant experiments.
Problem
Designing tax policies that balance equality and productivity is difficult because taxation can improve redistribution while weakening incentives to work.
Method
The AI Economist uses two-level deep reinforcement learning in economic simulations where agents and a social planner learn and adapt from observable data.
Results
16%: AI-driven tax policies improve the equality-productivity trade-off compared with the prominent Saez tax framework.
Takeaways & Limitations
The learned policy performs well despite strategic tax avoidance and remains effective with human participants without additional recalibration.
Takeaways & Limitations
The simulations cover a relatively small economy and omit human-behavioral factors, other-regarding utilities, and richer representations of real-world skill and pay.
Abstract
from arXiv · showhide
Tackling real-world socio-economic challenges requires designing and testing economic policies. However, this is hard in practice, due to a lack of appropriate (micro-level) economic data and limited opportunity to experiment. In this work, we train social planners that discover tax policies in dynamic economies that can effectively trade-off economic equality and productivity. We propose a two-level deep reinforcement learning approach to learn dynamic tax policies, based on economic simulations in which both agents and a government learn and adapt. Our data-driven approach does not make use of economic modeling assumptions, and learns from observational data alone. We make four main contributions. First, we present an economic simulation environment that features competitive pressures and market dynamics. We validate the simulation by showing that baseline tax systems perform in a way that is consistent with economic theory, including in regard to learned agent behaviors and specializations. Second, we show that AI-driven tax policies improve the trade-off between equality and productivity by 16% over baseline policies, including the prominent Saez tax framework. Third, we showcase several emergent features: AI-driven tax policies are qualitatively different from baselines, setting a higher top tax rate and higher net subsidies for low incomes. Moreover, AI-driven tax policies perform strongly in the face of emergent tax-gaming strategies learned by AI agents. Lastly, AI-driven tax policies are also effective when used in experiments with human participants. In experiments conducted on MTurk, an AI tax policy provides an equality-productivity trade-off that is similar to that provided by the Saez framework along with higher inverse-income weighted social welfare.
1 Introduction
The paper addresses the unresolved challenge of balancing equality and productivity by introducing an adaptive, simulation-based AI system that learns tax policies alongside agent behavior. Its AI-driven policies improve this trade-off over the Saez framework and remain effective amid strategic tax avoidance and human participation.
- Motivation: The paper frames optimal taxation as a difficult trade-off because redistribution can reduce labor incentives and productivity.Tax policy is important for reducing inequality, but higher taxation may discourage work, especially among more productive workers.
- Approach: The AI Economist uses two-level deep reinforcement learning to train adaptive social planners and economic agents in simulation.The planner learns tax policies while agents learn behaviors in response to the active policy, using observable data without prior economic assumptions.
- Approach: The simulation framework enables large-scale testing of economic policies with learned agent behavior and multiple social-outcome metrics.It is designed to model competitive pressures, trade, and resource scarcity while supporting comparisons across many economic designs.
- Results: 16%: AI-driven tax policies improve the equality-productivity trade-off compared with the prominent Saez tax framework.The learned policies use tax-rate schedules that differ from baseline policies.
- Results: AI agents learn tax-avoidance behaviors, yet the AI Economist’s tax schedule performs well despite this strategic behavior.Agents can modulate their incomes across tax periods in response to taxation.
- Results: The learned policy remains effective with human participants without additional recalibration, achieving a competitive equality-productivity trade-off and higher inverse-income weighted social welfare.The reported human-participant experiments provide preliminary evidence for applicability beyond the simulated setting.
2 Economic Simulations: Learning in Gather-and-Build Games
The Gather-and-Build simulation models adaptive agents that gather, build, trade, and learn individual policies under heterogeneous skills, resources, and labor costs. Learned behavior produces specialization and unequal incomes without imposing roles directly.
- Economic simulation framework: The framework studies economic design through simulations in which AI agents learn behaviors and social outcomes in dynamic environments.The section focuses on a no-tax setting to illustrate outcomes that taxation may address.
- Learning agent behavior: Agents share policy parameters during training but retain distinct behaviors because their observations and hidden states differ.The policy uses agent-specific observations and hidden states while sharing weights for learning efficiency.
- Environment rules and dynamics: Agents operate in a two-dimensional world where they collect wood and stone, build houses for coins, and trade resources.Resources regenerate stochastically, and harvested tiles remain empty until new resources spawn.
- Environment rules and dynamics: Agent heterogeneity arises from differing building and collection skills, initial locations, income opportunities, and labor costs for movement, gathering, trading, and building.Building skill affects house income, while collection skill affects bonus resources from harvesting.
- Emergent specialization: Learned agents develop a division of labor in which some gather resources, another builds houses, and another changes strategy over time.Low-skilled agents shift away from building and earn income by selling resources to higher-skilled builders.
- Emergent specialization: Specialization emerges from agents optimizing individual objectives and balancing income against effort, while free-market incomes remain highly unequal.The emergent roles are consistent with economic intuition that agents specialize in more efficient means of converting labor into income.
3 Machine Learning for Optimal Tax Policies
The paper frames optimal taxation as a trade-off between equality and productivity, then uses bracketed income taxes, social welfare objectives, and two-level reinforcement learning to optimize policies in a dynamic economy.
- Taxation and social objectives: Taxes can redistribute wealth and improve equality, but may reduce productivity by discouraging labor, especially among highly skilled workers.The resulting equality-productivity trade-off makes optimal taxation a constrained optimization problem.
- Emergent specialization: In the no-tax rollout, lower-skilled agents gather and sell resources while higher-skilled agents build houses and achieve greater utility.The figure reports payoffs of 11.3, 13.3, 16.5, and 22.2 coins per house across agents sorted by building skill.
- Tax policy representation: The AI Economist learns periodic income-tax schedules with fixed brackets, marginal rates, and lump-sum redistribution to agents.The planner chooses rates for each bracket, while collected revenue is evenly redistributed at the end of each tax period.
- Taxation and social objectives: Social welfare is primarily defined as the product of equality and productivity, while the framework also supports alternative utility-weighted objectives.Equality is based on the complement of the normalized Gini index of post-tax cumulative wealth, and inverse-income weighting emphasizes agents with lower endowments.
- Two-level reinforcement learning: The two-level reinforcement-learning framework trains economic agents in an inner loop and a social planner in an outer loop.Agents learn labor and trading behaviors under taxes, while the planner adapts the tax policy to optimize its chosen social objective.
4 Improved Social Outcomes with AI Agents
The AI Economist achieves the strongest equality–productivity outcomes among the evaluated tax models, while producing distinct redistribution patterns and remaining effective under strategic tax gaming. Its learned policy also reflects behavioral differences in how agents specialize, trade, and respond to taxes.
- Baseline methods: The evaluated comparison includes free-market, US federal, Saez, and AI Economist tax models using common bracketed schedules.The Saez framework is a prominent analytical baseline, but applying it in practice requires estimating income-tax elasticity, which is highly non-trivial.
- Overall economic outcomes: 16% improvement in the equality-productivity trade-off over the next-best Saez model, with the smallest productivity loss among taxed treatments.The AI Economist also improves equality by 47% relative to the free market at an 11% productivity decrease.
- Overall economic outcomes: The simulations reproduce an equality–productivity trade-off: redistribution raises equality but reduces productivity, with the severity depending on the tax schedule.The framework uses learned agent responses to taxes to expose this trade-off.
- Tax schedules and redistribution: The AI Economist combines progressive and regressive features, including a higher top tax rate and higher subsidies for low-income agents than baseline policies.Lower-skilled agents receive higher net income under the AI Economist than under the other models.
- Tax schedules and redistribution: The AI Economist’s stronger outcomes arise partly from behavioral effects: its policy changes specialization, resource collection, building, and trading relative to Saez taxation.Under Saez taxes, altered specialization and reduced building weaken productivity, while the AI Economist produces stronger redistribution through trading behavior.
- Strategic behavior: AI agents learn tax-avoidance behavior by alternating high and low incomes across periods, yet the AI Economist remains effective despite this strategic response.This behavior appears under both the Saez and AI Economist models, which have lower top tax rates.
5 Improved Social Outcomes with Human Participants
The AI Economist tax policy transferred to human-participant experiments with minimal calibration and achieved competitive equality-productivity outcomes. It also produced higher inverse-income-weighted social welfare, despite behavioral differences and experimental modifications.
- Transfer: The AI Economist tax model transferred from the AI-only setting without extensive recalibration or fine-tuning, requiring only income-bracket scaling by a factor of three.
- Experimental setting: Human experiments disabled trading, increased building costs, used five-minute episodes, and extended episodes to 3000 timesteps.
- Experimental setting: The human-participant experiments used US-based MTurk participants in groups of four, with four-episode HITs and tutorials before each episode.
- Results: The Camelback tax schedule achieved an equality-productivity trade-off comparable to Saez and better than US Federal and free-market approaches.
- Discussion: Human behavior differed substantially from AI behavior, including more adversarial blocking behavior and lower utility in some trials.
- Discussion: The authors do not endorse the learned schedule for direct use in the real economy, while identifying potential for future AI-driven tax models.
6 Conclusion
The conclusion presents AI-based economic simulation as a promising way to study and transfer economic policies, while emphasizing that the demonstrated setting remains limited. It identifies both transfer potential and substantial fidelity constraints.
- The approach aims to study economic-policy impacts at a complexity that traditional economics research cannot easily address.
- Model-free reinforcement learning allows flexible planner rewards and does not require prior world knowledge to find a well-performing tax policy.
- AI-based economic simulators produced agent behaviors consistent with economic intuition and policies that transferred to human-participant settings within a limited problem setting.
- The simulations do not yet model many human behavioral factors, interactions between people, or the scale and complexity of real-world skill and pay relationships.
- Future work should use real-world economic data and larger-scale reinforcement learning to improve simulation fidelity and scope.
7 Ethics and Normative Aspects
The ethics discussion emphasizes that the current AI Economist is only a limited representation of reality and could create risks if future systems are manipulated or trained on biased data. The authors pair these concerns with transparency and review measures.
- The current AI Economist provides only a limited representation of the real world, creating potential for manipulation that could increase inequality while obscuring the action behind an AI system.
- Biased or incomplete training data could produce biased AI-driven tax recommendations, especially when communities or workforce segments are under-represented.
- The simulation is not currently an actual tool for maliciously reconfiguring tax policy.
- The authors recommend model cards and data sheets to document ethical considerations and improve transparency and trust.
- The research was supported by expert consultation, publication for broad debate, and a timed open-source release of the environment and sample training code.
A Details of Environment
The environment models economic activity on a two-dimensional grid with resource gathering, building, trading, and periodic bracketed taxation. Agents and the planner receive different observations, while the planner changes tax rates only at tax-period boundaries.
- World dynamics: The Gather-and-Build environment uses a 2D grid containing agents, resources, houses, water, and fixed resource-source cells.
- Observations: The planner observes the full world state, while agents receive narrower egocentric spatial observations and information about their own inventories and skills.
- Agent mechanics: Agents navigate, gather resources, build houses, and trade through a continuous double auction using bids and asks.
- Market observations: Market observations include outstanding bids and asks, cumulative activity from other agents, average trading prices, and trade counts by price level.
- Taxation and redistribution: At the start of each tax period, the planner selects marginal rates for seven income brackets from discretized action subspaces.
- Taxation and redistribution: Planner observations include current tax rates, current-period marginal rates, temporal progress, and previous-period incomes and tax rates.
- Planner control: Action masks restrict the planner to NO-OP actions outside tax-period starts, allowing rich temporal observation while limiting tax actions to period boundaries.
B Training Hyperparameters and Experiment Settings
Training used parallel environment replicas to collect on-policy experience, while separate passages identify the training hyperparameters and environment settings used in the AI experiments.
- Training setup: Each training iteration collected 200 timesteps from each of 60 environment replicas, totaling 12000 sampled timesteps.Experience collection was parallelized over 60 replicas using the latest policy parameters.
- Experiment settings: Tables 2 and 3 report the training hyperparameters and environment settings used in the AI experiments.
C Details on Experiments with Human Participants
Human-participant experiments used a lobby and tutorial, an in-game interface, and a post-experiment survey, alongside the paper’s inner-outer reinforcement-learning framework for jointly learning agents and a social planner.
- Experiment modules: Human-participant experiments were organized into a lobby and tutorial, main graphical interface, and post-experiment survey.These modules are shown and described in Figures 19, 20, and 21.
- Learning procedure: The inner-outer learning algorithm trains economic agents and the social planner simultaneously using separate policy weights and transition buffers.The algorithm specifies a sampling horizon, tax-period length, on-policy learning algorithm, and stopping criterion.
- Learning procedure: The agent policy samples actions each timestep, while the planner samples marginal tax rates at the start of each tax period and updates taxes at its end.The environment computes state transitions, pre-tax rewards, planner rewards, taxes, and post-tax rewards within this loop.
- Participant interface: The game interface displayed agent endowment, episode time, bonus, world state, tax information, construction progress, and controls.After an episode, participants returned to the lobby until four episodes had been seen, then proceeded to the survey.
- Participant feedback: The post-experiment survey recorded collected resources, coin, and bonus, and gathered experience, strategy, feedback, and technical-issue reports.Participants received a confirmation code to verify successful completion on Amazon Mechanical Turk.