Heuristics in experiments with infinitely large strategy spaces
We introduce a new methodology that enables detection of the onset of convergence towards Nash equilibria in simple repeated games with infinitely large strategy spaces, thereby revealing the heuristics used in decision-making. The method works by constraining on a special finite subset of strategies, called decoupled strategies. We show how the technique can be applied to understand price formation in financial market experiments by introducing a predictive measure {\Delta}D: the different between positive decoupled strategies (recommending to buy) and negative decoupled strategies (recommending to sell). Using {\Delta}D we illustrate how the method can predict (at certain special times) participants' actions with a high success rate in a series of experiments
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingSimilar Papers 제목 키워드 기반
Efficient Strategy Synthesis for MDPs with Resource Constraints
We consider qualitative strategy synthesis for the formalism called consumption Markov decision processes. This formalism can model dynamics of an agents that operates under resource constraints in a stochastic environme…
Unbiased Methods for Multi-Goal Reinforcement Learning
In multi-goal reinforcement learning (RL) settings, the reward for each goal is sparse, and located in a small neighborhood of the goal. In large dimension, the probability of reaching a reward vanishes and the agent rec…
Multi-Goal Reinforcement LearningQ-Learningreinforcement-learningReinforcement Learning+1Offline Reinforcement Learning for Large Scale Language Action Spaces
Training a task-oriented dialogue agent can be naturally formulated as offline reinforcement learning (RL) problem, where the agent aims to learn a conversational strategy to achieve user goals, only from a dialogue corp…
Language ModelingLanguage ModellingOffline RLreinforcement-learning+2Towards Less Constrained Macro-Neural Architecture Search
Networks found with Neural Architecture Search (NAS) achieve state-of-the-art performance in a variety of tasks, out-performing human-designed networks. However, most NAS methods heavily rely on human-defined assumptions…
GPUNeural Architecture SearchCommunication in the Infinitely Repeated Prisoner's Dilemma: Theory and Experiments
So far, the theory of equilibrium selection in the infinitely repeated prisoner's dilemma is insensitive to communication possibilities. To address this issue, we incorporate the assumption that communication reduces -- …