Skip to content

Open Research Questions

  • Investigate an objective-aware Nash-Q equilibrium selector that prefers stage-game equilibria compatible with every agent's temporal objective when one exists. The current payoff_dominant selector ranks learned Q payoffs; it does not inspect monitor acceptance or return an infeasibility result.