Computer poker opponents can now play at the level of the best research programs, and a designer building one has to decide how much of that strength a human should face. In most card games on computers, the opponent across the table followed rules that a designer wrote by hand. Between 2015 and 2019, 4 research programs reached and then passed the level of elite professionals, and the last of them learned its strategy in about a week on one server.Scripted OpponentsA scripted opponent works from fixed rules. It raises with hands above one threshold, calls with hands above a lower one, folds the rest, and bluffs at a rate the designer chose in advance. Difficulty settings usually move those thresholds up or down.The weakness is that the rules never change. A human who plays a few hundred hands against a scripted opponent can learn when it bluffs and when it folds, and from then on the opponent is a fixed puzzle with a known solution. Nothing in the script lets the program notice that it is being exploited. Designers have usually hidden this with randomness, adding a chance that the program ignores its own rule on a given hand, which makes it harder to read without making it any stronger.Counterfactual Regret MinimizationThe research programs of the last decade replaced hand-written rules with self-play. The central method, counterfactual regret minimization, has the program play against itself over an enormous number of hands. After each decision it measures how much better every alternative action would have done, which the researchers call regret, and it gradually favors the actions with the most accumulated regret. In a 2-player game the average strategy moves toward an equilibrium that no opponent can profitably exploit.Bluffing comes out of this process without anyone programming it. Michael Bowling of the University of Alberta, whose group built 2 of the programs described below, has said that bluffing follows from the rules of the game itself, since a strategy that never bluffs would be easy to read.The Decision to Withhold PluribusWhen Carnegie Mellon University and Facebook AI published the 6-player program Pluribus in 2019, they declined to release its source code and included pseudocode for the major components instead. Noam Brown, one of its developers, said a release "could be very dangerous for the poker community." The concern was public money games on the internet, where a program can take a seat without the other players knowing it is software, and where a strategy strong enough to beat professionals would take money from everyone else at the table.A commercial game opponent never raises that problem, because the player already knows the opponent is a machine. Home video games can use the research freely, and the limit on those opponents is the strength the designer chooses.Solving Heads-Up Limit Hold'emHeads-up limit hold'em was the simplest version of the game played for money in casinos and in online poker at the time, and it became the first target. The milestone came in January 2015, when Michael Bowling, Neil Burch, Michael Johanson and Oskari Tammelin reported in Science that heads-up limit hold'em, a game with more than 100 trillion distinct decision points, had been essentially solved. The team had about 4,000 processors training the University of Alberta's program, Cepheus, at the same time, and the resulting strategy is so close to optimal that the best possible counter-strategy wins only 0.000986 big blinds per game against it. An earlier program from the same field, Polaris, had beaten 6 human professionals in 2008 without solving anything.No-Limit Against ProfessionalsNo-limit betting multiplies the number of possible situations, and the programs that handled it arrived almost together. DeepStack, built by the Alberta group with colleagues at the Czech Technical University in Prague, played 33 professionals from 17 countries in late 2016, with cash prizes of up to $5,000 for the top 3 humans. In March 2017 Science published the results of a system that beat professional players by looking only a few actions ahead and relying on a trained network to judge the positions beyond.Carnegie Mellon's Libratus took the other route, with heavy computation during play. From January 11 to 30, 2017, at Rivers Casino in Pittsburgh, it beat its 4 human rivals, Jason Les, Dong Kim, Daniel McAulay and Jimmy Chou, over 120,000 hands. The final margin was $1,766,250 in chips, which were not real money. During the match the program ran on roughly 600 of the 846 compute nodes of the Bridges supercomputer at the Pittsburgh Supercomputing Center, and it used the quiet hours to study the bet sizes the professionals had tried against it.A 6-Player TablePluribus extended the approach from 2 players to 6, where no equilibrium guarantee exists. It won against 5 professionals at a time over 10,000 hands, and separately against Darren Elias and Chris Ferguson, each of whom played 5,000 hands against 5 copies of the program. Reporting at the time stressed its relentless consistency and how modest its needs were. It was trained on an ordinary 64-core server over about a week, at a cost of roughly $144 at cloud computing rates, and could play on a machine with 2 CPUs and 128 gigabytes of memory. In 2019, 2 years after Libratus needed a supercomputer, a stronger multiplayer program ran on a single workstation.Pluribus treats bets of similar size as the same bet, so a bet of anything between roughly $70 and $130 counts as a $100 bet, which keeps the calculation small. It also varies its own bet with the same hand in the same spot, so a human cannot learn a single pattern to exploit.Designer Choices After PluribusPluribus played a fixed strategy that ignored the observed tendencies of the people across the table. Its authors explained the choice by the cost of exploitation, since a program that bends its play toward one opponent's weaknesses opens itself to exploitation in turn, and the known techniques needed too many hands to pay off against humans. A home game built on that method would beat a beginner slowly and offer that beginner no pattern to study, because the program's play would not respond to anything the beginner did.The practical opponent for a video game is therefore a strong base strategy with deliberate, readable weaknesses layered on top, tuned to a difficulty setting. The research made the strong part cheap, and the weak parts are now a design decision.A Question for Game DesignersThat leaves designers with a question the research teams never had to answer. Would players rather face a machine that plays close to perfectly and never explains itself, or one built to show them, hand by hand, which of their habits it chose to punish?Poker AI Opponents FAQHow do modern AI poker opponents work?Modern poker AI programs use counterfactual regret minimization to play millions of self-play hands, arriving at game theory optimal equilibrium strategies without human rules.What was the significance of the Pluribus poker AI?Pluribus defeated elite human professionals in six-player no-limit hold’em using modest computational resources, proving multi-player poker could be solved without supercomputers.Why were some poker AI programs withheld from the public?Developers withheld source code for elite AI to prevent automated software from taking seats in public online poker money games undetected.How does artificial intelligence impact card game strategy?Artificial intelligence evaluates massive decision trees and optimal betting ranges to fundamentally transform modern strategy. Learn more about computer poker players.
Reacciones
0
0 Comentarios
Se el primero en comentar
Nada aquí todavía.