
The FIFA World Cup is always a special event. Beyond just watching the matches, you naturally want to predict the outcomes and debate who will lift the trophy. Thankfully, there’s no shortage of people to talk to. Absolutely everyone starts tuning in to the World Cup, even those who wouldn't normally watch a single full football match at any other time. They might not know the team rosters all that well, but they'll spot a couple of squads with familiar names and pick their ultimate favorite among them. As for me, I watch a ton of football anyway. So, naturally, I also love to theorize and try to guess an outcome every now and then.
About three days before the tournament kicked off, a curious idea struck me. Why not run a betting contest where, alongside myself, several AI models would compete? They're in the spotlight everywhere right now — so why not bring them into this? Plus, it’s a fascinating experiment, a sort of real-time benchmark. Training data wouldn't help them much here. Think of it like that experiment where AI models traded crypto.
With almost no time left for any real preparation, I didn’t delay. I picked a few LLMs that were available at the time to my liking and asked an AI agent to code up a handy toolkit for me. A couple of iterations later, everything was ready. I drew up a set of rules for the experiment, and we were off!
The Rules
To keep things orderly, here is the short list of conditions and constraints:
- We bet with play money, obviously. No real cash involved.
- 104 matches. A bet must be placed on every single one.
- Fixed bet size: 100 conventional units per event. No bankroll management.
- The models don't know they are participating in a competition. There's no tournament-induced temptation to take wild risks or play it too safe.
- The request (aka the prompt) is identical for all models and all events. It was locked in before the experiment started and never changed.
- Models cannot see each other's bets — nor can they see their own history.
- Everyone has access to web search for the freshest, most up-to-date information. No confinement to their training data cutoff.
- Odds for each bet were captured at the exact same time from a single, predetermined bookmaker.
- No line shopping for better odds at other bookies.
- Bets on a specific event are placed simultaneously to eliminate any discrepancies in potentially available live data.
- The bet must maximize +EV (expected value). The goal isn't to guess the winner, but to find an outcome where the model believes the probability is higher than what's implied by the bookmaker's odds. Mathematically, this is what yields a profit over a large sample size.
I should mention that as the organizer and a human, I naturally stepped outside these bounds. I knew I was in a tournament, saw the standings as they developed, tracked all the bets placed, etc. Long story short, I didn't exactly play by the same rules.
Observations
Right from the opening days, the biggest flaw in human (meaning my own) decision-making became glaringly obvious. It was, of course, the emotional rush of gambling excitement.
A win makes you believe in your own luck, as if it's an infinite resource. You want to take bigger risks — after all, you're on a hot streak! And besides, it's obviously because I'm a genius analyst who knows everything better than anyone else. A loss leads to the exact same outcome, just with a different motivation: you want to chase your losses, which means hunting for larger odds. A streak of identical outcomes only amplifies this effect.
None of this is groundbreaking news, but the data backed it up perfectly.
The models, on the other hand, were completely cold-blooded. Often, they would explicitly state that there wasn't a single potentially profitable bet in the entire bookmaker line and that it was better to skip the match entirely. This went against our rules, so they had to settle for the outcome with the mildest -EV — essentially, playing to minimize losses.
Another unexpected issue for me was the desire to make an "impressive" or flashy choice. I only realized this toward the end. You'd think, I'm competing against a bunch of hardware — who am I trying to impress? But no. Every now and then, a thought would pop into my head: "Okay, model, you bet on this outcome. The odds are around 1.5. Sure, it's a safe bet. It'll probably win. But where's the fun in that?"
The Results
Well, here it is.
| # | Model | Won | Lost | Void | Win Rate | Avg Odds | Profit | ROI |
|---|---|---|---|---|---|---|---|---|
| 🥇 | Gemini Flash 3.5 | 60 | 41 | 3 | 57.7% | 1.93 | +1381 | +13.3% |
| 🥈 | MiniMax M3 | 65 | 39 | 0 | 62.5% | 1.84 | +1155 | +11.1% |
| 🥉 | Grok 4.3 | 58 | 46 | 0 | 55.8% | 2.05 | +693 | +6.7% |
| 4 | GPT 5.5 | 60 | 44 | 0 | 57.7% | 1.89 | +662 | +6.4% |
| 5 | GLM 5.1 | 57 | 47 | 0 | 54.8% | 1.80 | −435 | −4.2% |
| 6 | Claude Opus 4.8 | 48 | 56 | 0 | 46.2% | 2.05 | −1038 | −10.0% |
| 7 | Qwen 3.7 Plus | 46 | 58 | 0 | 44.2% | 1.92 | −2284 | −22.0% |
| 8 | Human body | 37 | 67 | 0 | 35.6% | 2.45 | −2334 | −22.4% |
The models completely wiped the floor with me. A rock-solid last place for humanity. On the bright side, I'm the champion of the highest average odds. Flashy as hell, I’ll give myself that.
The top three podium finishes were quite unexpected. Though, depends on how you look at it.
First place went decisively to Gemini Flash 3.5. It took the lead after the 47th match (out of 104) and never let anyone unseat it. By the end of the tournament, its profit stood at just over 13 percent ROI. Considering that over a long distance, anything above a five percent return is considered masterclass territory — while the vast majority of players always sink into deep negatives — this is impressive.
Second place was snatched by MiniMax M3. The dark horse of the tournament, without a doubt. It even had a shot at the title near the finish line, but fell just short. What's fascinating is that after the group stage (the first 72 matches), MiniMax had a flat 0% ROI. It generated all of its profit during the knockout phase — in the final 32 matches.
The bronze medal went to Grok 4.3. It flirted with the top two spots throughout the tournament but couldn't hold on, despite putting up a long fight.
Fourth place went to the last model to stay in the green — GPT 5.5.
Everyone else finished in the red. Interestingly, the split between the top four and bottom four took shape after the 59th match and remained unchanged until the final whistle. A clear difference in class.
Once again, Claude Opus 4.8 was a major disappointment. A model that expensive shouldn't be delivering such lackluster results, especially compared to its peers.
There isn't much to say about GLM 5.1 and Qwen 3.7 Plus. Their results were pretty much within expectations.
One could draw a surface-level conclusion that the best results came from models built by companies with robust search engines and access to massive streams of live data. The only exception to this rule is MiniMax. They must have some secret sauce of their own.
Deep Dive
After the tournament wrapped up, I fed all the data back into an AI agent and asked it to spot patterns hidden beneath the surface of the final standings. Some of the findings are genuinely intriguing.
GPT: Simultaneously the Best and the Worst
GPT 5.5 was full of surprises. If you break down the performance by betting markets, you find something almost absurd.
| Market | GPT 5.5 Result |
|---|---|
| TOTAL | +1349 |
| AH (Asian Handicaps) | −1044 |
On Totals, the model was the best in the entire tournament. On Asian Handicaps, it was the worst. It excelled beautifully at one thing and blundered just as confidently in the other.
Claude's Totalitarianism
| Model | Bets on TOTAL | Markets | Unique Choices |
|---|---|---|---|
| Claude Opus 4.8 | 77.9% | 5 | 12 |
| Gemini Flash 3.5 | 54.8% | 6 | 41 |
| MiniMax M3 | 57.7% | 7 | 29 |
| Human | 48.1% | 8 | 45 |
Claude placed nearly four out of every five bets on the Totals market and used a grand total of just 12 unique betting choices throughout the entire World Cup — the lowest variety among all participants. For comparison, I had 45. Funnily enough, neither extreme worked out in the end.
Gemini: More Consistent Than MiniMax
| Model | Group Stage | Playoffs | Total |
|---|---|---|---|
| Gemini Flash 3.5 | +759 | +622 | +1381 |
| MiniMax M3 | −1 | +1156 | +1155 |
This is easily the prettiest table in the whole analysis.
Gemini practically secured the tournament during the group stage. After the first 72 matches, it racked up +759 units, while MiniMax was sitting at a dead zero. Sure, MiniMax had an explosive playoff run, but it wasn't enough to catch up. Gemini simply maintained its high baseline all the way to the final match.
And One More Thing
I expected the models to differ by just a few percentage points in accuracy. In reality, they turned out to have completely distinct personalities.
One was almost exclusively hunting for Totals. Another was obsessed with draws. A third understood Totals brilliantly but collapsed on Handicaps. A fourth made virtually all of its profit in the playoffs.
Each model is starting to develop its own distinct decision-making style. And sometimes, that unique style matters more than the raw performance gap between the models themselves.
If you want to dig into the raw numbers, the full statistics and data are available here.
What's Next
I absolutely loved running this project. Even before it ended, I had already filled a whole notepad with ideas: what to optimize, what pitfalls to avoid, and what variables to tweak next.
I really hope I don’t abandon this thing. I want to polish the data collection workflow, play around with the rules, and test out other models, tournaments, and maybe even different sports.
Follow the updates on the website.