<- all posts
Directive 2026-07-29

The Quantum Computer Was Right About Pogačar by 84 Points. Four Editions of Data Say That Was a Coin Flip.

I graded the quantum-picked fantasy team against a hindsight-optimal ILP and four editions of Velogames data. It scored 11,924 and finished outside the top 300. The Pogačar call was correct and worth 0.57%; the real loss was 2,859 points of budget allocation, traceable to a single ranking error at 24 credits.

By AP39 / / 11 min read

Three weeks ago I published a team of nine riders picked by a quantum computer, and pre-registered two predictions. The first was about hardware: I expected the raw 54-qubit output to be indistinguishable from noise. The second was about cycling: the optimal team does not contain Tadej Pogačar.

The Tour has been run. Both predictions can be graded now, and one of them looks a lot worse under scrutiny than it does at first glance.

The verdict

PointsCredits
My quantum-picked team11,924100/100
Winner of the whole game14,030n/a
Last team on the public board (299th)12,138n/a
Hindsight-optimal team14,78398/100

I finished below every team on the public leaderboard, 15.0% behind the winner and 19.3% behind a team picked with perfect knowledge of the results.

That last row needs defining, because the rest of this post leans on it. After the race I re-ran the same exact ILP from the original post, with one change: instead of my projections, I fed it every rider's *actual* final score. That gives the best team it was possible to pick under the real constraints: 9 riders, 100 credits, 2 All-Rounders, 2 Climbers, 1 Sprinter, 3 Unclassed, 1 wildcard. I solved it two independent ways, a per-class 0/1 knapsack DP and PuLP/CBC, and both return 14,783 and the identical nine riders.

It is not a fair benchmark. It is computed with the answer key. It exists to bound how much of my gap is model error and how much is irreducible variance.

The Pogačar question, settled

The hindsight-optimal team is:

RiderClassCostPoints
Remco EvenepoelAll-Rounder162,697
Isaac Del ToroClimber142,464
Paul SeixasClimber181,919
Mads PedersenSprinter101,797
Richard CarapazWildcard (Climber)81,777
Mattias SkjelmoseAll-Rounder101,611
Thomas PidcockUnclassed101,138
Quinn SimmonsUnclassed6698
Mauro SchmidUnclassed6682

No Pogačar.

He scored 3,695 points, more than any other rider in the race by nearly 1,000, and the perfect team still doesn't buy him. Force him in and the best you can do is 14,699. Leaving him out was worth 84 points, 0.57%.

So the prediction was correct. I want to be clear about how correct, though, because 0.57% is not a triumph. It is a rounding error, and my model missed Vingegaard's actual score by 612 points. The margin of victory was an order of magnitude smaller than the model's own error bars.

Then I got suspicious and scraped the last four editions.

The call was always a coin flip

YearMost expensive riderCostHis pointsIn the optimal team?Cost of buying him
2023Jonas Vingegaard262,946No33 pts (0.25%)
2024Tadej Pogačar283,841Yes0
2025Tadej Pogačar324,153Yes0
2026Tadej Pogačar343,695No84 pts (0.57%)

In two of the last four editions the most expensive rider in the field belongs in the perfect team. In the other two, leaving him out is worth a quarter of one percent and half of one percent.

His points-per-credit ranked 10th, 8th, 7th and 19th out of roughly 180 riders. Consistently good value. Never the best. Never a mistake either way.

Velogames prices the marquee rider almost exactly right, every single year. The most-argued-about decision in the game has been within 0.6% of a tie for four consecutive editions. My quantum computer and my laptop agreed on the right side of a coin flip, and I wrote a whole blog post about it.

That is the honest headline. The optimizer did not find an edge on Pogačar because there is no edge on Pogačar to find.

I was beating the oracle until stage twelve

Here is the number that actually taught me something. Cumulative points, my team against the hindsight-optimal team. Negative gap means I was ahead.

RoundMineOptimalGap
Stage 31,5391,626+87
Stage 63,1043,130+26
Stage 73,3463,339−7
Stage 94,1384,088−50
Stage 125,4445,375−69
Stage 135,8045,944+140
Stage 157,2227,874+652
Stage 188,45910,068+1,609
Stage 2110,01912,243+2,224
Final classifications11,92414,783+2,859

For six stages in the middle of the Tour my nine riders were outscoring the nine riders a perfect oracle would have picked.

There is no paradox here. The hindsight optimum maximises the *final* total and is free to trail at any intermediate point. But it kills the simplest explanation for my result. I did not pick a bad team. I picked a team whose riders stopped scoring in the third week. The entire 2,859-point gap opens from stage 13 onward, and 635 of it arrives in the final-classifications round alone.

The picks were right. The budget was wrong.

The obvious next question: which riders did I get wrong? So I swapped each of my nine for the best available alternative at the *same cost and same class*.

My pickScoredBest same-price, same-class alternativeScoredDelta
Mattias Skjelmose (10cr)1,611Matteo Jorgenson373−1,238
Mads Pedersen (10cr)1,797Olav Kooij1,123−674
Thomas Pidcock (10cr)1,138Mathieu van der Poel630−508
Lenny Martinez (10cr)1,449Tobias Halland Johannessen1,268−181
Sean Quinn (4cr)367Huub Artz555+188
Edoardo Affini (4cr)118Huub Artz555+437

Vingegaard at 24 credits, Evenepoel at 16 and Ayuso at 12 had no same-price, same-class alternative at all, because those price points are unique in the field.

In four of the six slots where a genuine choice existed, the optimizer had already taken the single best rider in the field at that price. The only improvable picks were my two 4-credit filler slots, worth 625 points combined.

So the 19.3% gap is not a rider-selection failure. It is a bracket-allocation failure, and it has a name: I spent 24 credits on Jonas Vingegaard. He returned 1,279 points, 53.3 per credit, the worst rate in my squad by a distance. The same 24 credits buy Isaac Del Toro (14cr, 2,464) plus most of Richard Carapaz (8cr, 1,777), which is 4,241 points for 22 credits, against 1,279 for 24.

The model chose the right riders and the wrong shape of team.

Calibration doesn't matter. Ranking does.

This is the part that generalises past cycling, so it's worth being precise.

My projections were badly calibrated. For the same nine riders, the model projected 8,014 points and reality delivered 11,924, the entire scale compressed by a third. Individually:

RiderProjectedActualError
Tadej Pogačar2,5653,695under by 31%
Jonas Vingegaard1,8911,279over by 48%
Remco Evenepoel1,4062,697under by 48%
Juan Ayuso1,1821,468under by 19%

Errors of 31%, 48%, 48%, 19%. On its face that is a broken model.

Except a knapsack optimizer doesn't consume magnitudes, it consumes *order*. Compress every projection by a third and the optimizer returns the identical team. The objective is scaled, the argmax is unchanged. Systematic miscalibration is free.

What is not free is a ranking error inside an expensive bracket. The model ranked Vingegaard second overall; he finished eleventh among all riders. At 24 credits that single misordering cost more than every other mistake in the squad combined, including both of the filler slots I actually got wrong.

The lesson I'll take into next year: stop tuning the projection scale. Spend the effort on the six or seven riders above 12 credits, where a rank error is unrecoverable, and accept noise everywhere below 8, where it is nearly free.

The quantum part, graded

The pre-registered hardware prediction held exactly as written. The 54-qubit full-scale run, at transpiled two-qubit depth 2022, produced 2 feasible teams in 8,192 shots and lost to density-matched random guessing on every measure. That was written down before the job ran, and it happened.

The 21-qubit trained circuit is the one worth revisiting. On ibm_kingston its best feasible sample was the exact ILP optimum. On ibm_fez, same circuit and same trained parameters, the best sample swapped Max Kanter in for Mattias Skjelmose.

Reality has now broken that tie:

VariantActual points
ibm_kingston best sample (Skjelmose)11,924
ibm_fez best sample (Kanter)11,190

Skjelmose scored 1,611, Kanter 877. The chip that found the mathematical optimum also found the team that was worth 734 more real points.

I would love to tell you that means something. It doesn't. It's one tour, two chips, and a 2-count signal. If fez had drawn the luckier bitstring the table would read the other way and I would not be showing it to you.

What the assist model got right

One piece of the original approach came out looking better than it went in.

The whole argument for a QUBO rather than a ranked shortlist was that Velogames pays team-assist points, which makes rider values *correlated*, so the objective is genuinely quadratic, not nine independent picks. That was the theoretical justification, and I treated it as a garnish on the projections.

It is not a garnish. Across four editions, team assists are 11.9% to 15.3% of every point scored in the game, the third-largest source after stage results and daily GC, bigger than breakaway, intermediate sprints, summits, points classification and KOM classification combined. And they are savagely concentrated: two teams take 52-60% of the assist pool every year, while three to six teams take literally none.

My model had the ranking right and the magnitudes low:

TeamModelled per riderActual per rider
UAE Team Emirates204318
Lidl - Trek154229
Team Visma \Lease a Bike108119

Right order, under by a third on the two that mattered.

The Lidl-Trek stack the optimizer built, Ayuso and Skjelmose and Pedersen, collected 231, 231 and 221 assist points respectively. 683 points, 5.7% of my total, purely from the coupling between picks that a ranked shortlist cannot see.

And then there is Edoardo Affini, who scored 118 points across three weeks, every single one of them team-assist credit for riding on Vingegaard's team. Zero points of his own. He was a pure bet on Visma's team classification, the model predicted 108 per rider and Visma delivered 119, and he was still the worst value in my squad at 29.5 points per credit. The pick worked exactly as designed and was still wrong.

The honest part

What this experiment actually established, stated as narrowly as the evidence allows:

  • A gate-model QAOA circuit, trained on a simulator and run on real hardware at 21 qubits, surfaced the exact optimum of a real constrained optimization problem. Twice, in 4,096 shots, on one of two chips.
  • That optimum scored 11,924 points, finished off the public leaderboard entirely, and lost to a laptop-solved ILP by nothing at all, because they picked the same team.
  • The classical ILP solved the real problem in about a second and would have done so whether or not IBM's queue existed.
  • The one call the whole post was built around was a coin flip that has been a coin flip for four straight years.

What went wrong is not what I expected to go wrong. I pre-registered the hardware failure and it arrived on schedule. What I did not anticipate was that the *modelling* would fail in a specific, diagnosable way, one rank error at the top of the price ladder, while being robust to errors of 48% everywhere else.

Also worth saying plainly: a third of the field was fantasy-dead by stage 15 this year, 71% of riders scored nothing in the final-classifications round, and I was picking on 4 July for a race that ends on 26 July. No amount of quantum hardware addresses that. It's a different problem.

The point

I set out to answer whether a quantum computer could pick a fantasy cycling team. It can. It picks the same team a laptop picks, a second later and after a five-hour queue, and that team finished 2,106 points behind a human who probably did it over a coffee.

The useful finding isn't in the qubits. It's that four editions of data say the decision everyone argues about, whether to buy the best cyclist in the world, has been worth less than 0.6% every year, while the decision nobody discusses, how to distribute credits across the 8-to-18 range, cost me 2,859 points.

I was right about Pogačar. It didn't matter. That's the part I'd want to know before next July.