Why more puzzles stopped fixing your blunders
What puzzle sites actually train, what they leave out, and how to rebuild tactics practice out of the mistakes in your own games.
The short answer
Puzzles stop helping because they tell you a tactic exists. Your games never do that, and spotting that something is there is the part you are failing at, not the calculation once you know.
When you open a puzzle, you already know four things: it is your move, there is a forcing solution, it is probably short, and it probably starts with a check, capture, or threat. Strip those four away and the same position becomes much harder. That gap between "solve this" and "play this" is where puzzle training that works has to live.
What puzzles do train
This is not an argument that puzzles are useless. They are the cheapest way to build a stock of patterns: back rank, smothered mate, deflection off a defender, the knight fork geometry that lets a knight on e5 hit c6 and f7 at once. You need those in memory, retrievable in under a second, and repetition is how they get there.
They also train calculation under time pressure, which matters in blitz. And they are the only tactics practice you can do on a phone in a queue.
What they do not train is detection in a position that has not been pre-labelled, and they do not train the negative case: looking at a quiet position, checking, and correctly concluding that there is nothing. Most positions in your games are the negative case. Puzzle sets almost never contain one, because a puzzle with no solution is a bug.
Where your errors actually are
Across 155,054 engine-classified moves from 4,959 analysed games on Mated, 3.8% of moves played are outright blunders and another 4.2% are mistakes. That is roughly one real error every 13 moves, before you count the 18.4% classed as inaccuracies. In a 40-move game you are producing three moves that change the evaluation meaningfully.
In the same dataset, those errors are not spread evenly. 9.0% of middlegame moves are a mistake or worse, against 2.3% of endgame moves. The middlegame is where the piece count is high, the position is unforced, and nobody has told you to look.
One more number from that dataset: the average move gives away 94 centipawns against the engine's best, but the median is far lower, because a small number of very bad moves carry most of the damage. That shape matters more than the average. You are not losing games by being a bit worse than the engine on every move. You are losing them on two or three moves where you drop a piece, and your remaining forty moves are fine.
A caveat worth saying out loud: these are the games of people who signed up to have their chess analysed. That is not a random sample of everyone online, and the rating spread is not evenly distributed. Treat the shape of the finding as solid and the exact percentages as belonging to this group.
A concrete example of what a puzzle would not have shown you
You are Black in something Ruy Lopez shaped. Your knight sits on c6. You play ...Bb4, pinning a white knight on c3, because pinning knights is a thing you do. White plays Qa4. The queen hits the bishop on b4 along the fourth rank and the knight on c6 down the a4–e8 diagonal. You lose a piece.
Now imagine this as a puzzle from White's side. Title: double attack. You would find Qa4 in three seconds, because the site told you a tactic was there and you know a queen move on move nine is usually the point.
The failure in the actual game was on Black's move, one ply earlier, and it was not a calculation failure. You did not calculate anything, because nothing announced itself. Your bishop and your knight were both loose and both on light-square-adjacent squares, and no alarm went off. No amount of solving White-to-play double-attack puzzles installs that alarm. What installs it is seeing your own ...Bb4 punished, then being shown the position again next week and asked what is wrong with it.
How to turn your own games into the puzzle set
The method is simple and slightly tedious, which is why almost nobody does it.
Take your last ten games. Run them through any engine analysis — the free one on the Lichess analysis board is fine. Find every position where the evaluation swung by more than about 150 centipawns on one of your moves. For each one, do not look at the engine line yet. Set the board to the position before your move and try to find what you missed, with a clock running, say two minutes. Then check.
Then, and this is the part that separates it from post-mortem browsing, keep the position. Come back to it in three days and again in two weeks. You are not trying to remember the move. You are trying to reach the point where the pattern in that position triggers something before you have decided what to play.
Mated automates exactly this loop — it pulls your recent games from the Chess.com and Lichess APIs, runs Stockfish over every position in the browser, and builds a fifteen-minute session out of your own errors. But there is no secret in the software. If you are willing to keep a text file of FEN strings and revisit them on a schedule, you get most of the same thing for nothing.
The detection drill puzzles cannot give you
Because your own error positions come with no label, you can train the negative case with them too. Mix in positions from the same games where you played a good move. Now you are being asked "is there something here?" rather than "find the thing here", which is the actual question at the board.
A few things that make this work better in practice, from watching how people run it:
- Start from the position before your mistake, not after. The mistake itself is a hint.
- Give yourself real clock time. Solving in ten seconds trains a different skill than the one that failed.
- Write down, in words, why you played the losing move. "I was pinning the knight" is a more useful note than the engine's refutation.
- Weight your middlegame positions. That is where 9.0% of moves in the dataset above were a mistake or worse.
- Do not skip positions where you were already losing. The errors that turned -1 into -6 are the same errors that turn +1 into -1.
What nobody has measured well
How much tactics practice transfers to over-the-board play is genuinely unsettled. There is decent evidence that deliberate, effortful practice beats casual play for skill growth in general, but for chess specifically nobody has run the clean study: matched groups, one solving generic puzzles, one solving their own errors, blunder rate measured over a few hundred subsequent games.
So treat the claim in this article as a mechanism argument, not a proven result. The mechanism is that puzzle sets remove the cue-detection step and your games do not, and that the errors costing you rating are concentrated in a handful of moves rather than smeared across all of them. Both of those are things you can check in your own data this week.
The measurable version: count your blunders per game over your next thirty games and compare it to your last thirty. Not your puzzle rating. Your puzzle rating measures how good you are getting at puzzles.
The test that tells you which problem you have
Before you change anything, find out whether your tactics are failing on detection or on calculation. Take five of your own blundered positions and show them to yourself with the label attached — "White to play, there is a tactic". If you find four out of five, your pattern stock is fine and your problem is that you are not looking. More puzzles will not fix that. A checking habit will: after you pick a move and before you touch the mouse, name every one of your pieces the opponent's next move could hit.
If you find one out of five even with the label, you have a genuine pattern gap, and the boring answer is right. Go do a thousand puzzles, sorted by theme, until forks and deflections are recognition rather than calculation. Then come back to your own games.
Most players between 800 and 1800 are in the first group and train as though they were in the second, because the second is more pleasant. Solving puzzles feels like progress. Looking at the move where you hung your bishop, again, on a Tuesday, does not.
Your puzzle rating measures how well you solve positions somebody has already told you are solvable. Your blunder rate measures the thing that actually costs you games — go count that one instead.
Questions
- How many puzzles a day should I do?
- If your pattern stock is the problem, twenty to thirty untimed puzzles a day is plenty and more has diminishing value. If detection is the problem, the number is close to irrelevant — ten minutes on three positions from your own games will do more than a hundred puzzles. Test which group you are in first.
- Is Puzzle Rush or Puzzle Storm bad for me?
- It trains fast recognition of patterns you already know, which is useful for blitz. It trains nothing about positions where you do not yet know a tactic exists, because the format punishes you for thinking. Treat it as a warm-up, not as practice.
- Do I need software to do this?
- No. The Lichess analysis board is free, gives you engine evaluations move by move, and lets you copy a FEN out of any position. A text file of FENs plus a calendar reminder is a complete version of this method. Software saves you the sorting and the scheduling, not the thinking.
- My puzzle rating is 2200 and my game rating is 1300. What is going on?
- Almost certainly detection and time management, not pattern knowledge. Puzzles hand you a cue and a guarantee; your games hand you a quiet position on move 18 with two minutes on the clock. The 9.0% middlegame error rate in the dataset above comes from exactly those positions, not from missed mates in three.
See this in your own games
Mated reads your last games, scores them across five categories and builds a fifteen-minute session out of what it finds. No card.