What an engine can't tell you about your own chess
The specific coaching questions an evaluation bar will never answer, and a protocol for answering them yourself after your next ten games.
The engine knows the position. A coach knows the player
An engine evaluates positions. A coach diagnoses habits. Stockfish can tell you that 24.Bxh6 loses a piece; it cannot tell you that you play a sacrifice like that every time you have been passive for fifteen moves and can't stand it any longer.
That gap is the whole reason "analysed my game and got nothing out of it" is such a common complaint. You get a list of thirty-one accurate judgements about thirty-one positions you will never see again. What you needed was one sentence about you.
Everything useful a coach does falls into four buckets an evaluation bar structurally cannot fill: grouping your errors into causes, telling you which of them matters at your level, deciding what to do next week, and asking what you were actually thinking when you played the move.
How often you actually go wrong, and where
Some numbers, so we are arguing about the same reality. Across 94,566 engine-classified moves from 3,013 games analysed on Mated, 3.7% of moves played are outright blunders and another 4.2% are mistakes. That is roughly one real error every 13 moves, before you count the 18.3% of moves classified as inaccuracies.
One error every 13 moves sounds survivable until you notice it is not distributed evenly. In that same set of games, 8.9% of middlegame moves are a mistake or worse, against 2.2% in the endgame. Your endgame technique is probably not your problem. The moves you play between move 12 and move 30, when the position stops resembling anything you have seen, almost certainly are.
The third number is the one that changes how you should read a game report. The average move in that set gives away 94 centipawns against the engine's best, but the median is far lower, because a small number of very bad moves carry most of the damage. Your games are not slowly leaking. They are fine, fine, fine, catastrophe.
That has a direct practical consequence. Chasing your average accuracy percentage upward is chasing the wrong quantity. If two or three moves per game account for most of the loss, the only question worth asking is what those two or three moves have in common — and that is a question about you, not about the position.
Ask for these five things by name
If you are paying a human, or prompting a language model, or picking software, these are the outputs to demand. None of them appear on an evaluation graph.
- A cause, not a classification. "Blunder" is a score. "You move the attacked piece instead of checking whether the attacker can be taken" is a cause you can train.
- A repeat count. Did this happen once, or in six of the last ten games? One instance is noise. Six is your curriculum.
- A ranking. You have maybe eight identifiable weaknesses. Which one costs the most rating points per hour of work? A coach will guess; an engine will not even attempt it.
- A stop-doing instruction. Advice you can follow while your clock is running: "before every capture, look for the recapture you have not counted." Not "improve your calculation."
- The question the engine can't ask: did you see the move and reject it, or never see it at all? These are completely different diseases and they have completely different treatments.
One move, three diagnoses
You are in a Italian-ish middlegame, roughly equal, both kings castled. You play a knight to the rim where it can be won by a pawn push. Engine: blunder, −3.1. That is all it says, and everything you need to know is in what happened before you released the piece.
Case one: you never saw the pawn push. That is a scanning failure. The fix is mechanical — a fixed pre-move check on every non-forced move, one loop over your opponent's pawn moves and checks. Boring, cheap, works.
Case two: you saw the push and calculated that you had a defence, and the defence had a hole in it. That is a calculation failure and the fix is different: slow, deliberate calculation training on positions where the refutation is three moves deep, not scanning drills.
Case three: you saw it, knew it was bad, and played it anyway because you had 40 seconds and 22 moves to make. That is not a chess problem at all. That is a time management problem, and no amount of tactics puzzles will touch it.
Same move, same evaluation drop, three completely different weeks of work. An evaluation bar gives you one output for all three. A coach's first question — "what did you see?" — separates them in about ten seconds. This is why annotating your own games while you still remember what you were thinking beats analysing them a week later.
What "AI chess coach" currently buys you, honestly
Two different things get sold under that phrase, and they fail differently.
The first is an engine with a report layer: Stockfish plus classifications, accuracy percentages, maybe a phase breakdown. This is genuinely valuable and it is also mostly solved. What it adds over raw engine output is aggregation — seeing that the same error type shows up across games rather than in one. That is the part worth paying for, and it is the part that answers the question "which of my problems is biggest." It is roughly what Mated is built to do: run the engine over your recent games, group what it finds into categories, and turn the biggest group into fifteen minutes of drilling from your own positions.
The second is a language model that writes prose about your moves. Useful for one specific thing: turning a position you already understand into words you can remember. Unreliable for the thing people want it for. Ask it to calculate and it will produce plausible lines containing illegal moves, and it will describe pieces that are not on the board. If you use one, keep an engine open beside it and treat every concrete claim as a suggestion to verify. Prose about a position is not the same skill as evaluating one.
Nobody has published a good measurement of whether either kind of tool actually improves rating faster than just playing longer games and reviewing them yourself. Be suspicious of anyone, us included, who tells you otherwise with a number attached.
The version you can do for free this week
You do not need a product for the diagnosis step. You need a text file and about twelve minutes per game.
Take your last ten losses. For each, find the two moves where the evaluation fell furthest — not every inaccuracy, just the two worst, because that is where the damage is concentrated. Write one line per move in this format: move number, what you played, and which of the three cases it was (didn't see it, miscalculated it, played it under time pressure).
After ten games you have twenty lines. Sort them. Whatever category has the most lines is what you work on, and you will probably be surprised, because it is rarely the thing you tell your friends you are bad at. Most players at 1200 say they need openings and have twenty lines that say they need to look at their opponent's checks.
Do this once a month, not once. The value is in the second sort, when you can see whether last month's category shrank.
The thing worth logging that no engine can see
Record your clock, not just your moves. Next to each of those twenty error lines, write how much time you had left. Middlegame error clustering and time trouble are hard to separate — the middlegame is both the most complex phase and the phase where a 10+0 player is starting to feel the clock, and the numbers above cannot tell you which of those is driving your 8.9%.
But you can tell, for yourself, in a sample of one. If most of your worst moves came with under two minutes on the clock, your calculation is probably better than your results suggest and you are simply playing a time control that does not let you use it. That is a fifteen-second fix — change the time control — and it is invisible to every analysis tool that only reads the moves.
Your rating is not being drained evenly; it is being taken by two or three moves a game, and the useful question about those moves is not what you should have played but what you were thinking when you played them.
Questions
- Can I just paste my game into ChatGPT and ask what I did wrong?
- For explanations of ideas you half-understand, yes, and it is often clearer than an engine. For anything concrete — whether a piece was defended, whether a line ends in mate, what the best move is — verify it against an engine. Language models generate illegal moves and describe pieces that have already been captured, and they do it confidently. Use one for the sentence, not for the calculation.
- Should I learn the engine's best move in each position?
- Usually no. The engine's move is often correct for reasons that will not repeat and that require calculation you cannot yet do. What transfers is the cause of your error, not the refutation of it. If the engine says you missed a queen sacrifice, that is interesting trivia. If it shows you left a knight loose for the third game running, that is the lesson.
- How many games do I need to analyse before patterns show up?
- Ten losses is enough to see the biggest category. Nobody has measured the right number properly, but the reasoning is simple: at roughly one real error every 13 moves, ten games gives you dozens of errors, and you only need the two worst per game to find a repeat. Analysing fifty games at once mostly produces a spreadsheet you never open.
- Do I need a human coach at 1400?
- Not to find your weaknesses — you can do that with a text file and a free engine, as described above. A human is worth the money for two things: telling you which weakness to ignore, and noticing the psychological pattern you will defend rather than admit. That second one is genuinely hard to get from software, including ours.
See this in your own games
Mated reads your last games, scores them across five categories and builds a fifteen-minute session out of what it finds. No card.