MATED
← Blog

Most AI chess tools analyse. Few tell you what to fix.

A concrete test for whether an AI chess platform is doing analysis or doing coaching, and what your own error data looks like once somebody aggregates it.

·7 min read·Includes figures from Mated’s own move data

The difference, in one sentence

A platform that runs an engine gives you a verdict on a position. A platform that reads your games gives you a verdict on you — which of your errors repeat, where in the game they happen, and what you should do in the next fifteen minutes because of it.

Both use the same engine. Stockfish is free, and running it in a browser is no longer hard. The engine is not the product. What is done with the output is the product.

What the engine actually hands over

Per move, an engine gives you a number: the evaluation before your move, the evaluation after, and the best line it found at whatever depth it had time for. Everything else on an analysis page is derived from that. "Blunder", "mistake", "inaccuracy" are just thresholds on the difference. The accuracy percentage is a formula over the same differences.

That is genuinely useful and you should use it. But notice what it cannot tell you. It does not know that you had 40 seconds left. It does not know you played the move because you were following a plan you formed six moves earlier and never rechecked. It does not know that you have made the same category of error in eleven of your last twenty games, because it is only looking at one game.

Concrete version. You drop a bishop on move 24. The engine flags move 24 as the blunder, shows you the capture, and moves on. The real event was move 19, when you decided to push on the queenside and stopped looking at the diagonal your bishop was sitting on. A single-game engine report will never say that. A tool that has seen your last forty games can at least say: your blunders cluster three to five moves after you commit to a plan.

What your error rate looks like when somebody adds it up

Across 83,933 engine-classified moves from 2,655 analysed games on Mated, 3.6% of moves played are outright blunders and another 4.1% are mistakes. That is roughly one real error every 13 moves, before you count inaccuracies, which run at 18.2%.

Read that as a workload, not a grade. In a 40-move game you are making around three moves that change the evaluation materially. Three. Not thirty. If a platform hands you a game report with fifteen red marks and no ordering, it has given you a list where you needed a shortlist.

That is the first thing aggregation buys you: it makes the shortlist obvious. Three errors a game across forty games is roughly 120 mistakes, and they are not 120 different mistakes. They are a handful of patterns repeated.

The 93-centipawn number and why averages lie here

In the same set of 83,933 moves, the average move gives away 93 centipawns against the engine's best. The median is far lower, because a small number of very bad moves carry most of the damage.

This matters for how you read any AI chess platform that leads with an average. Average centipawn loss and accuracy scores compress a lopsided distribution into one number, so a game where you played 38 clean moves and threw a piece on move 39 can score similarly to a game where you drifted for 40 moves. Those two games need completely different homework.

If a report shows you one number per game, it is telling you the temperature of the room while the fire is in one corner. Ask for the distribution: how many moves cost more than 200 centipawns, and what did they have in common.

Errors are not spread evenly across a game

From the same 83,933 moves: 8.7% of middlegame moves are a mistake or worse, against 2.1% in the endgame. The middlegame is where the damage happens.

Do not read the endgame number as a clean bill of health. Fewer of your games reach an endgame at all, and the ones that do are often already decided, with forced or obvious moves. The endgame sample is easier than the endgame you actually fear. What the split does justify is prioritisation: if you have fifteen minutes a day and you spend all of it on rook endgames, you are practising the phase where 2.1% of your moves go wrong and ignoring the phase where 8.7% do.

This is the second thing aggregation buys you. Phase-level error rates are invisible in a single game report and obvious across two hundred of them.

Five questions that separate the two kinds of tool

Run these against whatever you are considering, including us.

  • Does it read your actual games from Chess.com or Lichess, or do you paste in PGN one game at a time?
  • Does it group errors by cause — hanging pieces, missed opponent threats, bad trades, time trouble — or only by move number?
  • Does the output change what you do tomorrow, or does it stop at a report?
  • Does it show you the distribution of your losses, not just an average or an accuracy percentage?
  • Can it tell you which phase and which time control your errors concentrate in?

Where a product is genuinely doing something you cannot

Reading forty games move by move, classifying every move, tagging causes, and turning that into a queue of positions is a few hours of work by hand. Doing it again every week is not going to happen. That loop is the honest case for a tool: Mated pulls your recent games through the public Chess.com and Lichess APIs, runs Stockfish over every position in your browser, and builds a fifteen-minute session out of the positions where you actually went wrong rather than a generic puzzle set.

What no tool can do is tell you what you were thinking. The engine sees the move, not the reason. The cheap version of this is a two-column note: the position where you went wrong, and one sentence on why you played it. "I assumed he couldn't take." "I was down to 20 seconds." "I never looked at his last move." Twenty of those and the pattern is usually embarrassing and specific, and you did not need to buy anything.

Also worth saying plainly: nobody has measured well whether drilling your own mistakes beats drilling a good generic curriculum at your rating. It is a reasonable bet — you are spending time on positions you demonstrably get wrong — but it is a bet, not a measured result. Anyone quoting you a rating gain figure for personalised training is making it up.

A test you can run this week

Take your last ten losses. For each, find the single move that cost the most, by evaluation swing. Write down the move number and one sentence on the cause. Ten lines, twenty minutes.

Now look at the sheet. If eight of the ten are between moves 15 and 30, you have confirmed the middlegame clustering in your own games and you know which phase to spend your time on. If seven of them say some version of "I didn't look at what he was threatening", you have found a habit, and no amount of opening study will touch it.

That sheet is what a coaching platform is trying to produce automatically. If a tool cannot produce something at least that useful, it is an engine with a nice front end, and you already have one of those for free.

The useful question is not which platform has the strongest engine — they all have the same one. It is whether anything in the product has looked at more than one of your games at a time, because every pattern worth fixing is invisible inside a single game report.

Questions

Is an AI chess platform just Stockfish with a chatbot on top?
Sometimes, yes. Stockfish is free and open source, so the engine is not what distinguishes products. Judge the layer above it: whether it reads your whole game history, whether it groups your errors by cause, and whether it produces work for you to do. Language-model commentary on positions is separate again and is frequently wrong about concrete lines, so treat its explanations as prompts for your own thinking, not as verdicts.
If one move in 13 is a real error, should I be analysing every game?
No. Three errors a game means most of the moves in a full-game review are moves you played fine, and you will spend your review time confirming that. Analyse in batches instead: look at ten games at once and ask what the errors have in common. The pattern across ten games is more actionable than a deep dive into one.
Does higher accuracy mean I'm improving?
Not reliably. Accuracy is derived from centipawn loss, and in our data the average loss of 93 centipawns per move is dragged up by a small number of very bad moves while the median sits far lower. Two games with the same accuracy can contain completely different problems. Track the count of moves that cost more than 200 centipawns instead — it moves for reasons you can actually explain.
My endgame error rate looks low. Can I stop studying endgames?
Careful. The 2.1% endgame error rate in our data comes from the endgames people actually reached, many of them already won or lost with forced moves. It does not mean your endgame technique is sound in a balanced rook ending. It does mean that if you only have fifteen minutes, the middlegame is the better bet at 8.7%.

See this in your own games

Mated reads your last games, scores them across five categories and builds a fifteen-minute session out of what it finds. No card.