MATED
← Blog

What an AI chess coach is actually good for

A week-by-week plan for using engines and AI tools on your own games, including the parts where the machine is wrong.

·8 min read·Includes figures from Mated’s own move data

The short answer

Hand an AI your own games and ask it to find the pattern in your mistakes. Ignore it when it explains why a move is good in a position no human at your level will ever have to hold.

That is the whole split. Engines and language models are excellent at labelling what happened and terrible at knowing which of the things that happened mattered to you. Your job in the loop is to supply the filter: which mistakes recur, which cost games, which happened when you had eight minutes left and which happened when you had eighty seconds.

Hand it your losses, not your positions

The instinct is to paste a position into a chatbot and ask what the best move is. That is the least useful thing you can do with it. Stockfish already answers that question better and for free, and knowing the best move in one position teaches you almost nothing about the next one.

The useful input is a batch. Ten games, engine-classified, with the clock times attached. Then the question is not "what was best here" but "where in these ten games did I lose more than 150 centipawns, and what do those moments have in common".

You can do this manually. Load ten games into Lichess's analysis board, note every blunder and the move number, and write the list in a text file. It takes about ninety minutes. Mated does the same pass automatically over your last games through the Chess.com and Lichess APIs, but the reason to use a tool here is time, not insight — the insight is in reading the list.

What the mistake data looks like across a lot of games

Across 83,933 engine-classified moves in 2,655 games analysed on Mated, 3.6% of moves played are outright blunders and another 4.1% are mistakes. That is roughly one real error every 13 moves, before you count the 18.2% classed as inaccuracies. If your games run 40 moves, you are making three or four genuine errors a game. Not one. Not ten.

That number reframes what "cleaning up" means. You are not hunting a single leak. You are trying to move three errors a game down to two, and the way to do that is to know where they happen.

In the same dataset, they do not happen evenly. 8.7% of middlegame moves are a mistake or worse, against 2.1% of endgame moves. The middlegame is where the damage is concentrated — which is inconvenient, because the middlegame is the hardest phase to study with flashcards and the phase where most published training material is thinnest.

One more from the same 83,933 moves: the average move gives away 93 centipawns against the engine's best, but the median is far lower. A small number of very bad moves carry most of the loss. This is the single most useful fact for deciding what to work on. You do not have a general accuracy problem. You have a handful of catastrophic moments per game, surrounded by moves that are basically fine.

Week one: build the list

Do not train yet. Spend the first week just producing an honest inventory.

Play your normal games at your normal time control. After each one, run the engine and record every move that lost 100 centipawns or more. For each, write down four things: move number, phase, how much time you had, and — in your own words, before you look at the engine line — why you played it.

That last field is the one AI cannot fill in for you. "I thought the knight was defended." "I saw the fork and stopped looking." "I was winning and got careless." "I had no idea what to do so I moved a rook." These are different problems with different fixes, and the engine's output looks identical for all four.

By the end of the week you will have twenty to thirty entries. Now the AI is useful: paste the list of your own reasons into a language model and ask it to group them. It is good at that. It will tell you that eleven of your thirty entries are some version of "I checked whether the piece was attacked but not whether the defender was pinned".

Week two to four: train the top two, ignore the rest

Take the two largest groups. Ignore everything else for a month, including the interesting-looking ones.

If your biggest group is missed defensive resources, your daily work is puzzles where the answer is a quiet defensive move, plus a habit: before every capture, name the piece that recaptures. If it is time-pressure blunders, your work is not tactics at all — it is playing 15+10 for three weeks and blitzing less. If it is "no idea what to do", you need plans in your own openings, and the fix is looking at ten master games in your structure, not a puzzle streak.

This is where most AI chess advice goes wrong. It will happily generate you a comprehensive study plan covering openings, middlegame strategy, endgame technique, and calculation. Given the middlegame concentration in the data above, a plan that spends a quarter of its time on endgames is misallocating your month. Push back on the model. Ask it to cut the plan in half, then in half again.

A fifteen-minute daily session drawn from your own recorded mistakes is enough at this stage. Mated builds that session for you and scores you across five categories so you can see whether the group you picked is actually shrinking; you can also build it yourself from your text file and a puzzle site with a custom set. Either works. The list is what matters.

What to ignore the machine about

Some engine output is actively misleading at 800 to 1800, and you should learn to skip it rather than argue with it.

The general rule: an evaluation is a fact about the position, not a fact about you. A move that is objectively second-best but keeps the position simple is often the right move for you to play, and no engine will ever say so.

  • Openings where the engine prefers a move by 0.2. At your level the difference between the first and fourth choice in a mainline is noise compared with whether you know the resulting plans.
  • Long forcing lines you would never find. If the refutation is eight moves deep with two only-moves, the lesson is not the line. The lesson might be that you should not have entered the position.
  • "Inaccuracy" labels in quiet positions. 18.2% of moves in the dataset above are inaccuracies. You cannot fix 18% of your moves and you do not need to.
  • Evaluations in dead-drawn endgames. Stockfish will tell you a move was a mistake in a position both players will draw regardless.
  • Chatbot explanations of concrete tactics. Language models are unreliable at reading a board. Trust the engine for lines and the model only for grouping and phrasing.
  • Anything about a piece sacrifice the engine says is fine. It is fine for the engine.

What the end of the month actually feels like

Probably not a rating jump. A month is 40 to 100 games depending on your time control, and rating moves slowly and noisily at that sample size. Anyone promising you 200 points in four weeks is selling something.

What you should notice instead is a change in the texture of your losses. Before, you lose games and cannot say why. After, you lose a game and can say the sentence out loud on move 22: I did this again, I stopped calculating when I saw a good move for myself. That sentence is the actual product of the month. It is also the thing that makes the next month cheaper, because you already have the list and you only have to update it.

Honest caveat: nobody has measured well how quickly this kind of self-directed error-grouping translates into rating. The mistake-frequency numbers above describe what is happening in games; they do not prove that any particular training routine changes them faster than another. Treat the plan as a way to stop wasting effort, not as a guaranteed rate of improvement.

The thing worth automating is the bookkeeping, not the thinking: let the machine count your errors and sort them, then decide for yourself which two are worth a month, and accept that you are ignoring the other five on purpose.

Questions

Can I just ask ChatGPT what move to play?
You can, and it will often be wrong. Language models do not calculate; they produce plausible chess prose. Use an actual engine for evaluations and lines. Use a language model for reading your own notes back to you and finding the repeated words in them — that is a text task, and it is genuinely good at text tasks.
How deep should I let the engine analyse?
Whatever your browser gives you in a few seconds per move is fine for finding blunders. Depth matters for correspondence play and for judging near-equal moves, neither of which you are doing here. You are looking for the moves that dropped 200 centipawns, and those show up at almost any depth.
Do I need a product for this?
No. A text file, Lichess's analysis board, and the discipline to write down why you played the move will get you the whole way. Tools save you the ninety minutes a week of clicking through games and keep the score for you. If you will do it by hand, do it by hand.
My blunders are all in time trouble. Is that a calculation problem?
Usually it is a decision-speed problem, not a calculation problem. Check your clock times against your error list. If most of your losses cluster under two minutes remaining, more tactics puzzles will not help much — playing a slower time control for a few weeks and watching whether the error rate moves will tell you more.
Should I work on my openings if the middlegame is where the errors are?
A little, and specifically the part of the opening that decides what middlegame you get. Learning 20 moves of theory does not help if you then have no plan on move 21. Learning the three typical pawn structures your opening produces does, because that is where 8.7% of your moves are going wrong.

See this in your own games

Mated reads your last games, scores them across five categories and builds a fifteen-minute session out of what it finds. No card.