The keyword tool still works. That is the confusing part. You open it, you get terms, volumes, difficulty scores, a clean export you can sort and hand to a writer. Nothing about the interface has broken.
What has quietly stopped working is the inference you were drawing from it. High volume plus achievable difficulty used to be a reasonable forecast of attention. Against answer engines, the same numbers predict very little, and it took a while for anyone to notice because the spreadsheet kept looking healthy.
The number was always a proxy for something else
Volume never measured demand. It measured how often people typed a particular string into a particular box.
That was a good enough stand-in while the box was the only entrance, and the strings were stable because everyone had been trained by a decade of typing short phrases to get useful results. People compressed their situation into three words because three words worked.
Take the training wheels off and the compression stops. Someone asking an assistant writes four lines: their team size, the tool they already run, the thing that went wrong last time, and only then the question. That request represents exactly the same demand it always did. It contributes nothing to any volume figure anywhere.
So the proxy did not become inaccurate. The behaviour it was standing in for moved somewhere the proxy cannot see.
A prompt does not decompose into terms
The obvious workaround is to extract the keywords from the prompt and carry on as before. It does not survive contact with the data.
A prompt carries a shape as well as a subject. Compare these two. Rank these for my situation. Explain this to someone who has never heard of it. Tell me what I am missing. The subject might be identical across all four and the answer you would need to be included in is completely different each time.
It also carries the constraint, which is usually the part that decides the outcome. “For a small team without a data engineer” is not a modifier on a keyword. It is the whole question. Anyone tracking how AI is changing search discovery ends up in roughly the same place, which is that the unit worth researching got bigger and less countable at the same time.
Twenty phrasings of the same situation resolve to one answer. One phrasing with a different constraint resolves to a completely different one. Neither of those relationships is visible in a term list.
Ranking and being used are separate events
Here is the part that breaks the forecast most cleanly. Position one for a term does not reliably mean inclusion in the answer built around that term.
An answer engine gathers material, reconciles it, and writes. What it reaches for skews toward whatever is easiest to use, and how the sources AI trusts most are distributed across the web looks nothing like a ranking table. Community threads, documentation, independent comparisons, review platforms, the occasional trade publication.
You watch it happen in reverse too. A page that ranks nowhere in particular gets quoted repeatedly because one paragraph in it answers a specific question outright. The page has no keyword story at all. It just happens to contain the exact sentence the answer needed.
Once both of those are true at once, the correlation between rank forecast and visibility forecast is loose enough that you cannot plan on it.
Difficulty scores measure the wrong contest
Difficulty estimates your odds against the pages currently holding a position. It assumes the contest is a queue and you are trying to move up it.
The contest in an answer is not a queue. It is a selection of material, and the constraint is not how strong your competitors’ backlink profiles are. It is whether your description of the topic is clear, consistent with what else exists, and specific enough to be worth including in a sentence.
Which produces a strange result that most teams hit within a month of measuring properly. You find high-difficulty topics where you get included easily because nobody has written a clean explanation. And you find low-difficulty topics where you are invisible because your own descriptions of yourself contradict each other across four different pages.
None of that shows up on a difficulty score, because a difficulty score was never trying to tell you.
What keyword data still earns its place for
Throwing the tool away would be an overcorrection, and the teams who did it in the first wave mostly quietly reopened it.
- It still describes a real population of people who type short phrases, and that population has not disappeared
- It is still the fastest way to see which topics exist in a category you do not know well
- It still shows you competitive movement over time, which is a decent early warning system
- It is still useful for the classic informational pages that get retrieved constantly precisely because they are clear and well organised
What it no longer does is predict. Treat it as a map of one territory rather than a forecast of a different one, and it stops being misleading.
The replacement for prediction is not a better score. It is running the actual questions your buyers ask, reading what comes back, and doing it often enough to tell a pattern from a bad week. Slower, less tidy, considerably closer to what is happening.

