Back to Blog
What Keyword Volume Had That AI Prompt Volume Doesn't: A Number You Could Check

Jack Westbrock
Senior Agentic Engineer
5 min read

Keyword volume ran content planning for twenty years and earned the job. A marketer looks up how many people searched a phrase last month, sorts the calendar by it, and defends the plan with it in a review. Most teams bought that number from Ahrefs or Semrush or pulled it from Keyword Planner, and it held up well enough to build a discipline on.
Two things made it work. Keywords repeat, so when a few thousand people type the same phrase in the same order, counting them is a well-defined operation on one string. And the number could be checked against something the marketer already owned: Search Console impressions, ad spend, click data. Most practitioners understood the figure was modeled, and the modeling was tolerable because the model could be caught when it drifted.
Prompt volume carries the same promise into answer engines, and it runs into a harder measurement problem.
Most prompts aren't repeated word for word. Some people do type “best crm software” into ChatGPT, but plenty type three sentences about their team size, their budget, and the tool they're trying to leave, and the next person phrases the same need differently again. So counting is no longer a tally of one string. Someone has to decide which differently worded questions belong together, group them into a theme, and scale that group up from whatever population they were able to observe to the market a brand actually sells into. Two modeling steps where search had roughly one.
And the check is mostly gone. Some first-party reporting is starting to appear: Google now shows impressions from AI features inside Search Console, and Perplexity pays publishers in its program based on how often their content gets cited. Neither is a count of how often a question was asked. So a marketer holding a prompt volume figure still has no first-party demand number to hold it against.
Which puts a lot of weight on one question: how far off can the number be?
TL;DR
A prompt volume that reads 1,000 might really be 1,050. It might also really be 500. The engines don't publish what people ask them, so every one of these figures is an estimate. What a planner needs to know is how far off it can be.
AthenaHQ calls its prompt volume “95%+ accurate” without saying accurate to within what, or measured against which data. No dataset, no panel, and no provider is named on any of their public pages.
Their own model page lists “Volume estimates with confidence intervals” as something the product outputs. We checked six of their pages on July 31 and found no interval on any of them, including on the 95% itself.
The dollar value they attach to a prompt is four estimates multiplied together. A prompt shown at $10,000 could sit anywhere between roughly $4,100 and $20,700, and not one of the four inputs is published with a margin of error.
Their public credit calculator prices usage by prompts, locations, models, and cadence. Nothing in it accounts for repeat readings of the same prompt, which tells you each figure on the screen is one reading rather than a distribution.
The specific case
AthenaHQ's prompt volume product runs on what they call the Query Volume Estimation Model. Their materials describe it as “95%+ Accuracy Rate: Validated against real-world data from multiple AI platforms,” and the FAQ repeats it as “95%+ accuracy rate when validated against actual platform data.” Forecasting is claimed separately at “85% accuracy for established query patterns and 70% for emerging topics” six months out.
They deserve credit for publishing more than most of this category does. The value formula is out in the open, the data sources are described in a table, and in May they published a validation study for a different model of theirs.
On this model, two things a planner would need are missing.
A tolerance. Take an estimate of 1,000 uses a month. If it is 95% accurate, does the true figure sit between 950 and 1,050, or between 500 and 2,000? Both get called accurate in ordinary business language, and the difference decides whether two topics can be ranked against each other or only sorted into rough tiers. The public materials don't say which.
A ground truth. “Actual platform data” is the most specific description offered. The underlying data is grouped into three buckets, of which the most concrete is “Premium data providers, industry partnerships.” No provider, dataset, panel, sample size, or holdout is named.
There is also a line worth reading closely. Under Output Generation, their model page lists “Volume estimates with confidence intervals” as something the model produces. We looked for one on the model article, the prompt volume product page, the plans page, the credit calculator, the estimation article, and the content scoring study. On those six pages, checked July 31, no interval appears, including none attached to the 95%, the 85%, or the 70%.
NOISE REPORTED AS SIGNAL — What the materials say: prompt volume at 95%+ accuracy, with confidence intervals listed as a model output. What the published record shows: no stated tolerance, no named ground truth, and no interval on the six pages checked July 31. A figure that arrives without a range tends to get planned against as though its error were zero. That is the difference between an estimate and a measurement.
What lands on the screen
When a marketer hears prompt volume, the picture is Search Console for AI: a list of individual prompts, each with a number beside it. That picture is roughly right, and the granularity does a lot of the selling. It looks like the kind of detail a content plan can be built on.
Per AthenaHQ's own copy, here is what a buyer gets. A monthly figure per prompt with a dollar value next to it, updated “within 24-48 hours,” and no range on either number. The meter is one credit per AI response. Their public credit calculator sizes usage across “prompts, locations, models, and cadence,” with no input anywhere for repeat readings of the same prompt.
Cadence is worth pulling apart, because it looks like it covers this and it doesn't. Cadence buys the same prompt checked again tomorrow. It does not buy the same prompt sampled several times and reported with a range. So each figure on that screen is one reading rather than a distribution, and the tool is priced in a way that makes repeated readings cost more.
That is the part to sit with, because per-prompt detail invites per-prompt decisions. Keyword volume counted one string typed by thousands of people. This estimates how often a cluster of differently worded questions gets asked, from a single pass, with nothing published around it.
Four estimates wearing one number
AthenaHQ publishes the formula that converts volume into a dollar value, which is the figure most likely to reach a business case:
Prompt Value = (Search Volume × Intent Score × Competition Factor) × Bid Price Estimate
Search Volume is the model's own output. Intent Score and Competition Factor are derived scores, described as user intent strength and competition density, and neither is directly observable. Bid Price Estimate carries the word estimate in its name; an earlier AthenaHQ article describes it as “Combining volume estimates with keyword bid prices to calculate potential value,” which is a paid-search bid price carried into a channel that doesn't price its answers by keyword.
The arithmetic matters here. If one input is off by 20% and the other three are exact, the dollar figure is off by 20%. If all four drift 20% in the same direction, the product runs from 0.41 to 2.07 times the number on screen, so $10,000 sits somewhere between roughly $4,100 and $20,700. That range is the arithmetic corner, not a claim about AthenaHQ's actual error. The actual error isn't knowable from the published materials, because none of the four inputs carries a published bound.

Fig 1. Where a $10,000 prompt value can actually sit: with all four inputs drifting 20%, the number ranges from roughly $4,100 to $20,700.
That's the whole issue in one line. Not that the number is wrong, but that its error is undisclosed, and an undisclosed error tends to get treated as a small one.
The other way to build it
There is a version of this where the uncertainty is part of the product rather than an omission from it.
Every figure IQRush reports arrives with a 95% confidence interval on the reading itself, not buried in a methodology appendix, because a citation share of 43% means one thing at plus or minus 2 points and something else at plus or minus 11. Two further checks sit on top. Stability asks whether enough responses have accumulated for the rankings a marketer is about to act on to have settled. Sufficiency asks whether the signal has separated from the noise floor, which is what tells a team that collecting more data wouldn't change the decision. An accuracy percentage can't answer either question. Those two can.
We've extended that same discipline to prompt demand. General availability at the end of August.
Frequently asked questions
Is AthenaHQ's prompt volume inaccurate?
Nothing in the public record shows that, and nothing in it allows anyone to verify the accuracy claim either. What's established is narrower: no stated tolerance, no named ground truth, and no interval displayed on the pages we checked. Unverifiable and inaccurate are different findings, and only the first one is supported.
Isn't every demand number modeled anyway?
Yes, and that's the point rather than a defense. Keyword volume was modeled too. What made it usable was that a marketer could check it against their own reporting. Prompt demand removes that check, which raises the burden on the vendor to publish the error instead of asking the buyer to assume it.
What would change this assessment?
A named ground-truth dataset, a stated tolerance, a validation design with a holdout, or an interval shown on a volume estimate inside the product. Any two of those four would move it a long way.
Back to Blog
What Keyword Volume Had That AI Prompt Volume Doesn't: A Number You Could Check

Jack Westbrock
Senior Agentic Engineer
5 min read

Keyword volume ran content planning for twenty years and earned the job. A marketer looks up how many people searched a phrase last month, sorts the calendar by it, and defends the plan with it in a review. Most teams bought that number from Ahrefs or Semrush or pulled it from Keyword Planner, and it held up well enough to build a discipline on.
Two things made it work. Keywords repeat, so when a few thousand people type the same phrase in the same order, counting them is a well-defined operation on one string. And the number could be checked against something the marketer already owned: Search Console impressions, ad spend, click data. Most practitioners understood the figure was modeled, and the modeling was tolerable because the model could be caught when it drifted.
Prompt volume carries the same promise into answer engines, and it runs into a harder measurement problem.
Most prompts aren't repeated word for word. Some people do type “best crm software” into ChatGPT, but plenty type three sentences about their team size, their budget, and the tool they're trying to leave, and the next person phrases the same need differently again. So counting is no longer a tally of one string. Someone has to decide which differently worded questions belong together, group them into a theme, and scale that group up from whatever population they were able to observe to the market a brand actually sells into. Two modeling steps where search had roughly one.
And the check is mostly gone. Some first-party reporting is starting to appear: Google now shows impressions from AI features inside Search Console, and Perplexity pays publishers in its program based on how often their content gets cited. Neither is a count of how often a question was asked. So a marketer holding a prompt volume figure still has no first-party demand number to hold it against.
Which puts a lot of weight on one question: how far off can the number be?
TL;DR
A prompt volume that reads 1,000 might really be 1,050. It might also really be 500. The engines don't publish what people ask them, so every one of these figures is an estimate. What a planner needs to know is how far off it can be.
AthenaHQ calls its prompt volume “95%+ accurate” without saying accurate to within what, or measured against which data. No dataset, no panel, and no provider is named on any of their public pages.
Their own model page lists “Volume estimates with confidence intervals” as something the product outputs. We checked six of their pages on July 31 and found no interval on any of them, including on the 95% itself.
The dollar value they attach to a prompt is four estimates multiplied together. A prompt shown at $10,000 could sit anywhere between roughly $4,100 and $20,700, and not one of the four inputs is published with a margin of error.
Their public credit calculator prices usage by prompts, locations, models, and cadence. Nothing in it accounts for repeat readings of the same prompt, which tells you each figure on the screen is one reading rather than a distribution.
The specific case
AthenaHQ's prompt volume product runs on what they call the Query Volume Estimation Model. Their materials describe it as “95%+ Accuracy Rate: Validated against real-world data from multiple AI platforms,” and the FAQ repeats it as “95%+ accuracy rate when validated against actual platform data.” Forecasting is claimed separately at “85% accuracy for established query patterns and 70% for emerging topics” six months out.
They deserve credit for publishing more than most of this category does. The value formula is out in the open, the data sources are described in a table, and in May they published a validation study for a different model of theirs.
On this model, two things a planner would need are missing.
A tolerance. Take an estimate of 1,000 uses a month. If it is 95% accurate, does the true figure sit between 950 and 1,050, or between 500 and 2,000? Both get called accurate in ordinary business language, and the difference decides whether two topics can be ranked against each other or only sorted into rough tiers. The public materials don't say which.
A ground truth. “Actual platform data” is the most specific description offered. The underlying data is grouped into three buckets, of which the most concrete is “Premium data providers, industry partnerships.” No provider, dataset, panel, sample size, or holdout is named.
There is also a line worth reading closely. Under Output Generation, their model page lists “Volume estimates with confidence intervals” as something the model produces. We looked for one on the model article, the prompt volume product page, the plans page, the credit calculator, the estimation article, and the content scoring study. On those six pages, checked July 31, no interval appears, including none attached to the 95%, the 85%, or the 70%.
NOISE REPORTED AS SIGNAL — What the materials say: prompt volume at 95%+ accuracy, with confidence intervals listed as a model output. What the published record shows: no stated tolerance, no named ground truth, and no interval on the six pages checked July 31. A figure that arrives without a range tends to get planned against as though its error were zero. That is the difference between an estimate and a measurement.
What lands on the screen
When a marketer hears prompt volume, the picture is Search Console for AI: a list of individual prompts, each with a number beside it. That picture is roughly right, and the granularity does a lot of the selling. It looks like the kind of detail a content plan can be built on.
Per AthenaHQ's own copy, here is what a buyer gets. A monthly figure per prompt with a dollar value next to it, updated “within 24-48 hours,” and no range on either number. The meter is one credit per AI response. Their public credit calculator sizes usage across “prompts, locations, models, and cadence,” with no input anywhere for repeat readings of the same prompt.
Cadence is worth pulling apart, because it looks like it covers this and it doesn't. Cadence buys the same prompt checked again tomorrow. It does not buy the same prompt sampled several times and reported with a range. So each figure on that screen is one reading rather than a distribution, and the tool is priced in a way that makes repeated readings cost more.
That is the part to sit with, because per-prompt detail invites per-prompt decisions. Keyword volume counted one string typed by thousands of people. This estimates how often a cluster of differently worded questions gets asked, from a single pass, with nothing published around it.
Four estimates wearing one number
AthenaHQ publishes the formula that converts volume into a dollar value, which is the figure most likely to reach a business case:
Prompt Value = (Search Volume × Intent Score × Competition Factor) × Bid Price Estimate
Search Volume is the model's own output. Intent Score and Competition Factor are derived scores, described as user intent strength and competition density, and neither is directly observable. Bid Price Estimate carries the word estimate in its name; an earlier AthenaHQ article describes it as “Combining volume estimates with keyword bid prices to calculate potential value,” which is a paid-search bid price carried into a channel that doesn't price its answers by keyword.
The arithmetic matters here. If one input is off by 20% and the other three are exact, the dollar figure is off by 20%. If all four drift 20% in the same direction, the product runs from 0.41 to 2.07 times the number on screen, so $10,000 sits somewhere between roughly $4,100 and $20,700. That range is the arithmetic corner, not a claim about AthenaHQ's actual error. The actual error isn't knowable from the published materials, because none of the four inputs carries a published bound.

Fig 1. Where a $10,000 prompt value can actually sit: with all four inputs drifting 20%, the number ranges from roughly $4,100 to $20,700.
That's the whole issue in one line. Not that the number is wrong, but that its error is undisclosed, and an undisclosed error tends to get treated as a small one.
The other way to build it
There is a version of this where the uncertainty is part of the product rather than an omission from it.
Every figure IQRush reports arrives with a 95% confidence interval on the reading itself, not buried in a methodology appendix, because a citation share of 43% means one thing at plus or minus 2 points and something else at plus or minus 11. Two further checks sit on top. Stability asks whether enough responses have accumulated for the rankings a marketer is about to act on to have settled. Sufficiency asks whether the signal has separated from the noise floor, which is what tells a team that collecting more data wouldn't change the decision. An accuracy percentage can't answer either question. Those two can.
We've extended that same discipline to prompt demand. General availability at the end of August.
Frequently asked questions
Is AthenaHQ's prompt volume inaccurate?
Nothing in the public record shows that, and nothing in it allows anyone to verify the accuracy claim either. What's established is narrower: no stated tolerance, no named ground truth, and no interval displayed on the pages we checked. Unverifiable and inaccurate are different findings, and only the first one is supported.
Isn't every demand number modeled anyway?
Yes, and that's the point rather than a defense. Keyword volume was modeled too. What made it usable was that a marketer could check it against their own reporting. Prompt demand removes that check, which raises the burden on the vendor to publish the error instead of asking the buyer to assume it.
What would change this assessment?
A named ground-truth dataset, a stated tolerance, a validation design with a holdout, or an interval shown on a volume estimate inside the product. Any two of those four would move it a long way.
AI search visibility you can defend
Whether you're building, buying, or briefing on AI search, get decision-grade data that holds.
© 2026 IQRush. All Rights Reserved.
© 2026 IQRush. All Rights Reserved.
Site by ONBOX
AI search visibility you can defend
Whether you're building, buying, or briefing on AI search, get decision-grade data that holds.
Resources
AI search visibility you can defend
Whether you're building, buying, or briefing on AI search, get decision-grade data that holds.
© 2026 IQRush. All Rights Reserved.