Back to Blog

Grade Your Own Dashboard in 2026: The IAB’s 3 Measurement Tiers, Applied in 20 Minutes

Tracie Kambies

Cofounder

5 min read

IQRush cover graphic: “Grading AI Visibility — a practical guide to evaluating AI-Visibility dashboards.” A mock dashboard shows an overall grade of B, 82 out of 100, alongside scores for three IAB tiers: exploratory 87, directional 82, decision grade 78.

If you are a CMO, marketing director, or analytics lead who has been handed an AI visibility dashboard and asked whether the numbers on it can move budget, the IAB just gave you a ruler. Almost nobody is using it. This post is the twenty-minute test you can run on your own vendor, today, without a statistician in the room.

TL;DR

  • The IAB drew a line on August 3. Its guidance, Measuring Visibility in the AI Era, sorts AI visibility data into three tiers: exploratory, directional, and decision grade. Only the third tier is fit to allocate budget against.

  • Under 50 queries is exploratory. The IAB says that volume cannot meaningfully characterize a category. If your vendor tracks 40 prompts, you own a conversation starter, not a measurement program.

  • Directional is not a soft version of decision grade. It is a different product with a different job: early signal, internal briefings, competitive awareness. It is explicitly not for budget allocation.

  • Twenty-plus vendors, twenty-plus methods. The IAB counted more than 20 companies selling AI visibility measurement, each on its own approach. Only 16% of brands track AI visibility at all.

  • Score your own tool in twenty minutes. The 12 questions below map directly onto the tier definitions. Six answers put you in a tier. The other six tell you how far you are from the next one.

We participated in the IAB working group discussions, and we have written elsewhere about what the 4 P’s mean for a 2026 measurement stack. This post is not that. We are not going to argue for the standard here. We are going to hand you the ruler and get out of the way.

What the IAB actually published on August 3

Measuring Visibility in the AI Era is guidance, not a ratified standard, and the IAB is careful about the difference. Caroline Giegerich, the IAB’s VP of AI, framed the gap plainly: consumers are increasingly discovering and considering brands inside AI platforms, but measurement frameworks have not kept pace. She has also said the thing that would keep her up at night is factual inaccuracy, which is the part of the guidance most marketers skim past.

The document does two useful things. It proposes a shared metrics hierarchy, the 4 P’s: Presence, Prominence, Portrayal, and Persuasion. And it sorts the data underneath those metrics into tiers by how much decision weight the data can carry.

The working group was not a vendor huddle. Walmart, Acxiom, Microsoft, WPP Media, eMarketer, and Tinuiti were all in it. That matters when you take this into a vendor conversation, because the tiers are buy-side language, not seller language.

The three tiers, in the language your vendor will not use

Here is the whole framework, translated.

Exploratory. Fewer than 50 queries in the measurement program. The IAB’s position is that this volume cannot meaningfully characterize a category. Useful for a hunch. Not useful for anything with a dollar attached.

Directional. Enough to answer trajectory questions. Fit for early signal detection, internal briefings, and competitive awareness. The guidance is explicit that directional data is not sufficient for budget allocation. That word, “sufficient,” is doing a lot of work, and it is the word your vendor will not repeat back to you.

Decision grade. Clears a higher bar across eight things at once: sample size, query volume, prompt-type coverage, testing cadence, reproducibility, and documented method. Fit for budget allocation, provider selection, and executive strategy.

Notice what is missing from all three definitions. Engine coverage. Pricing tier. Integration count. Dashboard design. Those are table stakes in 2026, common to every vendor in the category, and the IAB did not use a single one of them to define a tier. Neither should you.

The coin flip that explains why the tiers exist

Think of flipping a fair coin 50 times. Sometimes you’ll get 23 heads. Sometimes 27. Sometimes 30. Same coin. Just different runs. The variation isn’t because the coin changed, it’s because 50 flips is a small sample, and small samples bounce around a lot. AI visibility measurement at 50 prompts works the same way.

That is the entire argument for tiering, and it is why the IAB drew its floor at 50. We have watched the same brand sit inside an answer at 9:00 AM and vanish from the same prompt by 9:05. Sometimes the difference between two runs is cosmetic. Often it is not. The engine did not change its mind about your brand. You just took two draws from a distribution that was always wider than the dashboard admitted.

Once you see the tiers as a statement about how wide that distribution is, the twelve questions below stop feeling like a compliance checklist and start feeling like the only questions worth asking.

The 12-question scorecard

Score each one Yes or No. No partial credit, and no credit for a roadmap promise. Demo floor only: if the vendor cannot show it to you in the product, it does not count.

#

Question to ask your vendor

Tier it tests

1

How many distinct prompts are in my measurement program right now?

Exploratory floor

2

How many times is each prompt run inside one measurement window?

Directional floor

3

Show me the plausible true range around any number on this screen.

Decision grade

4

Does every rank position carry its own range, separate from the share range?

Decision grade

5

Does the product refuse to compute a week-over-week delta before the baseline has settled?

Decision grade

6

What prompt types are covered, and which buyer intents are missing?

Directional to decision grade

7

How often does the program re-run, and on what schedule?

Directional to decision grade

8

If you re-ran my program tomorrow, what would change, and by how much?

Decision grade

9

Where is the written method document, and may I keep a copy?

Decision grade

10

How does the math handle several citations landing in one answer?

Decision grade

11

How many prompts would I need to detect a five-point move in my category?

Decision grade

12

Which domains in my report are frequent enough to be real, and which are drifting noise?

Decision grade

Questions 1 and 2 are the gate. Everything from 3 down assumes you cleared it.

How to score it

Nine or more Yes, including 3, 5, 8, and 9. Decision grade. Move budget on it. Ask for the method document in writing anyway.

Five to eight Yes. Directional. Genuinely useful for a Monday standup and a competitive read. Take it out of the board deck, and take it out of the reallocation conversation.

Four or fewer, or a No on question 1 or 2. Exploratory. Whatever else the product does well, the number on the screen is not carrying the weight you are putting on it.

Most tools we have audited land in the five-to-eight band. That is not an insult. Directional measurement is a real product with a real job, and a lot of the category builds it well. The failure is the labeling, not the tier: directional data sold with decision-grade confidence, wrapped in a case study, and put in front of a CMO.

I was slow to see this myself. For most of 2025, I thought the category’s problem was engine coverage, and my co-founder Todd and I said so more than once. Coverage was never the constraint. More prompts alone will not fix it either, and neither will faster dashboards, broader engine lists, or overlaying agentic workflows or recommendations on top of a number that was never stable or accurate to begin with.

Marketer-Side Impact

Take a mid-size content program at roughly $750 per briefed piece. A single quarter’s reallocation off a directional reading, say 60 pieces pointed at the wrong query cluster, is about $45,000 of production spend chasing a number that could not tell movement from noise. That is before agency rework, before the credibility cost of walking a number back in a QBR, and before the client d/sat that follows when the second run does not reproduce the first. Across the category, we estimate the industry burns roughly $2 billion a year acting on AI visibility numbers that do not hold up. The IAB’s own count says only 16% of brands are tracking this at all, which means most of that waste is still ahead of us, not behind us.

Three questions for your next vendor call

  • “Which IAB tier does this dashboard sit in, and which specific criteria put it there?” A vendor who has read the guidance will answer in about fifteen seconds. A vendor who has not will pivot to engine coverage.

  • “If you re-ran my exact program tomorrow with no changes on my side, how much would this number move?” There is a correct answer, and it is a range, not a reassurance.

  • “Can I have the method document, in writing, before we sign?” Reproducibility and documented method are two of the eight decision-grade criteria. A vendor that cannot produce a document has already answered question 9 for you.

Final thought... We are taking bets that by the end of 2027, tier language will appear as a scored line item in AI visibility vendor selection RFPs, and the vendors who cannot answer question 9 in writing will stop making shortlists. If that has not happened by December 2027, this post was wrong. What do you think?

Frequently asked questions

Is the IAB framework a standard I have to comply with?

No. It is guidance, published August 3, 2026, and the IAB is deliberate about that distinction. Nobody audits you against it. Its value is as shared vocabulary: it lets you and your vendor argue about the same thing.

My vendor tracks 200 prompts. Does that make them decision grade?

No. Prompt count clears the exploratory floor and nothing else. Sample size is one of eight decision-grade criteria. A program with 200 prompts, run once, with no ranges on screen and no written method, is a large directional program.

What if my vendor scores directional? Do I have to fire them?

No, and we would not advise it. Directional data does a real job. Change what you use it for. Keep it for competitive awareness and internal briefings, and stop putting it in the slide where you justify a budget shift.

Can I run this scorecard on a tool we already bought?

Yes, and that is the better use of it. Renewal conversations are where this scorecard earns its keep, because you can score the tool against its own last twelve months of reporting rather than against a demo.

How is this different from just asking for a bigger sample?

Sample size fixes one of the eight criteria. It does nothing for reproducibility, documented method, prompt-type coverage, or whether the product will stop you from reading a delta that has not settled yet. You can have a very large program that is still not decision grade.

Back to Blog

Grade Your Own Dashboard in 2026: The IAB’s 3 Measurement Tiers, Applied in 20 Minutes

Tracie Kambies

Cofounder

5 min read

IQRush cover graphic: “Grading AI Visibility — a practical guide to evaluating AI-Visibility dashboards.” A mock dashboard shows an overall grade of B, 82 out of 100, alongside scores for three IAB tiers: exploratory 87, directional 82, decision grade 78.

If you are a CMO, marketing director, or analytics lead who has been handed an AI visibility dashboard and asked whether the numbers on it can move budget, the IAB just gave you a ruler. Almost nobody is using it. This post is the twenty-minute test you can run on your own vendor, today, without a statistician in the room.

TL;DR

  • The IAB drew a line on August 3. Its guidance, Measuring Visibility in the AI Era, sorts AI visibility data into three tiers: exploratory, directional, and decision grade. Only the third tier is fit to allocate budget against.

  • Under 50 queries is exploratory. The IAB says that volume cannot meaningfully characterize a category. If your vendor tracks 40 prompts, you own a conversation starter, not a measurement program.

  • Directional is not a soft version of decision grade. It is a different product with a different job: early signal, internal briefings, competitive awareness. It is explicitly not for budget allocation.

  • Twenty-plus vendors, twenty-plus methods. The IAB counted more than 20 companies selling AI visibility measurement, each on its own approach. Only 16% of brands track AI visibility at all.

  • Score your own tool in twenty minutes. The 12 questions below map directly onto the tier definitions. Six answers put you in a tier. The other six tell you how far you are from the next one.

We participated in the IAB working group discussions, and we have written elsewhere about what the 4 P’s mean for a 2026 measurement stack. This post is not that. We are not going to argue for the standard here. We are going to hand you the ruler and get out of the way.

What the IAB actually published on August 3

Measuring Visibility in the AI Era is guidance, not a ratified standard, and the IAB is careful about the difference. Caroline Giegerich, the IAB’s VP of AI, framed the gap plainly: consumers are increasingly discovering and considering brands inside AI platforms, but measurement frameworks have not kept pace. She has also said the thing that would keep her up at night is factual inaccuracy, which is the part of the guidance most marketers skim past.

The document does two useful things. It proposes a shared metrics hierarchy, the 4 P’s: Presence, Prominence, Portrayal, and Persuasion. And it sorts the data underneath those metrics into tiers by how much decision weight the data can carry.

The working group was not a vendor huddle. Walmart, Acxiom, Microsoft, WPP Media, eMarketer, and Tinuiti were all in it. That matters when you take this into a vendor conversation, because the tiers are buy-side language, not seller language.

The three tiers, in the language your vendor will not use

Here is the whole framework, translated.

Exploratory. Fewer than 50 queries in the measurement program. The IAB’s position is that this volume cannot meaningfully characterize a category. Useful for a hunch. Not useful for anything with a dollar attached.

Directional. Enough to answer trajectory questions. Fit for early signal detection, internal briefings, and competitive awareness. The guidance is explicit that directional data is not sufficient for budget allocation. That word, “sufficient,” is doing a lot of work, and it is the word your vendor will not repeat back to you.

Decision grade. Clears a higher bar across eight things at once: sample size, query volume, prompt-type coverage, testing cadence, reproducibility, and documented method. Fit for budget allocation, provider selection, and executive strategy.

Notice what is missing from all three definitions. Engine coverage. Pricing tier. Integration count. Dashboard design. Those are table stakes in 2026, common to every vendor in the category, and the IAB did not use a single one of them to define a tier. Neither should you.

The coin flip that explains why the tiers exist

Think of flipping a fair coin 50 times. Sometimes you’ll get 23 heads. Sometimes 27. Sometimes 30. Same coin. Just different runs. The variation isn’t because the coin changed, it’s because 50 flips is a small sample, and small samples bounce around a lot. AI visibility measurement at 50 prompts works the same way.

That is the entire argument for tiering, and it is why the IAB drew its floor at 50. We have watched the same brand sit inside an answer at 9:00 AM and vanish from the same prompt by 9:05. Sometimes the difference between two runs is cosmetic. Often it is not. The engine did not change its mind about your brand. You just took two draws from a distribution that was always wider than the dashboard admitted.

Once you see the tiers as a statement about how wide that distribution is, the twelve questions below stop feeling like a compliance checklist and start feeling like the only questions worth asking.

The 12-question scorecard

Score each one Yes or No. No partial credit, and no credit for a roadmap promise. Demo floor only: if the vendor cannot show it to you in the product, it does not count.

#

Question to ask your vendor

Tier it tests

1

How many distinct prompts are in my measurement program right now?

Exploratory floor

2

How many times is each prompt run inside one measurement window?

Directional floor

3

Show me the plausible true range around any number on this screen.

Decision grade

4

Does every rank position carry its own range, separate from the share range?

Decision grade

5

Does the product refuse to compute a week-over-week delta before the baseline has settled?

Decision grade

6

What prompt types are covered, and which buyer intents are missing?

Directional to decision grade

7

How often does the program re-run, and on what schedule?

Directional to decision grade

8

If you re-ran my program tomorrow, what would change, and by how much?

Decision grade

9

Where is the written method document, and may I keep a copy?

Decision grade

10

How does the math handle several citations landing in one answer?

Decision grade

11

How many prompts would I need to detect a five-point move in my category?

Decision grade

12

Which domains in my report are frequent enough to be real, and which are drifting noise?

Decision grade

Questions 1 and 2 are the gate. Everything from 3 down assumes you cleared it.

How to score it

Nine or more Yes, including 3, 5, 8, and 9. Decision grade. Move budget on it. Ask for the method document in writing anyway.

Five to eight Yes. Directional. Genuinely useful for a Monday standup and a competitive read. Take it out of the board deck, and take it out of the reallocation conversation.

Four or fewer, or a No on question 1 or 2. Exploratory. Whatever else the product does well, the number on the screen is not carrying the weight you are putting on it.

Most tools we have audited land in the five-to-eight band. That is not an insult. Directional measurement is a real product with a real job, and a lot of the category builds it well. The failure is the labeling, not the tier: directional data sold with decision-grade confidence, wrapped in a case study, and put in front of a CMO.

I was slow to see this myself. For most of 2025, I thought the category’s problem was engine coverage, and my co-founder Todd and I said so more than once. Coverage was never the constraint. More prompts alone will not fix it either, and neither will faster dashboards, broader engine lists, or overlaying agentic workflows or recommendations on top of a number that was never stable or accurate to begin with.

Marketer-Side Impact

Take a mid-size content program at roughly $750 per briefed piece. A single quarter’s reallocation off a directional reading, say 60 pieces pointed at the wrong query cluster, is about $45,000 of production spend chasing a number that could not tell movement from noise. That is before agency rework, before the credibility cost of walking a number back in a QBR, and before the client d/sat that follows when the second run does not reproduce the first. Across the category, we estimate the industry burns roughly $2 billion a year acting on AI visibility numbers that do not hold up. The IAB’s own count says only 16% of brands are tracking this at all, which means most of that waste is still ahead of us, not behind us.

Three questions for your next vendor call

  • “Which IAB tier does this dashboard sit in, and which specific criteria put it there?” A vendor who has read the guidance will answer in about fifteen seconds. A vendor who has not will pivot to engine coverage.

  • “If you re-ran my exact program tomorrow with no changes on my side, how much would this number move?” There is a correct answer, and it is a range, not a reassurance.

  • “Can I have the method document, in writing, before we sign?” Reproducibility and documented method are two of the eight decision-grade criteria. A vendor that cannot produce a document has already answered question 9 for you.

Final thought... We are taking bets that by the end of 2027, tier language will appear as a scored line item in AI visibility vendor selection RFPs, and the vendors who cannot answer question 9 in writing will stop making shortlists. If that has not happened by December 2027, this post was wrong. What do you think?

Frequently asked questions

Is the IAB framework a standard I have to comply with?

No. It is guidance, published August 3, 2026, and the IAB is deliberate about that distinction. Nobody audits you against it. Its value is as shared vocabulary: it lets you and your vendor argue about the same thing.

My vendor tracks 200 prompts. Does that make them decision grade?

No. Prompt count clears the exploratory floor and nothing else. Sample size is one of eight decision-grade criteria. A program with 200 prompts, run once, with no ranges on screen and no written method, is a large directional program.

What if my vendor scores directional? Do I have to fire them?

No, and we would not advise it. Directional data does a real job. Change what you use it for. Keep it for competitive awareness and internal briefings, and stop putting it in the slide where you justify a budget shift.

Can I run this scorecard on a tool we already bought?

Yes, and that is the better use of it. Renewal conversations are where this scorecard earns its keep, because you can score the tool against its own last twelve months of reporting rather than against a demo.

How is this different from just asking for a bigger sample?

Sample size fixes one of the eight criteria. It does nothing for reproducibility, documented method, prompt-type coverage, or whether the product will stop you from reading a delta that has not settled yet. You can have a very large program that is still not decision grade.

AI search visibility you can defend

Whether you're building, buying, or briefing on AI search, get decision-grade data that holds.

Book a demo

© 2026 IQRush. All Rights Reserved.

Site by ONBOX

AI search visibility you can defend

Whether you're building, buying, or briefing on AI search, get decision-grade data that holds.

Book a demo

© 2026 IQRush. All Rights Reserved.

Site by ONBOX

AI search visibility you can defend

Whether you're building, buying, or briefing on AI search, get decision-grade data that holds.

Book a demo

© 2026 IQRush. All Rights Reserved.

Site by ONBOX