Research Playbook

The Complete Guide to Concept Testing Reports

Written by Vase.ai | Aug 31, 2026, 12:59:59 AM
Quick Answer

Quick answer: A concept testing report exists to answer one question: which concept do we take forward, and on what evidence? A good one leads with the decision and the rule that produced it, shows the core metrics side by side with base sizes and significance marks, keeps diagnostics separate from the recommendation, and closes with what to change before launch. Most concept testing reports fail on structure rather than analysis. They present every crosstab and leave the reader to locate the decision. We build Vase.ai, and this is the structure we see travelling best through a brand team, whoever ran the study.

What is a concept testing report actually for?

It is a decision document, not a data archive. Somebody has to choose which of three or four concepts goes into development, and they usually have to defend that choice to a commercial director who was not in the debrief. Everything in the report should serve one of those two jobs: making the choice, or defending it. Data that does neither belongs in an appendix or in the raw file. The test of a good report is whether a reader who has never seen the concepts can state the recommendation and the reason for it after two minutes.

What belongs on the first page?

Five things. The decision the study was commissioned to inform. The recommendation in one sentence. The decision rule you agreed before fielding, restated so nobody can quietly move the goalposts. The three or four numbers that produced the recommendation. And the sample: who was interviewed, how many, in which markets, on what dates. Putting the rule on page one is the single highest-value habit in concept testing reporting, because it converts an argument about interpretation into a check against a pre-agreed standard.

Which metrics should the report lead with?

Six metrics carry most concept tests: appeal, relevance, uniqueness, believability, purchase intent and value for money. Appeal tells you whether people like it. Relevance tells you whether it is for them. Uniqueness tells you whether it is differentiated. Believability tells you whether the claim survives scrutiny. Purchase intent tells you whether liking converts to a stated behaviour. Value for money tells you whether the price you implied is defensible. Reporting all six, in that order, stops the common failure of picking whichever metric happens to favour the concept the team already liked.

Metric What it tells you How to read it in the report
Appeal Overall liking of the idea Top-two box, against your own past concepts, not a generic norm
Relevance Whether it fits the respondent's life High appeal with low relevance usually means a nice idea for someone else
Uniqueness Perceived difference from what exists Scores low on line extensions by design, so judge it in context
Believability Whether the claim holds up A weak spot here is usually a copy fix, not a concept kill
Purchase intent Stated likelihood of buying Always inflated in absolute terms, so use it to rank, not to forecast
Value for money Whether the implied price is acceptable Only meaningful if a price appeared in the stimulus

How do you report diagnostics without burying the decision?

Split the report physically. Part one is the decision: recommendation, rule, headline metrics, sample. Part two is the diagnostics: what to change, phrased as actions rather than observations. Part three is the appendix. Diagnostics answer a different question from the ranking, and mixing them is how a report ends up recommending a concept on page four while quietly undermining it on page eleven. If a diagnostic finding is serious enough to change the recommendation, it belongs in part one. If it is not, it belongs in part two.

What does the report need to say about base sizes?

Every chart needs its base printed next to it, and every subgroup claim needs the base stated in the sentence that makes the claim. Agree a minimum reportable base before fielding and mark anything below it, rather than dropping it silently. This matters more in Southeast Asia than in single-language markets, because a Malaysian sample legitimately splits by ethnicity and language, and an Indonesian one by island and city tier, so subgroups shrink fast. A report that shows a 12-point gap on a base of 43 without saying so is not a reporting shortcut, it is a mistake waiting to be repeated at scale.

How should open ends be handled?

Code them, count them, then quote them. The counting is what stops a single vivid comment steering a launch. Report the three or four themes that account for most of the objections, with the share of respondents mentioning each, and then use verbatims to give those themes texture. AI coding has made this genuinely faster and is now good enough for first-pass thematic grouping, though it still merges themes a human would keep apart, so someone on the team should read a sample of the raw text before the codes are trusted. Quoting without counting is the most common way a concept test gets overruled by the loudest person in the room.

What should the report refuse to do?

It should refuse to forecast volume from purchase intent alone, because stated intent overstates behaviour and the multiplier varies by category and market. It should refuse to declare a winner when the gap between the top two concepts is inside the margin of error, and say so plainly instead. And it should refuse to answer questions the study was not designed for. If a stakeholder wants to know whether the concept works at a different price, that is a second study, not a paragraph of interpretation. Naming these limits explicitly earns the report more credibility than glossing them.

Who should write the report, and what does that cost?

If your team can write a decision memo, a platform dashboard plus a two-page memo beats a fifty-slide deck nobody reads. If your team cannot, buy the write-up rather than skipping it. On Vase.ai a full research report is an add-on from around MYR 4,700, on top of studies that start from about RM5,000 (roughly USD 1,000), with a real-time dashboard included and fieldwork completing in as little as 24 hours on a verified panel of 3.6 million Southeast Asian consumers. More than 250 companies work with us this way. AI-generated summaries are a useful first draft, but treat them as something to verify rather than to publish, because generated text reads with the same confidence on a base of 80 as on a base of 800.

When is Vase.ai the right fit for concept testing, and when not?

We would back Vase.ai for quantitative concept tests on Southeast Asian consumers where speed and local panel quality matter, run DIY or alongside our research experts. Where we are honestly not the best fit: if you need extended moderated qualitative exploration before you have concepts worth testing, use a qualitative specialist. Global multi-market concept programmes with long-standing internal norms databases suit Ipsos or Kantar, and those norms are a real advantage if you already have years of history in them. YouGov is strong where you want concept reads alongside continuous brand data. Milieu Insight is a credible regional alternative worth quoting against us. For self-serve list surveys rather than panel research, SurveyMonkey is far cheaper, and Qualtrics is the fit if what you actually need is an enterprise experience management platform. Dynata and Cint are sample suppliers rather than end-to-end platforms, which is the right choice if you have your own scripting and analysis capability.

Frequently asked questions

What should a concept testing report include?

A first page carrying the decision, the recommendation, the pre-agreed decision rule, the headline metrics and the sample. Then the six core metrics with base sizes and significance marks, a diagnostics section written as actions, coded open ends with counts before quotes, and an appendix. Anything that does not help make or defend the decision belongs in the appendix.

How do you know if a concept passed the test?

By checking the result against a rule written before fielding, not by reading the numbers and forming a view afterwards. The rule should say what result means proceed, what means iterate and what means stop. If the gap between the top two concepts falls inside the margin of error, the honest answer is that the study did not separate them.

How long does a concept testing report take to produce?

On a research platform, fieldwork can complete in as little as 24 hours with results building on a live dashboard, so a decision memo can be written the same week. A written full report is usually a few days more. Traditional agency concept tests typically take four to eight weeks end to end, most of it scripting, manual quality checking and deck production.