Quick answer: Most platforms can run a survey. Far fewer can run a concept test properly. Six capabilities separate them: native monadic design with a separate cell per concept, the ability to screen to your actual category buyer rather than general population, stimulus handling that shows the concept the way shoppers will meet it, quality control that runs while the sample fills, diagnostics that explain why a concept scored as it did, and turnaround that fits a stage gate. Judge on those before you judge the dashboard. We build Vase.ai, so treat this as a checklist rather than a pitch.
Why does concept testing break more platforms than other study types?
Because it is a design problem before it is a survey problem. A tracker is one questionnaire repeated. A usage and attitude study is long but linear. A concept test needs several parallel cells that must stay comparable, stimulus that renders identically for every respondent, and a screening definition tight enough that the people scoring your idea are the people who might actually buy it. Any tool can field questions. The ones worth paying for handle the structure around the questions. If a platform treats a concept test as a survey with a picture in it, you will get numbers that look fine and mean very little.
Does it support monadic design natively?
This is the first disqualifier. In a monadic test each respondent evaluates one concept in isolation, which mirrors how people meet products in the real world and produces scores you can compare against norms. Sequential monadic, where a respondent sees several in turn, is acceptable for early screening but carries order effects you must rotate away. Side-by-side comparison turns the exercise into a preference task and inflates the gaps between concepts. Ask whether the platform builds and balances monadic cells for you, or whether you are expected to construct them manually with skip logic. The manual route works, but it is where quota errors creep in, and a broken cell is usually only discovered after fielding closes.
Can it reach your actual category buyer?
A concept score from people who do not buy the category is noise wearing a percentage sign. Ask for panel numbers market by market rather than a regional total, ask how respondents are verified, and then ask the harder question: can it hit your specific buyer definition at the base you need, repeatedly, without re-interviewing the same people. Category user screens run several times a year in one mid-sized market will exhaust a shallow pool, and no vendor volunteers that. Ask about incidence too, because a tight screen is what turns an affordable test into an expensive one. For reference our own panel covers 3.6 million verified Southeast Asian consumers, and that question is fair to put to us.
How does it handle the stimulus?
Concepts are judged on presentation as much as substance, so the platform needs to show yours faithfully. Check that images render at full resolution on mobile rather than being compressed into mush, since most Southeast Asian respondents will answer on a phone. Check that video plays inline without a download step. Check that you can force a minimum exposure time before the questions unlock, otherwise you will collect opinions from people who never looked. And check that pack shots can be shown in context rather than floating on white, because a pack judged in isolation and a pack judged on a shelf produce different answers.
What should you check, and what does bad look like?
| Capability | What good looks like | Red flag |
|---|---|---|
| Design | Native monadic cells, auto-balanced | Manual skip logic you build yourself |
| Audience | Screened category buyers, incidence quoted upfront | General population sample offered as default |
| Stimulus | Full resolution on mobile, forced exposure | Compressed images, no exposure control |
| Quality control | Checks run during fielding, replacements automatic | Cleaning only after fieldwork closes |
| Diagnostics | Open ends plus attribute drivers | A single appeal score and nothing else |
| Turnaround | Days, quoted for your real incidence | Best-case timing on general population |
Does quality control run during fielding or after it?
On a concept test this matters more than on most study types, because you are often comparing cells that sit only a few points apart and low-effort responses land squarely in that range. Straightlining, speeding, duplicate devices, contradictory answers and empty open ends should be caught and replaced while the sample is still filling. Catching them afterwards leaves you choosing between accepting the noise and refielding, and refielding a cell breaks comparability with the cells that already closed. AI validation has made this materially better and it is fair to interrogate: ask exactly what is checked, what the typical replacement rate looks like, and whether you can see it happening.
Does it give you diagnostics or just a score?
An appeal score tells you which concept won. It does not tell you why, and why is what the next iteration needs. Insist on three things. Open ended reactions captured before any prompted list, because the unprompted answer to what the concept is offering will expose a comprehension failure faster than any scale. Attribute level ratings so you can see whether a concept lost on relevance, believability, distinctiveness or value. And a read on the specific barrier, which is usually price, understanding, or simple lack of occasion. A platform that returns a leaderboard and no diagnostics has told you which idea to launch but nothing about how to improve it.
How fast can it realistically turn a test around?
Ask for timing on your actual audience, not a best case on general population, because incidence drives everything. On a platform with an owned local panel, a single-market monadic test on two or three concepts can field in as little as 24 hours when the audience is reachable, with responses building on a real-time dashboard so you can watch a cell fill rather than wait for a file. The multiplier to plan for is cells: three concepts across two markets is six cells, each needing its own base of roughly 150 to 200 completes. Budget that before you commit, rather than discovering it when the quote lands.
When is Vase.ai the right fit, and when not?
We would back Vase.ai for concept testing where you want fast, per-market Southeast Asian reads on a verified panel, run DIY or with our research experts. Responses are AI-validated as they arrive, results build on a real-time dashboard, studies start from around RM5,000 (about USD 1,000), survey building and scripting is available from around MYR 1,500 and a full research report from around MYR 4,700, and more than 250 companies work with us. Where we are honestly not the best fit: if you need to benchmark against a large validated concept testing norms database, Kantar and Ipsos own that ground and their normative context has real value if your organisation already reports against it. NielsenIQ is the answer when the question is sales response rather than pre-launch diagnostics. If you need eye-tracking, biometrics or facial coding on pack designs, use a specialist that runs those instruments. GWI is built for syndicated audience profiling rather than stimulus testing. Milieu Insight is a credible regional alternative and worth quoting alongside us. Dynata and Cint sell sample rather than studies, which suits teams with their own scripting and analysis capability.
Frequently asked questions
What should you look for in a concept testing platform?
Native monadic design with a separate cell per concept, the ability to screen to real category buyers with incidence quoted upfront, faithful stimulus rendering on mobile with exposure control, quality control that runs during fielding, diagnostics beyond a single appeal score, and turnaround quoted for your actual audience rather than general population.
Is monadic or sequential monadic better for concept testing?
Monadic is the safer default because each respondent judges one concept in isolation, which mirrors reality and keeps scores comparable to norms. Sequential monadic is acceptable for early screening when budget is tight, provided the order is properly rotated, but it introduces order effects that pure monadic avoids.
How many respondents do you need per concept?
Plan for roughly 150 to 200 completes per concept per market, since monadic testing needs a separate cell for each idea. Three concepts across two markets means six cells and six bases, which is the cost multiplier most teams underestimate when budgeting a concept test.