The most famous quality failure in consumer research history is also Q1’s clearest illustration of the gap. The New Coke decision tested cleanly in quantitative aggregation, with net preference breaking favorably and focus-group dissent treated as statistical noise. It shipped into one of the great category retreats of the century. The methodology was procedurally correct; the interpretation was catastrophically wrong. Procedural correctness, the case shows, is the failure mode the field has been mistaking for the solution. A food marketing academic surfaced the case in Q1 to make a specific argument: pre-register hypotheses, split datasets, treat qualitative dissent as signal rather than noise.8 The prescription is a guardianship argument applied to study design, and it lands four decades later because the gap has only widened. Today’s version of the New Coke decision is built faster, costs less, and breaks the same way.
The synthetic data debate reached a definitional breaking point in Q1. The founder of a synthetic-data platform spent much of his episode less advocating for his category than distinguishing it from what most of the market means when it uses the term.1 The conflation has consequences. “Synthetic data” now covers both causal AI that models relationships within genuine survey datasets and LLM-generated respondents that interpolate from training corpora, two capabilities with radically different epistemological legitimacy, evaluated in procurement as if they were the same thing.
The epistemological problem with LLM-generated synthetic respondents is specific and structural. These systems produce outputs that reflect the distribution of text that existed before the survey was designed, so they can’t surface genuinely novel consumer positions. And they fail in ways that are difficult to detect from the output alone: the same question posed in English versus Spanish activates different parts of the network and returns materially different answers, with no signal in the result that anything has gone wrong.1 The synthetic-data founder drew a clarifying contrast to coding, the domain where LLMs have had their most commercially visible success. Wrong outputs in software fail visibly, get tested, and get retried. Consumer research has no equivalent sandbox. The output looks plausible whether it’s accurate or not, because accuracy can’t be checked against a ground truth the research was designed to discover.1 An independent articulation of this failure mode appeared in Q4 2025 under the label “research theater,” the production of outputs that perform the function of research without its epistemological content. Q1 has now located that failure at its methodological root.
The affirmative case for legitimate synthetic data is narrower and more defensible: causal scenario modeling applied to existing real-respondent data to extend analysis into subgroups the original sample was underpowered to reach. The respondents are real; the extended analysis is synthetic.1 The distinction matters because it preserves the human signal at the base of the inference chain, the thing LLM-generated respondents lack by construction.
The underpowered study is the industry’s most persistent and least discussed quality failure. The same synthetic-data platform founder cited a striking figure: more than 50% of surveys, he claims, return no statistically valid results. If that’s accurate, the research industry’s quality problem lies primarily in study design and powering, not in analysis or visualization.1 This claim is flagged as unverified (see Notes & Methodology), but its directional force is consistent with what the quarter’s other methodological voices observed from different angles: studies are routinely commissioned that can’t, in principle, return actionable answers.
The “just-add-sugar” problem names the failure mode that follows. Regression finds that sweetness predicts preference, and it does. But the researcher is supposed to know that a product already at its formulation limit, or whose brand positioning requires restraint, can’t act on that correlation. The regression can’t make that judgment. When the researcher delegates interpretation to the algorithm, the category error becomes invisible.1
The CEO of a leading panel quality firm extended the epistemological crisis to sampling, the layer of the research stack least visible to clients and most vulnerable to degradation.6 Bot responses, professional survey farmers, and AI-assisted identity fabrication represent a systemic threat, not a nuisance. And the tools enabling faster research synthesis simultaneously lower the barrier to fabricating convincing-looking respondents. The industry is in an arms race it isn’t winning.6 Clients who treat sampling as a procurement decision, awarded to whoever delivers fastest at lowest cost, create the conditions that fraud exploits.
Where this leaves us is a line procurement standards will have to start drawing in 2026. The legitimate tier models on real data and instruments AI to read the real signal more cleanly; the illegitimate tier inflates a small sample to half a million synthetic rows when the original survey already failed, and adds nothing the original couldn’t deliver. The work is upstream of that. It takes a practiced human to tell apart a result that’s true from one that merely looks plausible, and that skill is becoming the scarce, appreciating asset, not the depreciating one.