Years ago I stood in a research facility while a baby formula was customised down to a single variant, for infants in a neonatal unit whose bodies could not tolerate the standard product. It is still my favourite thing about that job. You do not arrive at that variant from a dashboard. Somebody went and looked at real babies.
I thought about it this week for a strange reason.
A company called Simile raised a 200 million dollar Series B, five months after launching, to predict human decisions at scale. Built by Stanford researchers on their published work, it is a foundation model trained on real people and pointed at the decisions of all eight billion of us, with a second model that scores its own confidence in every answer (Simile, via Superhuman, August 3 2026).
Take the funding headline off and look at what is being sold. Most of the digital-human wave so far has been a synthetic presenter: an avatar to talk at you. This is the other side of the glass. A synthetic respondent. The survey panel, the concept screen, the claims test, the pack test. That whole stack is being rebuilt as a simulator.
If you own a P&L in consumer goods, this lands on the softest line in your budget. Research is slow. It is expensive. It delays launches, and it is the line every finance director already half-suspects of being insurance rather than insight. A tool that costs a fraction of a panel and answers in an afternoon will not need to be sold to your CFO. It will need to be defended against by your insights team.
So the useful question is not whether it works. It will work well enough, on some things, sooner than the insights function expects. The question is which decisions a simulated customer is allowed to make.
Here is the frame I keep returning to. These models are trained on what has already happened. That is how they work. They are very good at harnessing the past for the future. Point one at a population and ask what people like this have already chosen, and you get a fast and cheap read that is directionally sound. Ask it about something nobody has ever tasted, smelled, held, or been embarrassed to put in a trolley, and you are asking history to vote on the future.
That is the line, and it is drawable today.
On one side: screening fifty concepts down to eight, sizing a segment roughly, triaging pack variants, killing the obviously bad ideas before you spend panel money on them. Simulation does that work quickly and cheaply, and a research budget still paying full freight for it is wasting money that could fund the work that matters.
On the other side: anything where a real human reaction is the ground truth. Claims substantiation. Regulatory. Cultural nuance in a market where you are the outsider. The reason people actually stopped buying. No regulator has ever accepted a simulated respondent, and no shopper has ever been persuaded by one.
Two things make this harder than a demo suggests.
The first is representativeness, and it is the same argument I have been having about training data for years. One in three babies born in the United States has at least one parent of Hispanic heritage. If that is not true of the data underneath your simulator, your simulator does not know your consumer, and it will still answer you with total fluency.
The second is subtler. A confidence score tells you how sure the model is. It does not tell you whether the model is right. Those two come apart most in exactly the situations you bought the thing for: a new category, a new behaviour, the first genuinely novel product in a decade. You get a confident number about the thing you understand least.
And the failure does not announce itself. You optimise to a model of your customer, the model is built out of your customer's past, and everything looks fine until a concept that tested beautifully does not sell. By then the research line has been cut, so there is nothing left to check the model against.
One concrete thing worth doing this quarter. When a vendor gets in front of you, do not let them run your next question. Hand them three studies you have already paid for, where you know the answer and they do not, and have them predict the results. Then look at how the confidence scores behaved on the two they got right and the one they missed. That is a cheap test that tells you more than a six-month pilot, and it gives your insights team a number to argue with instead of a hunch.
The variant for those babies was never going to come out of a model of the past. Simulate the screening. Go and look for the rest.
-- Imteaz
KEEP READING
If this one landed, earlier pieces on the same thread:
FORWARD IT
This one is for whoever owns the insights line and is about to be asked why it costs what it costs. If someone forwarded this to you, subscribe here.
Or just hit reply and tell me one thing: which research decision in your business would you genuinely hand to a simulated respondent, and which one would you never? I read every one.
Your prompts are leaving out 80% of what you're thinking.
When you type a prompt, you summarize. When you speak one, you explain. Wispr Flow captures your full reasoning — constraints, edge cases, examples, tone — and turns it into clean, structured text you paste into ChatGPT, Claude, or any AI tool. The difference shows up immediately. More context in, fewer follow-ups out.
89% of messages sent with zero edits. Used by teams at OpenAI, Vercel, and Clay. Try Wispr Flow free — works on Mac, Windows, and iPhone.




