Synthesis·

How Synthesis Turns 5,000 Responses Into Actionable Insights

Reading 5,000 responses would take 12 hours. Sampling 100 introduces bias. Here's what synthesis actually does instead.

A

AskEveryone

Thousands of scattered response nodes with glowing threads connecting similar ones into clusters

Suppose you ask your audience one question and 3,000 people respond.

You could read them all. It would take about twelve hours, assuming you give each response thirty seconds — which isn't enough time to really absorb a thoughtful paragraph. You'd have a sense of the general shape of things, but you'd inevitably over-weight the responses you read last, or the ones that were unusually well-articulated, or the ones that confirmed what you already suspected.

Alternatively, you could read a random hundred and hope they're representative. Which they might not be — the patterns that matter most in qualitative data are often in the quiet majority, not the top of the list.

Neither option is good. For most of the history of audience research, this was a real constraint: qualitative data at scale was effectively unusable without a research team. That's what changed.


The scale problem in plain terms

Qualitative data — open-ended responses in people's own words — is rich in ways that numbers aren't. It contains the language people actually use, the emotional texture of their experience, the connections they make that a preset option could never capture.

It's also hard to analyze. You can average a star rating. You can't average a paragraph. Reading and making sense of hundreds of individual responses requires pattern recognition across the whole set — which is something human attention genuinely struggles to do at scale. We're good at reading individual responses deeply. We're not good at holding five hundred in memory simultaneously and finding what's consistent across them.

This is why structured formats — ratings, multiple choice, scales — have historically dominated audience research. Not because they produce better data, but because they produce data you can work with. The qualitative option was technically available; practically, it wasn't.

What synthesis does, step by step

The process, written for someone who doesn't work in data science.

First, each response is converted into a numerical representation called an embedding. This is a way of encoding meaning mathematically — responses that say similar things, even with different words, end up numerically close to each other. "I wish you'd covered this topic more deeply" and "there wasn't enough depth on that subject" are different sentences that land near each other in the mathematical space.

Second, responses are clustered by similarity. Groups of responses that share meaning get identified. These clusters are the raw themes — the things multiple people said independently, in different ways. The clustering is mathematical, not interpretive. It reflects the actual structure of the responses, not a researcher's intuition about how to categorise them.

Third, a language model reads each cluster and generates a description: what the theme is about, how many responses belong to it, representative quotes, and the emotional character of the responses. This is the step where AI adds its value — it can articulate in readable prose what the mathematical clustering found. But the themes themselves come from the math, not the language model.

What you get back is a structured view of what your audience said: themes ranked by frequency, with supporting quotes for each, and a sense of the distribution — what most people thought versus what a minority felt strongly about.

What this produces that reading can't

The most reliable insights from synthesis are often the quiet majority — the theme that 40% of responses touched on, without anyone making a dramatic point of it. These themes are nearly impossible to notice when reading individual responses, because they don't stand out. They're just consistent.

Reading selects for what's interesting. Synthesis selects for what's consistent. These aren't the same thing, and the difference matters. The response you'd notice while reading might be well-articulated but unusual. The pattern you'd miss might be the most important thing in the data.

We built AskEveryone around this process because we saw consistently that creators who could only read their responses were over-learning from the memorable ones and under-learning from the typical ones. Synthesis inverts that — you see the pattern first, then you can drill into the individual responses that exemplify it.

What synthesis can't do

This is worth being direct about.

Synthesis reflects the structure of the data. It doesn't evaluate whether that structure is important. If 40% of your audience says they want more interviews, the synthesis surfaces that clearly. Whether you should do more interviews — whether that's the right direction for your work, whether your audience is representative of the audience you're building toward — is still a judgment call.

There's also the self-selection limitation: the synthesis reflects what people who responded thought, not what your whole audience thinks. More responses improve the reliability of the themes, but the question of who chose to respond doesn't have a technical solution.

And the descriptive layer — where a language model writes the readable summary — can occasionally overstate confidence or miss nuance. Themes are real when they emerge from a large enough cluster with enough variation. The synthesis is a starting point for interpretation, not a substitute for it.

That said: the alternative — reading everything manually, or not collecting open responses at all because the data was unanalyzable — produces worse outcomes. Imperfect synthesis of rich data beats perfect analysis of data that never existed.

What creators actually do with this

The most common use is content direction — themes from audience responses become episode topics, newsletter angles, video briefs. The language in the themes becomes title candidates. The representative quotes become pull quotes or opening hooks.

The second most common use is product direction — understanding what your audience is missing, what they'd pay for, what they'd recommend to others if it existed.

The third is relational — sharing the synthesis results with your audience, closing the loop, showing them what they collectively told you and what you're doing with it.

All three depend on the synthesis making the responses usable. That's what it does.

Tags

synthesisqualitative analysisaudience insightsthematic analysis

Want insights like this for your audience?

Set it on autopilot. One question a week, every response analyzed into insights you can actually use.

Start free — no credit card