Features·

Comparing Two Audiences Is Harder Than Comparing Two Numbers

Two audiences can agree on what they care about and disagree on how they feel about it — or care about completely different things in overlapping words. Honest comparison needs structure.

Two sets of themes with connective lines showing overlap, divergence, and translation

You run a survey across two audience segments and get a clean result: Group A scored your content 4.2, Group B scored it 3.8. Group A is happier than Group B. Done.

Or are they? The problem with numeric comparison is that it flattens a complicated thing into a simple thing, and the simple thing often hides what's actually going on. Two groups can share a score and care about completely different things. Two groups can share a complaint phrased the same way and mean entirely different versions of it. Two groups can agree on everything surface-level and disagree in the places that matter most.

Comparing two audiences honestly requires holding both stories in view at the same time. That's harder than comparing two numbers, and most tools don't do it well.


The three failure modes of audience comparison

Most audience comparison work fails in one of three ways, and each one produces a different kind of wrong answer.

Flat comparison — two groups, one metric, one result. The output is "Group A is happier than Group B" or "Group A engages more than Group B." This is fast and clean and almost always too reductive. It tells you which group won on some dimension. It doesn't tell you what either group actually thinks.

Separate comparison — analyze each group independently, present the results side by side. This is better, but the reader has to do the comparison in their head. Themes for Group A, themes for Group B, and you're squinting between them trying to figure out which ones overlap and which ones diverge. By the time you've made sense of it, you've forgotten half the details.

Honest comparison — pool the responses from both groups into one shared topic model, so the themes are the same across both groups, and then measure how each group's responses distribute across those shared themes. Now you can say "both groups mentioned X, but Group A mentioned it three times as often" or "Group A mentioned Y at all, Group B barely did."

That third approach is what we built comparison synthesis to do. It's not just running two analyses and stapling them together — it's running one analysis that both groups participate in, so the comparison is structurally honest.

Why pooled analysis matters

When you analyze two groups separately, the clustering algorithm sees different data in each run. That means the themes that emerge from Group A might use slightly different clusters than the themes from Group B — not because the underlying signal is different, but because the clustering is responding to the specific responses it saw.

You end up with "Group A's top theme is 'lack of depth'" and "Group B's top theme is 'wanting more advanced content.'" Are those the same theme? Different? You don't know. The model didn't compare them — it ran twice on different inputs.

Pooled analysis fixes this. Both groups' responses are fed into one clustering run. The themes that emerge exist across both groups. Then, for each theme, you can look at how responses distribute: "of the 73 responses in the 'lack of depth' theme, 58 were from Group A and 15 were from Group B."

Now the comparison is structural. The themes are the same because they came from the same model. The differences are in how the groups weight the themes, which is what you actually want to know.

What divergence actually looks like

The output of comparison synthesis is a set of shared themes, each annotated with how each group contributed to them. There are three distinct patterns worth watching for.

Convergent themes — both groups mentioned this at similar rates. These are the things your full audience agrees on, and you can act on them without worrying about which group you're optimizing for.

Divergent themes — both groups mentioned this, but one group cared about it much more than the other. These are the places where you have to choose, or where you need to serve both groups deliberately.

Unique themes — a theme that one group mentioned and the other essentially didn't. These are the places where the two groups aren't disagreeing; they're living in different worlds. Group A is worried about something Group B doesn't even know to think about.

The third category is the one most tools miss, because aggregate analysis treats a theme mentioned by 5% of respondents as a weak signal. Segmented analysis can see that the 5% is actually 40% of one group and 0% of the other, which is a completely different signal.

The statistical honesty problem

Comparison synthesis produces visibly different numbers for different groups, which creates a temptation to over-interpret. "60% of paid subscribers mentioned X, versus 15% of free subscribers" sounds like a clear finding. But it's a finding about the people who responded in each group, not about the groups as a whole.

If 900 paid subscribers got the question and 180 responded, you're hearing from 20% of paid. If 18,000 free subscribers got the question and 600 responded, you're hearing from 3% of free. Both groups have the same self-selection problem, but the rates are different, which means the representativeness is different.

Honest comparison requires acknowledging this. "Of the people who responded, X% of paid raised this theme" is different from "X% of paid subscribers think this." The first is what the data says. The second is what the data might mean, and the path from one to the other involves judgment about who chose to respond and why.

Comparison synthesis doesn't automate that judgment. It presents the structural comparison cleanly so you can apply your own interpretation. That's the honest version.

When comparison is worth the complexity

Not every question needs segmented analysis. If your audience is homogeneous, aggregate analysis is fine. If the segments you're comparing are tiny, the noise in each segment will dominate the signal.

Comparison is worth the complexity when:

  • You have structurally different audience groups (paid vs free, enterprise vs individual, early adopters vs recent joins) and you suspect they might answer differently.

  • You're about to make a decision that affects one group more than another (pricing changes, format shifts, content direction pivots).

  • You've noticed disagreement in the aggregate data that you can't explain without more context.

  • You're running a second wave of a previous question and want to know how the response has shifted over time.

For those cases, the comparison view shows you something aggregate analysis can't: whether "your audience" is actually one audience or several, and whether the decisions you're making serve all of them or just the largest one.

What comparison synthesis doesn't do

It doesn't tell you which group's feedback matters more. That's a business question, not a data question.

It doesn't solve the self-selection problem. You're still hearing from the subset of each group that chose to respond, and those subsets aren't necessarily representative of the groups as wholes.

It doesn't replace reading individual responses. The comparison view shows you structure; individual responses show you the specific language and reasoning that structure is built from. Both matter, and good analysis involves moving between them.

What it does is remove the specific failure mode where you think two groups agree because the aggregate numbers look similar, when actually they're answering a different question in the same words. Catching that failure is often the difference between a decision that works and a decision that ages badly.

Tags

comparison synthesisaudience segmentationstatistical divergenceinsights

Want insights like this for your audience?

Set it on autopilot. One question a week, every response analyzed into insights you can actually use.

Start free — no credit card