← All writing

Positioning

How to know if your messaging will actually work (before you ship it)

By Nick Pham10 min read

TL;DR

Internal sign-off measures internal alignment, not market resonance. The False Confidence Filter is a two-day protocol of four tests, each run with five ICP-matched buyers who have never met you. The Stranger Test catches clarity gaps, the Priority Test catches relevance gaps, the Clone Test catches differentiation gaps, and the Champion Test catches whether a champion can resell your value unaided. Stranger and Priority failures force a rewrite. Clone and Champion failures are usually fixable.

Internal sign-off doesn't predict whether messaging works in market. It predicts that everyone in the room stopped objecting.

Picture the all-hands where the slide goes up and the presenter asks if anyone has questions. Nobody does. The silence gets read as agreement, and it almost never is.

A messaging review works the same way. By the time the words reach the homepage, every stakeholder has nodded twice.

Then the page goes live, demo requests stay flat, and nobody can explain it. The messaging tested well in the room. It never met anyone holding a budget.

There's a way out and it takes two days. We call it the False Confidence Filter. Four tests, five ICP-matched strangers, and a go/no-go decision built on evidence instead of nods.

Two votes, two ballots

Internal approval measures internal alignment. Whether the market cares is a separate question, and nobody in the room can answer it.

The team voting yes on the hero is voting on whether the words match the strategy we built over six months. The buyer landing on that page is voting on whether the words match the problem they woke up worrying about.

Two votes. Completely different ballots.

Pipeline360's 2026 survey of marketing leaders found that most call their content strategy advanced. Only 19.1% track whether it contributes to pipeline.

Forrester's State of Business Buying reports the cost of that gap year after year. Most B2B purchases stall, and most buyers end up unhappy with whoever they picked.

The team that wrote the words never sees any of it. We're close to the product, our buyer is far from it, and internal review only tells us how the words sound to people who already know the answer. Same dynamic we covered in the post on why your positioning sounds right but nobody is buying.

The comfortable fallbacks

Three answers come up every time someone raises this. None of them work before launch.

A/B testing needs volume and a live page. Most early-stage and mid-market teams can't reach significance on a single test, and by the time they can, the bad version already shipped. It also only reports which of two options did better, so we can iterate our way to a less-bad version of broken messaging and never learn it was broken.

"We'll fix it after launch" sounds cheap and isn't. If the launch teaches buyers we're generic, there's no clean reset, and we spend the next quarter fighting our own first impression.

Leadership approval is weakest of all. Every executive has ten times more context about the product than a buyer will ever have, and they've heard the strategy often enough that the words now sound inevitable. Asking them to stand in for someone deciding in 30 seconds whether we're worth the click is a category error.

Which leaves structured feedback from ICP-matched strangers.

The False Confidence Filter

The Filter is four tests, run against the same messaging, with five ICP-matched buyers, inside 48 hours. Each one catches a failure that internal review structurally cannot see.

TestWhat it catchesFailure mode
Stranger TestWhether someone outside the bubble understands what we do in 30 secondsClarity gap
Priority TestWhether the problem we describe is one our ICP prioritizes this quarterRelevance gap
Clone TestWhether our messaging is distinguishable from the three closest competitorsDifferentiation gap
Champion TestWhether a champion can re-explain our value to their CFO unaidedInternal selling gap

A win on one doesn't cover a loss on another. Clear but irrelevant gets ignored, and relevant but indefensible inside the buyer's own company gets killed in committee.

About the five. It comes from Nielsen Norman Group's usability model, where five users surface about 85% of problems.

Published research asks for more. Wynter recommends around fifteen and the peer-reviewed saturation work lands between nine and seventeen. Five is a smoke test, and a smoke test is what a launch decision needs.

The Stranger Test

Can someone who's never heard of us, who matches our ICP, say back what we do after 30 seconds on the homepage?

Recruit five buyers who fit. Not customers, not open prospects, not anyone in our network. Wynter, UserTesting, and Respondent all recruit ICP-matched B2B respondents.

Show them the hero, headline, subhead, and primary CTA for 30 seconds, then ask what the company does, who it's for, and what problem it solves. It passes when four of five get all three right. It fails when anyone guesses wrong, asks for clarification, or hands back a generic category.

Our own team can't run this, because they can't un-know the product. They read the homepage as a reminder, and strangers read it as an attempt to teach them something new.

In a 1997 Nielsen Norman Group study, 79% of test users always scanned any new page they encountered and only 16% read word by word. That one is nearly thirty years old and still holds, because it describes how eyes work rather than how a market behaved in one quarter. Our messaging has to survive a scan.

The Priority Test

Is the problem we describe one our buyer is trying to solve right now?

Same five-person recruit. Show the hero plus the value proposition section, then ask three questions:

  1. "Is this a problem you're trying to solve right now?"
  2. "If yes, where does it rank against everything else your team has this quarter?"
  3. "If you stopped solving it entirely, what would happen?"

It passes when three of five call the problem real, currently top-five, and expensive to ignore. It fails when someone calls it real but not urgent, or says it belongs to another team.

Think about the extended warranty at checkout. The clerk asks whether we want the protection plan on the dishwasher, and the risk is real. Appliances break.

We say no anyway. Politely, every time, because a real problem that isn't urgent loses to whatever we came in for. Most messaging fails exactly there, and "great pitch, we'll keep you in mind" is the sound it makes.

The Clone Test

Take our hero and value props, pull the same sections from the three closest competitors, strip every logo and visual cue, then show all four to five ICP-matched buyers and ask which one is for them and why.

It passes when four of five pick ours and their reasons cluster on the differentiator we're claiming. It fails when they pick at random, can't say why, or choose a competitor because those words are sharper.

This is where most messaging dies. We think the differentiation is obvious, then the logos come off and it evaporates. Same problem as how to differentiate when every competitor sounds the same, tested instead of argued.

The Champion Test

Can our champion re-explain the value to a CFO with none of us in the room?

The champion is rarely the only decision-maker. They carry our words into a meeting we'll never attend, which is the dynamic behind buying committee messaging. Messaging that can't survive a retelling doesn't close.

Find five recent champions across wins, losses, and stalled deals. Ask each one to pretend we're a skeptical CFO who's never seen the homepage and has ten minutes. Record it and don't interrupt.

It passes when four of five reproduce our core claim, our differentiator, and our category. They can use their own words as long as the meaning survives the trip.

It fails when they give a generic version, mix us up with a competitor, or fall back to feature lists. The Voice of Customer Research playbook is the slower version of the same exercise.

Fatal and fixable

Twenty data points, one decision.

A Stranger or Priority failure is fatal. If buyers can't say what we do, or don't care about the problem, no other strength compensates. Rewrite.

A Clone failure usually means the differentiator got buried. Look at which value props sent buyers elsewhere, then sharpen the hero and clean up the hierarchy.

A Champion failure usually means the messaging is buyer-facing but not committee-ready. Layer in CFO-facing proof and payback logic without losing the hook, the way we described in the messaging house framework.

Four tests at four of five or better is a green light. Ship it.

The rubric earns its keep because messaging debates spiral when the evidence is soft. "I think the headline could be sharper" goes nowhere. "Three of five ICP buyers couldn't explain what we do in 30 seconds" ends the argument.

The two-day version

Day one is recruiting in the morning, then Stranger and Priority back to back with the same five people in the afternoon. Day two is the Clone Test with five fresh respondents, then the Champion Test against five champions from the CRM.

By the end, every test is marked pass, partial, or fail, and the call is written. Ship, refine, or rewrite.

If it fails Stranger or Priority, the page we were going to launch is a draft again, and the Filter cost us two days instead of a quarter. The same protocol covers homepage rewrites and repositioning sprints, including why your SaaS homepage isn't converting.

Nobody in the room has questions.

Go find five people who do.

What to do next

If you're sitting on messaging that everyone approved and nobody outside the building has read, that's the moment the Filter exists for.

A Bare Strategy messaging sprint runs the four tests with real ICP-matched buyers, then hands back the rewrite the evidence asks for.

If that's where you are, start here. The first conversation is free.

Frequently asked questions

Five per test is a deliberate floor, not a research standard. It comes from [Nielsen Norman Group's usability model](https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/), where five users surface about 85% of problems, while [Wynter recommends closer to fifteen](https://wynter.com/post/a-good-sample-size-for-qualitative-research) and the peer-reviewed saturation work lands between nine and seventeen. Before a launch you're hunting fatal flaws, and those show up early. If two of five can't explain the message, ten more interviews only confirm it.

No, and the reason is structural. Surveys give you scaled answers to questions you already knew to ask, and validation needs unprompted reactions to the ones you haven't thought of. The Priority Test shows where your problem ranks against everything else on a buyer's quarter, which a survey can't capture without contaminating the question.

That's common, and it isn't a contradiction. Customers love the product because they went through onboarding and watched the value land, while strangers are reading 100 words on a page. The failure says the gap between the product experience and the homepage experience is too wide for a new buyer to cross, so leave the product alone and rewrite the page.

Run it with outbound. Identify ten ICP-matched contacts on LinkedIn, message them directly, and offer a $50 gift card for 20 minutes. The quality bar doesn't move, so never use customers, prospects in open deals, or anyone in your own network. A platform delivers respondents in a day and outbound takes a week.

Related reading

The author

Nick Pham

Founder of Bare Strategy. Twenty years in B2B marketing, the last decade in product marketing inside enterprise software.

More about the operator →

If this is where you are

Bring the problem, not a brief, and you'll leave the first conversation with something useful either way.

Start a conversation