PMaker home
Only real behavior records settle whether a product worksSmall samples misleadThey hand you the comfortable answer in the direction you were hoping forTwo of three clicked, and you read "sixty percent"Check the four criteriaBehavior over claims, paths over single points, enough sample size, and real scenariosWhen you have no dataInstrument before launch, ship a small version to real users, never use surveys as behaviorInsiders are people who already know they'll use it

Judging needs data, and data needs records. Without instrumentation and sample size, any "feels good" is just a feeling.

Judge by Real Data

Small samples and your own team's feelings don't count. Only the recorded behavior of real users in real scenarios does.

"I find it good to use," "my friend says it's nice," "our user numbers are going up"—none of these are data. Only one thing settles whether a product works: the recorded behavior of real users in real scenarios.

What you'll run into:

  • You conclude from "the team thinks it's good," and a few people liking it is enough to ship
  • A survey comes back with a few dozen responses, and you treat it as a rule
  • The feature shipped, but there's no behavior data, so you can't say whether it was used or where people stopped

Why small samples don't count

The problem with small samples and your own team's feelings isn't that they're "unreal"—it's that they lead you astray: they hand you the comfortable answer in the direction you were already hoping for.

  • Insiders are people who "already know they'll use it." Whatever you make, they say it's good, because what they back is you, not the product. What you actually need is the people who still need convincing.
  • Small-sample noise can support any conclusion. Two out of three people clicked a button and you read "sixty percent"—but that's just rolling the dice three times. Small enough, and numbers reflect luck, not rules.
  • Asked face to face, people say what they think you want to hear. People tend to perform cooperation, rationality, and helpfulness. "I really need this feature" doesn't mean they'll actually use it next month.
  • A short-term bump from a campaign isn't the product's doing. Movement has to be attributable. If you can't tell which link in the chain caused it, a rise can't be repeated.

What counts as real data

Numbers are numbers, but some count as data and some are just noise. Four criteria:

  • Look at behavior, not statements. What users actually did is far more reliable than what they said. Don't ask "do you need it"; count "who used it and how many times."
  • Look at the full path, not a single point. "Page views went up" means nothing; look at the whole chain from entry to completion and find where people stopped.
  • The sample needs scale. At least enough to separate fluctuation from trend. Single-digit samples expose problems but can't support conclusions. Note this is different from usability testing: a handful of users exposes most of the "where is it hard to use" problems, but judging whether the product has value and deserves more investment needs a real sample size.nng5
  • It has to be a real scenario. They use it on their own, with nobody watching, and no reward bait. Data from test or demo environments doesn't count.

Run these four over a claim like "page views went up" and it falls apart fast: up by how much, who rose, which entry they came from, where they left. Miss any link in that chain and the number can't drive a decision. Judging a product rests on records that survive this kind of questioning.

When you have no data

"I just don't have real data yet"—that's the most common situation, and it has an answer:

  • Instrument before you launch. A feature without tracking is a feature that never shipped. Treat tracking as part of the feature and walk through it together. See track before launch.
  • Ship a small version to real users first. Real behavior at small traffic beats the feelings of a hundred insiders. Narrow audiences and minimal versions exist to bring real data earlier. See thin slice.
  • Don't let surveys stand in for behavior. A survey tells you what they say they want; behavior tells you whether they actually use it. The two can complement each other, but a survey is not behavior data.
  • Decide what to validate before deciding what to watch. Once the north star is set, you know which numbers deserve tracking and which are just pretty. See building a metrics system.

References

  1. Why You Only Need to Test with 5 Users — Nielsen Norman Group