Almost every conversation-intelligence pilot is designed the same way. Pick ten stores, coach five, leave five alone, compare conversion at the end of the quarter. It feels like a controlled experiment. It is not one, and the number it produces will not survive a finance review.

Two stores are never the same experiment

Two doors in the same chain differ in footfall quality, catchment income, tenure mix, manager competence, local competition, mall anchor tenancy, and the accident of which two associates happened to leave in February. Any of those moves conversion by more than a coaching programme will in ninety days. When the treated stores come out ahead, you cannot separate the coaching from the fact that one of them sits next to a new cinema.

The problem is not that the effect is small. It is that the noise between doors is larger than the effect you are trying to detect, and no amount of averaging across ten stores fixes that when the differences are systematic rather than random.

The selection problem nobody admits to

Treated stores are rarely selected at random. They are selected because the district manager is enthusiastic, or because the store is already good, or because it is close to the office. Each of those makes the treated group different from the control group before the programme starts. A vendor who lets you pick your own pilot stores is not being accommodating; they are letting you contaminate your own measurement, usually in their favour.

If a vendor is comfortable with store-versus-store, ask why. The method’s weakness happens to favour whoever is selling.

What works instead: the matched cohort inside one store

Run the comparison inside a single door. Capture every conversation in the store, but coach only some of the associates. The coached and the capture-only group now share footfall, catchment, manager, promotion calendar, inventory position and weather. The only material difference between them is the coaching.

Then swap the groups halfway through. If the effect follows the coaching rather than the people, you have evidence rather than a coincidence. If the previously uncoached group improves after the swap and the first group holds, that is about as close to proof as a retail floor allows.

Three conditions that make the readout auditable

  • Point-of-sale access at ticket level. Without it you are measuring behaviour scores against behaviour scores, which proves only that coaching changes coaching.
  • Published capture coverage, weekly. If coverage in a store falls to forty percent in a given week, that week has to be visible rather than quietly averaged into a flattering total.
  • Cohort assignment fixed before the first day of counted measurement, and written down. Groups that get adjusted mid-pilot produce numbers that mean nothing.

Why this matters more than it sounds

A weak method does not just produce a weak number. It produces a number that gets defended for three years, because nobody wants to reopen a decision they already announced. The reason to insist on within-store cohorts is not statistical fastidiousness. It is that the alternative gives you a result you cannot revisit honestly when it turns out to be wrong.