The corpus consists of real conversations captured at designated counters in participating stores, under visible signage, a QR-accessible plain-language notice, and a verbal disclosure inside the greeting. No conversation was staged or solicited for research purposes.
Behaviours were defined before any scoring took place, derived from top-quartile conversations where the quartile was selected on point-of-sale outcome rather than on managerial reputation. Each definition carries an inclusion and an exclusion boundary, and an eligibility test so that conversations which never called for the behaviour are excluded rather than counted as failures.
Scoring accuracy was validated against human double-scoring and is reported separately by language and accent group. Definitions that produced repeated disagreement were rewritten. No voiceprints or voice embeddings were created at any stage, including model training; attribution was by capturing seat and roster shift.