Article summary
Audit a ‘100,000 to 5,000 high-converting users’ claim through sample flow, field lineage, denominators, controls, cost and reproducibility.
“We screened 100,000 phones into 5,000 high-converting users” is not a case-study conclusion. It is a claim to audit. Ask where the 100,000 came from, how 95,000 were excluded, whether “high-converting” was defined before screening, and whether an unscreened control exists. Without that evidence, precision in the headline makes the claim more suspicious, not less.
Artifact one: the sample flow
| Stage | Required disclosure | Red flag |
|---|---|---|
| Original 100,000 | Source, date, country, permission | Only “our database” |
| After format/dedupe | Excluded count and reason | Denominator suddenly changes |
| After TG task | Explicit, unknown and error | Unknown counted unregistered |
| Final 5,000 | Every rule and row count | Rules selected afterward |
Artifact two: the task schema
TG Registration returns phone and registration result. Activity may return UserID, username, offline time, active days, names, VIP and frozen status. Gender and Age adds avatar URL, age and gender. None of those fields is “high conversion.”
Test for circular definition
If the case defines recent activity as high conversion and then uses an activity field to prove it found high converters, the reasoning is circular. Conversion needs an independent business outcome such as qualified inquiry, payment or retention within a predefined window, collected separately from checking fields.
Do not mix four denominators
| Rate | Denominator | Question |
|---|---|---|
| Format pass | Original unique phones | Input quality |
| Platform registration coverage | Processable phones | Platform observation |
| Contact-eligible queue | Records with source and permission | Operating eligibility |
| Conversion | People entering the comparison | Business outcome |
Look for a control
The strongest design preregisters random assignment between an observation-assisted process and the existing process within one qualified sample, keeping message, team, time and offer consistent. Outcomes for only the final 5,000 cannot show whether those people would have converted anyway.
Audit temporal order
The checking observation must precede the business result, and selection rules must precede seeing outcomes. Choosing 5,000 from future purchasers and claiming they could have been identified earlier is data leakage. Require timestamps, versions and a frozen analysis plan.
Bring excluded people back into view
Report later outcomes for the 95,000, at least through a representative sample, to see whether many converters were missed. Looking only at selected records creates survivor bias. High precision with very low recall may discard most real opportunities.
Cost and negative events are evidence
List cleaning, task, human review, content, sending and governance cost, plus opt-outs, complaints, account restrictions and wrong merges. Revenue without incremental contribution margin cannot show that the case deserves replication.
Contents of a reproducibility package
Provide an anonymized sample description, TXT generation rules, field dictionary, batch manifest, exclusion code, message versions, control assignment, metric SQL and uncertainty. Personal data need not be disclosed, but an independent reviewer must be able to recalculate stage counts.
Six common red flags
No source, no unknowns, fields inconsistent with the product, only the best country shown, conversion definition changed after the result, or revenue substituted for incremental profit. Any one needs evidence; several make the case unsuitable for procurement or marketing claims.
How to write the audit finding
Rate data integrity, task validity, causal evidence, reproducibility and risk control separately, listing missing artifacts. A qualified case is not one with an impressive number. It lets a third party trace every step from 100,000 to 5,000 and confirm that the resulting value was incremental.
