How we measure whether any of this works
A slice of every account's budget runs three dumber versions of the same idea.
A targeting system that cannot be compared to a list is a marketing claim, not a product. So every account runs four arms against the same product and the same mailbox.
Arm one: a static firmographic list with a generic angle. This is the database baseline.
Arm two: companies with one obvious signal, usually recent funding, with an angle that references it. This is the single-signal baseline.
Arm three: a stacked rules score. Several signals, no time ordering, no learning.
Arm four: the system. Temporal sequences, state inference, and budget that moves based on resolved outcomes.
The reported numbers are positive reply rate, qualified conversation rate, meetings held, opportunities and revenue, each with a 95% credible interval from a beta posterior and the sample size printed next to it.
No hypothesis is killed before 40 resolved actions, and at least 15% of budget stays on non-leading arms, because a system that collapses onto its first lucky segment is a random number generator with good manners.
If lift against the static list is under 1.0, the screen shows it under 1.0.
If you want this run against your own product, the first brief is free.
Get my first brief