← Every exhibit
The Metric Hierarchy in Controlled Experiments
| Metric Class | Operational Purpose | Primary Examples | Decision Rule in Experimentation |
|---|---|---|---|
| 1. Overall Evaluation Criterion (OEC) | The unified strategic objective function that defines organizational success | Composite score balancing 30-day net customer retention and annualized margin contribution | Primary criterion for declaring experimental victory and rolling out code to production |
| 2. Direct Driver Metrics | Local transactional metrics measuring immediate behavioral response | Call-to-action click rate, checkout form start rate, onboarding completion rate | Diagnostic indicators to understand why the treatment impacted the primary OEC |
| 3. Guardrail / Health Metrics | Invariant system and business health indicators that must not be degraded | Server latency (p99), checkout error rate, customer support ticket volume, refund rate | Non-negotiable veto power: if any guardrail metric breaches thresholds, the test is aborted |
| 4. Long-Term North Star Metrics | Macro-economic business metrics tracking durable shareholder value | Customer Lifetime Value (LTV), Net Revenue Retention (NRR), brand equity indices | Evaluated retrospectively across long-term holdout cohorts to verify sustained impact |
Swipe or scroll horizontally if the table is wider than your screen.
Reference & Evidence
Source: Table from this essay. Sources and interpretation are given in the article.
Related exhibits
-
Guardrail Metric Verification and Downstream Economic Translation
From the essay What is an A/B test? randomized controlled trials, statistical power, and governance
-
The A/B test experimental architecture
From the essay What is an A/B test? randomized controlled trials, statistical power, and governance
-
Decision Errors and Statistical Power
From the essay What is an A/B test? randomized controlled trials, statistical power, and governance