Claude Code plugin
ab-test-playbook
An A/B test engine for Claude Code. It suggests test scenarios, designs them against a written methodology, audits plans you already have, interprets the numbers you paste in, and renders every scenario as a single-file HTML card.
What a scenario card looks like
Rendered below with no install and no build step: this is the real output of scripts/build_card.py run against the shipped example, not a mockup of one.
Two mockups differing in exactly one thing, the tested element boxed in red, and the three mandatory boxes. Open full size.
What makes it opinionated
- One variable. A variant pair changes exactly one thing. If you insist on more, the output says the result cannot be attributed.
- One primary KPI, always a guardrail. Five metrics of equal weight is not a measurement plan. Every scenario names the metric that decides, and at least one that must not break.
- No dark patterns, and protections are not test subjects. Fake scarcity, hidden totals and uncloseable modals are refused. So is treating bot checks, identity verification or legal consent steps as friction to remove.
- Evidence is labelled. Every proposal states whether it rests on your own data, an archive precedent, a sector observation, or intuition — and says so when it is weak.
- No sample promises without traffic. Duration and significance are not estimated from thin air.
Install
# as a Claude Code plugin
/plugin install ab-test-playbook
# or install the skills with the skills CLI
npx skills add ali-demirbas/ab-test-playbook --all
Then, inside Claude Code: /ab-test suggest, /ab-test design, /ab-test audit, /ab-test results.
What is inside
- 211 curated scenarios across journey stages, each already carrying the three boxes.
- A real stats engine — z-test, confidence interval, sample size, sample-ratio-mismatch check, all stdlib-only.
- Adversarial review — a methodology critic reads every generated scenario, and a mockup reviewer checks the two variants differ in exactly one thing, before anything reaches you.
- A portable schema — a test definition independent of any experimentation platform.