Actuals is a free plugin for Claude Code, Codex, Cursor, and 30+ other coding agents: it interviews you about your business, writes a versioned measurement spec, audits your dashboards against 20 named vanity-metric anti-patterns — and generates the instrumentation to measure what's real.
developers using AI in METR's randomized trial — while believing they were 20% faster
Microsoft's flat multiplier behind Copilot's "assisted value" dollar figure
of companies can attribute any profit impact to AI — McKinsey, State of AI
Each skill triggers on plain English too — say "are these numbers real?" and the audit just runs.
A structured interview — what the AI does, who benefits, and the decision your metrics must inform — becomes a versioned measurement spec with formulas, owners, and guardrails.
Point it at dashboards, tracking plans, SQL, or the ROI deck. Every finding cites a named anti-pattern, quotes your artifact, carries a published receipt, and ships a fix.
Events with typed constants, SQL/dbt models implementing each formula verbatim, a pre-launch baseline snapshot, and an eval harness with judge–human calibration built in.
Vetted MCP configs for the tools you already run — merged into your project non-destructively. Secrets stay as env-var placeholders you export yourself.
A self-contained dashboard.html — metrics vs baselines and targets, tripwires, open findings, the claims ledger. Renders recorded values only, and says so.
The audit's review mode checks spec drift, definition rot, stale owners, and overdue calibrations — the maintenance loop that keeps the spec from becoming the thing it replaced.
One interview produces metrics/MEASUREMENT.md — the single source of truth your reporting has been missing.
VM-02 · Minutes-times-Wage Dollarization where: dashboard-export.csv:4 quote: "1,240 hrs × $72/hr = $89,280 / month" receipt:[MSFT-COPILOT] flat-multiplier value math fix: cost per resolved ticket — constants live in the spec, re-measured quarterly
From the demo fixture that ships with the plugin — a SaaS with an AI support assistant and a very confident dashboard.
Stable IDs, detection signals, severity, published evidence, and a concrete fix for each.
"Devs report saving 11 hrs/week." Perception inverts reality — METR's trial participants were 19% slower while feeling 20% faster.
Actions × minutes × $/hour. Every constant is an assumption in a dollar costume — and it can break silently for months.
"10,000 AI conversations!" Out of how many tickets? A numerator without a denominator is a press release, not a metric.
A curated, verified MCP registry — written into your config on your confirmation. Env-var placeholders only, never literal keys.
Instrument publishes metric definitions as PostHog insights and dashboards — your data carries the live numbers, the spec stays the source of truth.
Run the deterministic checks from CI or any agent — no skills required, zero dependencies, stdio.
Actuals ships with its own measurement spec, written with its own method. GitHub stars and installs are listed there under "Vanity Metrics We Will Not Use." Its claims ledger commits to never claiming it improved your metrics without your baselines. There's no telemetry in the plugin — and this page loads no external resources and runs no analytics.
Same five skills everywhere — Actuals follows the open Agent Skills standard, so it installs into whichever agent you code with.
> /plugin marketplace add Hamza-Saraswat/actuals > /plugin install actuals@actuals-marketplace█
Then just ask: "what should we measure for our new AI feature?" — or hand it your dashboard and ask if the numbers are real.