The tests, the pull request they report on, and the tracking the same walk keeps correct. Then the tools your own model can call, what fails a build, and what we run for you. What is not in the box is near the bottom.
1 sentence per test5 verdicts, never blurred0 model calls on a replaythe MCP tools behind the testing and instrumentation half6 npx commands6 editors auto-wired
smolanalytics runs end-to-end tests that have no test code. You write one sentence describing what should work; an agent opens a real browser, reads the page through its accessibility tree so it picks an element rather than a pixel, and returns a verdict. A run that passes is recorded and replayed with no model call at all — measured on this site at 8.0s for the first run and 1.4s for the second — and the agent wakes again only when the recording stops fitting the app. On a pull request the whole suite reports as one comment, edited in place, with five verdicts kept apart on purpose: passed, failed (the app did not do what the sentence describes), flaky (it failed and then passed on a retry, so it is never counted as a pass), stale (a recording no longer fits, which a replay cannot tell from a rename) and errored (our runner could not run, never your app). Setup is a URL you already have and a sentence: no account for the first run, no GitHub App, no preview environment built, nothing written to your repository, and in CI it runs on your own Actions runner and comments with the token Actions provides free. The same walk through your product knows which user actions exist, so it also writes and maintains your analytics tracking inside the SDK you already use — Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment — and when a refactor deletes a tracking call it opens a pull request putting that exact line back. Two things need no account: npx smolanalytics audit reads the repo you are standing in and names the user actions nothing measures, with the file and the line and no network call, and GET /api/readability?url= checks whether an AI crawler can read a page. 14 days, no card. Then $19/mo with 100 tested pull requests, 10c each after.
your tests
There is no test code to maintain, because there is no test code. A test is one sentence in a markdown file, and the agent works out how to do it every time your UI changes.
$ npx smolanalytics test --url https://yourapp.com --test "the pricing page shows a monthly price"
The whole setup: a URL you already have and a sentence. Staging, a deploy preview, or localhost through a tunnel. No account, no GitHub App, no preview environment to build, and nothing written to your repository. Playwright is downloaded the first time this command runs, with a line saying what and why, and never for anyone who does not use it.
you ▸ (the agent reads the page, it does not squint at it)
It works from the accessibility tree — the role, name, value and state of everything on the page — and picks an element, so the click that follows is a real locator with actionability checks. Nothing here guesses a coordinate off a screenshot, which is how a run clicks the wrong thing and then blames the wrong feature. When the list is truncated it says so, rather than letting the agent conclude a button does not exist.
you ▸ (a run that passes is recorded, so the next one is free)
The steps that worked are compiled into a plan and replayed with no model call at all. Measured on this site: 8.0s for the first run with the agent, 1.4s for the second — one flow on one machine, not a benchmark; your app and your runner will give you your own numbers. The agent only wakes when the recording stops fitting the app, which is exactly when judgement is worth paying for.
you ▸ (passed · failed · stale · errored)
Five verdicts — passed, failed, stale, errored, flaky — and stale and errored are never worded or coloured as a failure. Failed means the app did not do what the sentence describes: that is a bug report. Stale means a recording stopped fitting, and a replay cannot tell a rename from a removal, so the agent goes and re-checks it. Errored means our runner could not run — no browser, no key, no network — and is never a statement about your app.
$ npx smolanalytics test --suite tests/ --url "$URL"
A folder of markdown files, one heading per test, the sentence underneath it. The heading is the test's identity and the recording's filename, so renaming a heading throws that recording away and the agent runs it fresh.
you ▸ (the flags worth knowing)
--headed watches it happen in a real Chromium window. --plan <file> replays one recording and only wakes the agent if it no longer fits. --plans <dir> is where recordings live, .smolanalytics/recordings by default. --max-steps raises the step budget for a long flow.
on every pull request
The CI half is a workflow file you copy in, not an application you install on your repositories. It runs on your own Actions runner, against a URL your host already built, and comments with the token Actions hands every job for free.
$ npx smolanalytics test --suite tests/ --url "$URL" --comment
One comment on the pull request, edited in place on every push rather than stacked. A row per test: the verdict, whether it replayed or woke the agent, and how long it took. Each failure is written out in full underneath the table, because the cell truncates it and the person reading was not watching the browser.
you ▸ (what the workflow asks of you)
Copy the template to .github/workflows/e2e.yml, add ANTHROPIC_API_KEY to that repository's secrets, and keep one of its three preview-URL steps. Its permissions are contents: read and pull-requests: write, and nothing else, so a compromised step in it cannot push code or reach another repository.
you ▸ (the URL comes from the host you already use)
The template ships the Vercel wait-for-preview step; Netlify and Render publish a preview per pull request too and have equivalent actions. The two alternatives are commented out beside it: point at a URL you already have, or build and serve the app inside the job. Anything reachable works, because a reachable URL is the entire input.
you ▸ (the recordings are cached between runs)
A CI runner starts empty, so with no cache every test on every pull request is a fresh agent run, forever. The template restores .smolanalytics/recordings before the suite and saves them after it with if: always() — the run with a failing test is the run that re-recorded the most, and a post-step that only saves on success throws exactly that work away.
you ▸ (who is skipped, and why it is a skip)
A fork's pull request and dependabot both run without repository secrets. The template skips them rather than running every test against an empty key and then failing the comment with a 403.
you ▸ (the exit codes, if you gate on this later)
0 every test passed, 1 a test failed, 2 this runner could not finish. Published as a contract, because a gate that cannot tell those apart turns an outage on our side into a bug report about your app. Week one the template sets continue-on-error: a new tool that puts a red X on a pull request before it has earned any trust gets uninstalled instead of read.
you ▸ (and the verdicts land on your project page)
Set SMOLANALYTICS_PROJECT and SMOLANALYTICS_WRITE_KEY in the job and every run is recorded: the suite, what each test last did, and how many runs needed a model. The POST authenticates with the write key because CI has no cookie, and a delivery that fails leaves the verdict standing — a test tool that fails a build over its own telemetry gets removed the same day.
your tracking
The same walk through your product knows which user actions exist, so it also writes and maintains the tracking for them — inside the SDK you already run. We do not replace your analytics; we keep its instrumentation correct.
$ npx smolanalytics audit
Reads the repo you are standing in and names the user actions nothing is measuring: payments, signups, logins, invites, deletions, shares, uploads, each with the file and the line. It counts the tracking you already have, in whichever SDK wrote it, so a working PostHog reads as working. No account, no key, and no network call on this path, so your code never leaves the machine.
you ▸ what does this app do that nothing is measuring?
instrumentation_coverage
The same question asked of a repo your agent already has open: form submissions, auth flows, payments, mutating API handlers. Absence is invisible in event data, because an action nobody instrumented looks exactly like one nobody performed.
you ▸ write the tracking this repo is missing, in the SDK it already uses
propose_instrumentation
The exact edits at the exact call sites, in Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment — whichever of them the repo already runs. A second SDK is never added beside a working one, because two SDKs on one action double-count it.
you ▸ prove each event is wired and firing
verify_instrumentation
A firing / wired / missing table per event, so you never have to trust that the edit worked. instrumentation_health does the same against your declared plan on demand, and event_source names the file and line an event fires from — a count that dropped is a symptom, a missing call site is a cause.
you ▸ (a tracked event goes silent: one PR puts the line back)
When an event in your tracking plan stops arriving but the traffic behind it doesn't, smolanalytics finds the commit that deleted the track() call and opens a pull request putting it back. The line comes back verbatim out of the parent blob, not generated, in the SDK it was deleted from. Off by default twice over: a deployment-wide switch, and a per-project setting that starts at off. Worst case is a duplicate track() call, undone by closing a pull request nobody merged.
you ▸ (what we can write, and what we can watch)
We write tracking for Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment. To also watch those numbers and tell you when one stops, we need to read them, and today that means PostHog — or our own ingest, which is included.
the rest of the cli
Six commands, and two of them are above: test and audit. These are the other four, plus the one endpoint that needs no account at all. Which of them needs a key is written on each row, because finding that out at the prompt is a bad way to learn it.
$ GET /api/readability?url=https://yoursite.com
Checks whether an AI crawler can physically read a page. GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot and Google-Extended run no JavaScript, so a client-rendered app that ranks fine on Google can be a blank page to them. Free, no auth, stores nothing.
$ npx smolanalytics plan check
Fails CI when an event your tracking plan declares stops firing, against your instance and key.
$ npx smolanalytics connect
Writes the MCP entry into whichever of Cursor, Claude Code, VS Code, Windsurf, Claude Desktop and Cline it finds installed. One organization token operates every project.
$ npx smolanalytics init
Wires the tracker into the app in this directory, for the engine half. It works out what your project is, names the one file it will edit before touching anything, and refuses to write a placeholder key. Testing needs none of it.
what runs without you
Everything above is something you invoke. These are not, and they need events rather than a repo. The worst case of each one is written beside it.
you ▸ (a day after you merge, a comment appears on the PR)
Your merge writes a deploy marker. A cron waits 20 hours for a real after-window, scores every candidate metric in your own events, and posts before and after per day onto the pull request. Significance beats magnitude, so a proven 8% outranks an unproven 40%. No computable movement, no comment. Claimed per commit before it posts, so a double-fired cron never comments twice.
you ▸ (a flag breaches its guardrail: it gets pulled)
A flag whose guardrail fails twice, past a warmup and with a gap between the two failures, is switched off without waiting for you. The reason is written onto your own event timeline as an event you can query.
you ▸ (every event, on a schedule, ranked by what it costs)
investigate
Every event you track is tested for step changes, not just the four you would have thought to check. Ranked by cost in dollars when the metric carries an amount and in people when it does not, with the segment carrying the change and the correlating deploy named. Benjamini-Hochberg correction, so sweeping many metrics cannot manufacture a discovery.
you ▸ I shipped a fix for that
mark_finding_acted
Mark it acted, from this tool, the terminal desk, or POST /v1/findings/acted, and it keeps watching. It upgrades to verified only if the metric recovered after the date you acted. Acted and still down says exactly that. A metric that recovered on its own retires as recovered rather than fixed.
you ▸ (and on a quiet week, nothing arrives)
Silence is the designed output. No comment on a PR with no computable movement, findings under 20 samples suppressed as noise, anything under 100 says small sample out loud, and a product below the detection floor gets a note saying what could not have been seen at that volume.
your agent
If the app you measure is itself an AI agent, the same engine computes its health from the same events. The one thing counting can't do is read what people said, so your own model does that part over MCP, and its inferences stay visibly separate from the computed counts.
you ▸ which tool is slowest, and how often does it error?
agent_tools
Per-tool health from agent_tool_call events: calls, error rate, latency p50/p90/p99, and the client split (cursor, claude, copilot, windsurf).
you ▸ what are my agent's errors, grouped by type?
agent_errors
Error taxonomy over the failing tool calls, grouped by error_type (timeout, rate_limit, bad_args), optionally scoped to one tool.
you ▸ what is my re-ask rate? do conversations resolve?
agent_conversations
Turns per conversation, re-ask rate, abandon rate, time-to-first-token, and resolution rate, computed only when you send a resolved bool, never invented.
you ▸ read a sample of my conversations so you can label them
sample_conversations
Hands your own model whole conversations, oldest turn first. Capped and deterministic: same data in, same sample out. A turn only carries text if your app sent a text property.
you ▸ label these by intent, sentiment and frustration
label_conversation
Your model writes its inference back as a new agent_label event; nothing already recorded is ever mutated. Naming the model is required, because a label is an inference, not a measurement.
you ▸ what are people actually asking my agent about?
agent_labels
Conversation counts per value of one label. Every result names the labeling model plus how many conversations are labeled and unlabeled. Nothing labeled yet gets an honest empty answer, never a fabricated split.
what it found
The findings reach you three ways: npx smolanalytics desk in the terminal, the tools below from your editor, and the morning email. There is no dashboard on your instance to go and look at, on purpose.
$ npx smolanalytics desk
What the investigation found, in the terminal you are already in, against your own instance and key. There is no screen on the instance to open instead: it is a data layer, and this is one of the three ways the findings reach you, beside the MCP tools and the morning email.
you ▸ what should I fix first, and what is it costing me?
investigate
The whole investigation in one call: the most expensive finding, the queue with its status chips (needs you, watch, fix first, recovered, verified, acted, auto-reverted), the quarter's movements with Benjamini-Hochberg correction, and below-detection-floor notes when a product is too small to call.
you ▸ when did this number change, and what shipped near it?
explain_change
The change day, the size, the segment carrying it, and the deploy markers that landed around it. Correlation, and the copy says so every time.
you ▸ what is that drop actually costing me?
Send an amount on your own track() call and findings on that metric are priced in dollars from the next event onward. Stripe, Lemon Squeezy, Polar and Dodo webhooks reach /v1/revenue/{provider} and do the same. Metrics without an amount are never given an invented price.
you ▸ show me the exact rows behind that number
rows_behind
The finding recomputed down to the rows it was derived from. A number you can open is debuggable; a number you cannot is a claim.
you ▸ what would this have told me last quarter?
backtest
Replay your history through the same investigator with the clock moved back: every finding dated by the day it would first have surfaced, with the detection lag printed. Each sweep sees only events strictly before its own date.
you ▸ how's it going? what's broken?
whats_notable
24h drops and spikes, the biggest drop-off in your auto-detected journey, the worst-converting segment through it, the week-over-week headline, a retention read.
you ▸ can I trust a verdict on 12 users?
No, and it won't give you one: findings under 20 samples are suppressed as noise, and anything under 100 says small sample out loud.
you ▸ (a daily email, no ask needed)
A daily job pulls each instance's own brief and emails every member of your org, each finding tagged [verified], [acted], [recovered] or [needs you]. Point a Slack or Discord webhook at it and it lands there instead.
you ▸ alert me if signups drop below 10 a day
create_alert
A threshold on a rolling-window count, checked every 5 minutes, fired to your webhooks. Last-checked value and last fire are inspectable.
you ▸ send alerts to Slack
add_webhook
Slack URLs are auto-detected. Everything else gets HMAC-SHA256-signed JSON (X-Smolanalytics-Signature) so you can verify it's really us.
you ▸ prove the Slack hook works
test_webhook
A real delivery down the exact same path alerts and the digest use, reporting the HTTP status back.
every tool, by category
One connection: the organization token at smolanalytics.com/api/mcp operates every project, and the read key stays server-side. Paste it into Claude Desktop, Claude Code, Cursor, Windsurf, VS Code / Copilot or Cline. Each tool is written as the question that invokes it.
instrumentation / plan
9 tools
propose_instrumentation
you ▸ read my repo and tell me exactly what to track
verify_instrumentation
you ▸ prove each event is wired and firing
suggest_instrumentation_fix
you ▸ this event isn't arriving, fix it
regenerate_plan_from_code
you ▸ rebuild the tracking plan from my track() calls
define_event
you ▸ name a business event from clicks I already capture
list_defined_events
you ▸ what defined events exist?
delete_defined_event
you ▸ remove that defined event
instrumentation_coverage
you ▸ what does my product do that nothing measures?
event_source
you ▸ where is this event fired from, and is it arriving?
deploys
4 tools
record_deploy
you ▸ mark that I shipped v2.1 just now
list_deploys
you ▸ what did I ship, and when?
delete_deploy
you ▸ remove that mis-recorded deploy
deploy_impact
you ▸ did last night's deploy move signups?
the prompts
Built into the MCP server. Instead of asking ten questions in a row, you invoke one name and your model runs the whole routine against your real code and your real tracking.
/instrument-my-app
pick events, wire tracking, set the plan, verify with instrumentation_health, set a drop alert
/did-my-deploy-break-anything
before/after on your last ship: which events moved, what regressed, what to roll back
what fails your build
$ npx smolanalytics test --suite tests/ --url "$URL" --comment
The suite, on the pull request, from a workflow file you copy in. Exit 1 is a test failing, exit 2 is our runner failing, and the template starts with continue-on-error so the comment lands before the gate does.
you ▸ (no ask needed: it's a test in our CI)
A test in our build asserts that the answer over MCP equals the answer the HTTP report returns. If the two ever diverge, our build fails. You can run the same comparison yourself on the proof page.
you ▸ declare the events this app should send
set_tracking_plan
Your tracking plan is a declaration of intent. Without it, an event that stopped firing and an event you retired on purpose are indistinguishable, which is why the unattended tracking restore refuses to act when no plan is declared.
$ npx smolanalytics plan check
Verifies live traffic against the plan and exits 1 on breakage. A deploy that silently kills your signup event fails the build instead of quietly zeroing a funnel.
you ▸ is my tracking broken?
instrumentation_health
Reality vs plan, on demand: which events are arriving, which are missing, which properties are missing, what's arriving unplanned.
privacy
$ smolanalytics.init(key, { anonymous: true })
Cookieless mode: nothing is stored on the device. The server derives a daily-rotating anonymous id, unlinkable across days, so funnels still work within a day and no consent banner is needed. Users who sign in keep full analytics.
you ▸ (safe by default: no ask needed)
Your instance answers no report without its key, its settings pages need the password, and it never phones home.
what we run for you
An isolated instance per project
Every project gets its own Fly app: a dedicated volume, a scale-to-zero machine, its own keys. A breach of one tenant is a breach of one tenant.
Events never touch the control plane
Your traffic goes straight to your instance. The cloud site has six subprocessors, every one named on /security, scrypt-hashed passwords, per-project scoped keys.
Your test runs, on the page that knows who you are
Every verdict the runner posts lands on your project page: the suite, what each test last did, and how many runs needed a model. That page also holds your repository, your keys and your billing, which is why the runs live there and not on the instance.
Exactly two trial emails, ever
One when the trial is ending, one when it ends. Deduped so you never get either twice.
Nothing deleted without notice
If a trial lapses: 7 days of grace, then the machine stops with data kept, an export notice at day 30, destruction at day 37 only if that notice went out at least 7 days earlier.
Teams
Orgs with owner, admin and member roles and invite links. Built for solo devs and small teams, never metered per seat.
fees, plainly
$19 / month
100 tested pull requests included
10c each
every tested pull request past the included ones
$6 / million events
past the 2M fair-use ceiling, metered to the cent
$0.68 of AI a month
past it, AI answers pause until the month resets; every computed report keeps working
One plan. No seats, no per-site fee, no tiers, nothing else. Annual is ten months for twelve. 14 days first, no card. The repo audit and the readability check are free and need no account at all.
not in the box
✕per-pull-request preview environments — today the agent runs against a URL you give it (staging, a deploy preview, or localhost over a tunnel)
✕writing the suite from your codebase — suggest drafts it from the running app instead, and you keep the sentences you agree with
✕mobile apps: the agent drives web browsers only
✕test data seeding — the agent uses whatever state the environment is already in
✕No screen on your instance. It ingests, answers /v1, serves MCP and reads your PostHog; the runs are on your project page and the findings come back through the terminal, your editor or email.
✕Does not replace PostHog, GA4, Amplitude or Plausible. It keeps their tracking correct, and can stand in when you have none.
✕Nothing useful from your events on day one, because a step change needs a baseline. The tests and the repo audit work on day one, because they read your app and your code instead.
✕No video session replay. The session inspector reconstructs a journey from events and never records the DOM.
✕Native mobile gets the event checks, not the readability one.
✕PostHog is the only live vendor read. Mixpanel is a history pull, the rest are file import.
✕No model of ours reads your conversations. Yours does, over MCP, which is why that part is free.
questions
What happens on day one, honestly?
The tests work in the first minute: npx smolanalytics test needs a URL, a sentence and your own Claude API key, and nothing else. npx smolanalytics audit and the readability check answer in seconds too, because they read code and HTML. The event side cannot say anything useful yet: a step change needs a baseline behind it. Bringing PostHog or Mixpanel history across with migrate_from shortens that wait.
Do I write anything in a test file besides the sentence?
A heading and the sentence under it. Optional frontmatter names the test and its criticality. Write what a careful person would check and say what you expect to SEE — name the page, the control and the evidence, because "checkout works" cannot fail usefully. Never put a real password or card number in a file you commit: point the tests at a seeded account on staging and a provider test card.
Do I have to learn all this?
No. A test is a sentence, and the CI half is one workflow file. Everything below that you ask in plain English from your editor, and the prompts exist so a whole routine, like checking what last night's deploy broke, is one name instead of ten questions.
How does the AI not hallucinate my numbers?
It never generates SQL or estimates anything. Every report is computed by the same deterministic engine whether you reach it over HTTP or over MCP, and a test in our build fails if the two ever disagree. You can run that same comparison yourself on the proof page. A test verdict is a different kind of claim: the agent says what it observed in the browser, and a run that observed nothing is reported as errored rather than as a pass or a failure.
If a model labels my conversations, are the numbers still computed?
The two halves are kept apart on purpose. The labels are an inference: your own model reads a sample over MCP and writes what it thinks the intent, sentiment or frustration was back as an append-only event, and every report names the model that wrote them plus how many conversations are labeled and unlabeled. The counts over those labels are computed by the same deterministic engine as everything else.