One sentence now, a verdict on the pull request.
End-to-end tests are the ones everybody agrees they should have and nobody keeps. You write them, a button moves, forty of them go red for no reason, and within a month the suite is muted and the next regression ships to customers instead.
So you write a sentence instead. An agent uses your app in a real browser, and one comment on the pull request says what broke. The same walk through your product also keeps your tracking correct in whatever analytics you already run.
how a test run works
Four steps, each one a mechanism you can check.
One heading and one sentence in a markdown file, and that is the whole test. The heading is the test's identity and the recording's filename. There is nothing else in the file: no selectors, no page objects, no fixtures, and therefore nothing that goes red when a button is renamed.
It works from the accessibility tree — the role, name, value and state of everything on the page — and picks an element, so the click that follows is a real locator with actionability checks. Nothing here guesses a coordinate off a screenshot, which is how a run clicks the wrong thing and then blames the wrong feature.
The steps that worked are compiled into a plan, and the next run replays it with no model call at all. Measured on this site: 8.0s for the first run, 1.4s for the second — one flow on one machine, not a benchmark; your app and your runner will give you your own numbers. The agent wakes again only when the recording stops fitting the app — which is exactly when judgement is worth paying for.
A row per test with the verdict, whether it replayed or woke the agent, and how long it took; each failure written out in full underneath. It is the same comment on every push rather than a new one, and it is posted by your own CI job with the token GitHub Actions already gives it.
**1 failed · 1 passed · 1 stale** in 24.0s Against `https://your-app-git-pr-482.vercel.app` · [run log](…) | | test | how | time | | --- | --- | --- | --- | | **fail** | An expired card is refused with a message that says so | agent | 21.4s | | stale | The cart survives a reload | recording stopped fitting | 1.2s | | pass | A shopper can add an item to the cart | replayed, no model calls | 1.4s | **An expired card is refused with a message that says so** — `tests/checkout.md` > The page showed "Something went wrong" and stayed on /checkout. The test asks for a > message naming the card as expired, and nothing on the page said so. --- 1 of 3 ran from a recording, with no model calls. Stale is not a failure: a recorded run stopped fitting the app, which a replay cannot tell apart from a rename.
The comment is written by your own CI job. Copy the workflow template into .github/workflows/e2e.yml, add ANTHROPIC_API_KEY to that repository's secrets, and keep one of its three preview-URL steps. It asks for contents: read and pull-requests: write and nothing else, it skips forks and dependabot because Actions withholds secrets from them, and week one it runs with continue-on-error so a new tool reports before it starts blocking merges.
where to start
Pick the row that sounds like you. The first one needs no account and takes about a minute.
I want to know whether my app still works
- 1Write one sentenceIn a markdown file, one heading per test and the sentence under it: "from the storefront, open the first product, add it to the cart, and check the cart shows one line at the price the product page listed." Say what you expect to SEE, because "checkout works" cannot fail usefully. No selectors, no page objects, no fixtures.
- 2Point it at a URL you already haveStaging, a deploy preview, or localhost through a tunnel. The agent opens a real browser, reads the page through its accessibility tree, does the thing and tells you what it observed. No account for this, no GitHub App, no preview environment to build, and nothing written to your repository. It runs on your own Claude API key.
- 3Put the same command in CICopy the workflow template to .github/workflows/e2e.yml, add ANTHROPIC_API_KEY to that repository's secrets, and keep one of its three preview-URL steps. From then on the suite runs on your own Actions runner and one comment on the pull request says what broke, edited in place on every push.
npx smolanalytics test --url https://yourapp.com --test "the pricing page shows a monthly price" npx smolanalytics test --suite tests/ --url "$URL" --comment # in CI
I already have PostHog, Mixpanel, GA4 or Amplitude
- 1Run the audit, before you sign up for anythingnpx smolanalytics audit reads the repo you are standing in and names the user actions nothing is measuring, with the file and the line. It counts the tracking you already have, in whichever SDK wrote it, so a working PostHog reads as working. No account, no key, no network call.
- 2Let your agent write the calls in your own SDKPoint us at the repo and the pull request we open writes posthog.capture(), mixpanel.track(), gtag("event", …) — whatever this codebase already runs. We never add a second analytics SDK beside a working one, because two SDKs on one action double-counts it.
- 3Connect the PostHog you already haveIn project setup, paste a read-scoped personal API key. We prove it against a real query before storing it, then read your project as bounded daily aggregates on a schedule. Your events never move, and when a refactor deletes a tracking call we open the pull request putting that exact line back, in the SDK it was deleted from.
npx smolanalytics audit # free, no account, reads whatever analytics you already runhow the tracking is written and proved →
I code in Cursor or Claude Code
- 1Connect oncePaste your organization's MCP token (from Settings) into your editor: one connection operates every project, you pass a project name to reach any of them and the read key stays server-side. Restart the editor.
- 2Ask for the work, not the chartpropose_instrumentation writes the tracking this repo is missing in the SDK it already uses, verify_instrumentation proves each event fires, and investigate returns the whole investigation in one call: findings, causes, costs, quarter movements. Your own model does the talking, so that part costs you nothing and there are no API keys to add here.
- 3Trust the number, then close the loopEvery answer is computed by the same deterministic report the HTTP API serves, so it cannot be hallucinated. When you ship a fix, mark_finding_acted records it, and the finding upgrades to verified when the metric recovers.
claude mcp add --transport http smolanalytics https://smolanalytics.com/api/mcp \ --header "Authorization: Bearer <your-org-token>" # cursor and the rest: same URL, same tokenthe editor-native guide →
who writes the tracking
The same walk through your product is what makes this half possible: an agent that has just used your checkout knows the checkout exists. It writes the track() calls at the right call sites in the SDK you already run — Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment — and proves each one fires with verify_instrumentation. We add no second SDK beside a working one, because two SDKs on one action double-count it. A committed tracking plan is then checked in CI with npx smolanalytics plan check, so a deletion is caught at the pull request instead of six weeks later, and npx smolanalytics audit names what nothing measures before you have signed up for anything. The full loop is on the instrumentation page.
if you run no analytics at all
A project can hold its own tracking: the instance that comes with it ingests events, answers /v1 and serves MCP, with no screen to go and look at. It is behind one door on the setup page and the tests need none of it. If you already run PostHog, GA4 or the like, the two sections above are the product.
who this is for
Not for a data team that wants a distributed event warehouse. That is a different tool.