smolanalytics
menu
log inStart trial
how it works

One sentence now, a verdict on the pull request.

End-to-end tests are the ones everybody agrees they should have and nobody keeps. You write them, a button moves, forty of them go red for no reason, and within a month the suite is muted and the next regression ships to customers instead.

So you write a sentence instead. An agent uses your app in a real browser, and one comment on the pull request says what broke. The same walk through your product also keeps your tracking correct in whatever analytics you already run.

smolanalytics runs end-to-end tests that have no test code. You write one sentence describing what should work — "a returning customer can check out with a saved card" — and an agent opens a real browser, reads the page through its accessibility tree so it picks an element rather than a pixel, does the thing, and reports what it observed. Setup is a URL you already have (staging, a deploy preview, localhost through a tunnel) and that sentence: no account for the first run, no GitHub App, no preview environment built, and nothing written to your repository. On a pull request the whole suite lands as one comment, edited in place on every push, from a workflow file you copy into your own repository; it runs on your own Actions runner and comments with the GITHUB_TOKEN Actions gives every job. A run that passes is recorded and replayed with no model call at all — measured on this site at 8.0s for the first run and 1.4s for the second — and the agent only wakes again when the recording stops fitting the app. Four verdicts are kept apart on purpose: passed, failed (the app did not do what the sentence describes), stale (a recording no longer fits, which a replay cannot tell from a rename, so it is never reported as a bug) and errored (our runner could not run, never your app). The same walk through your product knows which user actions exist, so it also writes and maintains your analytics tracking inside the SDK you already use — Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment — and when a refactor deletes a tracking call it opens a pull request putting that exact line back. Keep PostHog, Mixpanel, GA4 or Amplitude exactly where they are; if you have no analytics at all, the engine underneath can be it. 14 days, no card. Then $19/mo with 100 tested pull requests, 10c each after.

how a test run works

Four steps, each one a mechanism you can check.

You write the sentence

One heading and one sentence in a markdown file, and that is the whole test. The heading is the test's identity and the recording's filename. There is nothing else in the file: no selectors, no page objects, no fixtures, and therefore nothing that goes red when a button is renamed.

The agent reads the page, it does not squint at it

It works from the accessibility tree — the role, name, value and state of everything on the page — and picks an element, so the click that follows is a real locator with actionability checks. Nothing here guesses a coordinate off a screenshot, which is how a run clicks the wrong thing and then blames the wrong feature.

The run that passes becomes a recording

The steps that worked are compiled into a plan, and the next run replays it with no model call at all. Measured on this site: 8.0s for the first run, 1.4s for the second — one flow on one machine, not a benchmark; your app and your runner will give you your own numbers. The agent wakes again only when the recording stops fitting the app — which is exactly when judgement is worth paying for.

One comment, edited in place

A row per test with the verdict, whether it replayed or woke the agent, and how long it took; each failure written out in full underneath. It is the same comment on every push rather than a new one, and it is posted by your own CI job with the token GitHub Actions already gives it.

what lands on the pull request
**1 failed · 1 passed · 1 stale** in 24.0s

Against `https://your-app-git-pr-482.vercel.app` · [run log](…)

| | test | how | time |
| --- | --- | --- | --- |
| **fail** | An expired card is refused with a message that says so | agent | 21.4s |
| stale | The cart survives a reload | recording stopped fitting | 1.2s |
| pass | A shopper can add an item to the cart | replayed, no model calls | 1.4s |

**An expired card is refused with a message that says so** — `tests/checkout.md`

> The page showed "Something went wrong" and stayed on /checkout. The test asks for a
> message naming the card as expired, and nothing on the page said so.

---
1 of 3 ran from a recording, with no model calls. Stale is not a failure: a recorded run
stopped fitting the app, which a replay cannot tell apart from a rename.

The comment is written by your own CI job. Copy the workflow template into .github/workflows/e2e.yml, add ANTHROPIC_API_KEY to that repository's secrets, and keep one of its three preview-URL steps. It asks for contents: read and pull-requests: write and nothing else, it skips forks and dependabot because Actions withholds secrets from them, and week one it runs with continue-on-error so a new tool reports before it starts blocking merges.

where to start

Pick the row that sounds like you. The first one needs no account and takes about a minute.

path 1

I want to know whether my app still works

  1. 1
    Write one sentence
    In a markdown file, one heading per test and the sentence under it: "from the storefront, open the first product, add it to the cart, and check the cart shows one line at the price the product page listed." Say what you expect to SEE, because "checkout works" cannot fail usefully. No selectors, no page objects, no fixtures.
  2. 2
    Point it at a URL you already have
    Staging, a deploy preview, or localhost through a tunnel. The agent opens a real browser, reads the page through its accessibility tree, does the thing and tells you what it observed. No account for this, no GitHub App, no preview environment to build, and nothing written to your repository. It runs on your own Claude API key.
  3. 3
    Put the same command in CI
    Copy the workflow template to .github/workflows/e2e.yml, add ANTHROPIC_API_KEY to that repository's secrets, and keep one of its three preview-URL steps. From then on the suite runs on your own Actions runner and one comment on the pull request says what broke, edited in place on every push.
npx smolanalytics test --url https://yourapp.com --test "the pricing page shows a monthly price"
npx smolanalytics test --suite tests/ --url "$URL" --comment   # in CI
path 2

I already have PostHog, Mixpanel, GA4 or Amplitude

  1. 1
    Run the audit, before you sign up for anything
    npx smolanalytics audit reads the repo you are standing in and names the user actions nothing is measuring, with the file and the line. It counts the tracking you already have, in whichever SDK wrote it, so a working PostHog reads as working. No account, no key, no network call.
  2. 2
    Let your agent write the calls in your own SDK
    Point us at the repo and the pull request we open writes posthog.capture(), mixpanel.track(), gtag("event", …) — whatever this codebase already runs. We never add a second analytics SDK beside a working one, because two SDKs on one action double-counts it.
  3. 3
    Connect the PostHog you already have
    In project setup, paste a read-scoped personal API key. We prove it against a real query before storing it, then read your project as bounded daily aggregates on a schedule. Your events never move, and when a refactor deletes a tracking call we open the pull request putting that exact line back, in the SDK it was deleted from.
npx smolanalytics audit    # free, no account, reads whatever analytics you already run
how the tracking is written and proved →
path 3

I code in Cursor or Claude Code

  1. 1
    Connect once
    Paste your organization's MCP token (from Settings) into your editor: one connection operates every project, you pass a project name to reach any of them and the read key stays server-side. Restart the editor.
  2. 2
    Ask for the work, not the chart
    propose_instrumentation writes the tracking this repo is missing in the SDK it already uses, verify_instrumentation proves each event fires, and investigate returns the whole investigation in one call: findings, causes, costs, quarter movements. Your own model does the talking, so that part costs you nothing and there are no API keys to add here.
  3. 3
    Trust the number, then close the loop
    Every answer is computed by the same deterministic report the HTTP API serves, so it cannot be hallucinated. When you ship a fix, mark_finding_acted records it, and the finding upgrades to verified when the metric recovers.
claude mcp add --transport http smolanalytics https://smolanalytics.com/api/mcp \
  --header "Authorization: Bearer <your-org-token>"   # cursor and the rest: same URL, same token
the editor-native guide →

who writes the tracking

The same walk through your product is what makes this half possible: an agent that has just used your checkout knows the checkout exists. It writes the track() calls at the right call sites in the SDK you already run — Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment — and proves each one fires with verify_instrumentation. We add no second SDK beside a working one, because two SDKs on one action double-count it. A committed tracking plan is then checked in CI with npx smolanalytics plan check, so a deletion is caught at the pull request instead of six weeks later, and npx smolanalytics audit names what nothing measures before you have signed up for anything. The full loop is on the instrumentation page.

if you run no analytics at all

A project can hold its own tracking: the instance that comes with it ingests events, answers /v1 and serves MCP, with no screen to go and look at. It is behind one door on the setup page and the tests need none of it. If you already run PostHog, GA4 or the like, the two sections above are the product.

who this is for

Indie hackers and solo devs
You ship a lot and nobody owns the test suite, because there is nobody else. A sentence per flow is the amount of testing that actually survives contact with a one-person roadmap.
Teams whose e2e suite is already muted
You have Playwright. Keep it. This is for the flows nobody got round to covering and the ones that go red every time a button moves, where there is no selector to update because there are no selectors.
Agent-native builders
You live in Cursor or Claude Code. Your agent writes the app, writes the tracking, and now checks the app it wrote from the same window.
Anyone whose tracking quietly broke
A refactor deleted a track() call six weeks ago and the funnel has been wrong ever since. That is the second half of this product, and it works whether or not you ever write a test.

Not for a data team that wants a distributed event warehouse. That is a different tool.

questions

What does a test file look like?
A markdown file with a heading and a sentence under it, one heading per test. Optional frontmatter names it and its criticality. Write what a careful person would check and say what you expect to see: name the page, the control and the evidence. Never put a real password or card number in a file you commit — point the tests at a seeded account on staging and a provider test card.
We already have Playwright tests. Why would we pay for this?
Keep them. The question is not whether you can write end-to-end tests, it is whether anyone is still maintaining them in six months. This is for the flows nobody got round to covering, and for the ones that go red every time a button is renamed. When your UI changes, there is no selector to update — the agent looks at the page again and works it out.
What is the difference between failed and stale?
Failed means the app did not do what the sentence describes: that is a bug report and it is the only status that says something is wrong with your product. Stale means a recording stopped fitting the app, and a replay cannot tell a rename from a removal, so it is never coloured or worded as a failure — the agent re-checks it and rewrites the recording. Errored is a third thing again: our runner could not run, and it is labelled as ours.
Do I have to install a GitHub App?
No. The CI half is a workflow file you copy into your own repository. It runs on your Actions runner and comments with the GITHUB_TOKEN Actions hands every job, with permissions of contents: read and pull-requests: write and nothing else. The GitHub App exists for the separate tracking-restore feature and is off by default; testing never needs it.
What does a run cost me in model calls?
The first run of a test uses the agent. Every run after it replays a recording with no model call at all, until the recording stops fitting. In CI the template caches the recordings between runs, because a runner starts empty and without the cache every test on every pull request would be a fresh agent run forever.
What happens after I fix something?
For a test, the next run tells you: a passing run replaces the failure and re-records. For a tracking finding, npx smolanalytics desk marks it acted and then reports whether the metric recovered, instead of leaving you to remember.
Where do the results show up?
The pull request comment is the one nobody has to go and look at, and it is the point. Your project page holds the run history for a connected project: the suite, what each test last did, and how many runs needed a model. In the terminal, the exit code is the contract (0 passed, 1 a test failed, 2 our runner could not finish) and npx smolanalytics desk prints what the analytics half found.
What does it cost to start?
Nothing to try: 14 days at Pro limits, no credit card. After that it is $19/month with 100 tested pull requests included and 10c each after. A tested pull request is a pull request your suite ran against — counted once however many times you push to it, while terminal runs and staging crons are never counted at all. Events are a fair-use ceiling behind that rather than the price. The first test needs no account at all — just your own Claude API key, which is also what the CI workflow uses. Nobody gets charged without entering a card themselves.
What can it tell me on the first day?
About your app, plenty: a test is a sentence and the verdict comes back in under a minute. About your events, nothing, and it will say so rather than invent something — a step change needs a baseline behind it before any movement can be scored. npx smolanalytics audit and the readability check answer immediately because they read code and HTML instead. Bringing PostHog or Mixpanel history across with migrate_from shortens the wait on the event side; it does not remove it.
What can it not do?
It does not build a preview environment for you: it runs against a URL you give it. It does not write your suite from your codebase: npx smolanalytics suggest drafts one from the running app, and you keep the sentences you agree with. It drives web browsers only, so mobile apps are out. It seeds no test data, so it uses whatever state the environment is already in. And it does not replace PostHog, GA4, Amplitude or Plausible: it keeps their tracking correct, and can stand in only if you have none.
Start the 14-day trial
14 days, no card. Then $19/mo with 100 tested pull requests, 10c each after.