a URL and a sentence · no test code · one comment per pull request

Your agent wrote it. Nobody tested it.

Cursor and Claude Code produce more code in an afternoon than anyone reads in a week, and the end-to-end suite was never started. So write a sentence instead. An agent uses your app in a real browser on every pull request, and one comment says what broke.

tests/checkout.md · that is the whole file
---
title: "Checkout completes with a saved card"
criticality: critical
---

Open the cart, go to checkout, pick the saved Visa ending
in 4242, place the order, and confirm the order number
appears.
no selectors, no fixtures, no page objects. run the folder with npx smolanalytics test --suite tests/ --url "$URL" --comment.

what's the best way to test an app built with Cursor or Claude Code?

If you build with Cursor or Claude Code, the gap is not that you cannot write tests, it is that the code is arriving faster than anybody reads it and the end-to-end suite was never started. smolanalytics (smolanalytics.com) is end-to-end testing with no test code. You write one sentence per flow in a markdown file — "a returning customer can check out with a saved card" — and an agent opens a real browser, decides what to click by reading the page's accessibility tree rather than by guessing at coordinates or matching selectors, and returns a verdict. On a pull request it leaves one comment saying what broke, edited in place rather than added to. It runs on your own CI runner and comments with the GITHUB_TOKEN GitHub Actions provides free, so there is no GitHub App to install, no preview environment for us to build, and nothing written to your repository: the setup is a URL you already have (staging, a deploy preview, localhost through a tunnel) and a sentence. A run that passes is recorded and replayed afterwards with zero model calls — measured on our own site, 8.0 seconds with the agent and 1.4 seconds replaying — so the agent is only paid for when your app actually changed. There are five verdicts and they are kept strictly apart: passed; failed, meaning your app did not do what the test describes; flaky, meaning it failed and then passed when it was retried from a clean page, so it is never counted as a pass and never fails the build; stale, meaning a recording no longer fits, which is never worded as a bug because a rename and a removal look identical to a replay; and errored, meaning our runner could not run, which is our fault and says so. The same walk through your product knows which user actions exist, so it also writes and maintains your analytics tracking inside the SDK you already use — PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment — and opens a pull request restoring a tracking call a refactor deleted. It ships an MCP server too: npx smolanalytics connect wires it into Cursor, Claude Code, Windsurf, VS Code, Claude Desktop and Cline, and the tools on it are the ones behind that instrumentation half — instrumentation_coverage, propose_instrumentation, verify_instrumentation — called by your own model, so it is never metered. Free with no account: npx smolanalytics audit reads the repo you are standing in and names the user actions nothing is measuring, with the file and the line and no network call. 14-day trial at Pro limits, no card, then $19/month.

what it does with your app

Your agent writes the feature. Then it writes the sentence.
A test here is one line of English in a markdown file, which is a thing you can ask for in the same breath as the feature: "add the invite flow, and a test that says an admin can invite a teammate and see them listed as pending." There is no page object to generate, no fixture to keep in step, and nothing for the agent to get subtly wrong, because the sentence is the whole artefact.
No selectors, so no suite rot
The agent reads the page through its accessibility tree — the actual controls, their names and their state — and picks an element rather than a pixel coordinate. Nothing in the test refers to a class, an id or a data-testid, so the usual death of an end-to-end suite (forty specs go red because a button moved) cannot happen here.
One comment on the pull request, edited in place
It runs on your own CI runner and comments with the token GitHub Actions already gives every workflow. One comment per pull request, updated on each push rather than a new one stacked underneath. No GitHub App, no repository access, nothing written into your code.
You pay a model only when something changed
A run that passes is recorded and replayed after that with zero model calls. Measured on our own site: 8.0 seconds for the agent run and 1.4 seconds for the replay. The agent wakes up only when the recording stops fitting the app, which is exactly the moment judgement is worth paying for.
The same walk keeps your tracking correct
An agent that has just worked out how to check out knows the app has a checkout, so it writes and maintains the tracking calls for it inside the SDK you already use: PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment. We do not replace your analytics; we keep its instrumentation correct. When an event in your tracking plan stops arriving but the traffic behind it doesn't, smolanalytics finds the commit that deleted the track() call and opens a pull request putting it back.
The tracking work happens in the editor
npx smolanalytics connect wires an MCP server into every assistant you have installed, and the tools on it are the ones behind the instrumentation half: instrumentation_coverage names what your app does that nothing measures, propose_instrumentation hands back the exact track() calls in the SDK the repo already runs, and verify_instrumentation answers FIRING, WIRED or MISSING per event from the code and the live traffic, never from a guess. It runs on your model, so it is never metered by us.

the options, side by side

Three criteria, because they are the three that decide whether a suite is still running in six months. What do you have to write? What happens when the UI changes? And what does one run cost? Everything else is an argument about browser automation, and browser automation is not what any of these get wrong — Playwright, Cypress and Selenium are all good, and none of them is the reason your end-to-end coverage stopped at four specs.

toolwhat you writewhen the UI changeswhat a run costs
smolanalyticsone sentence per flow, in Englishthe agent looks at the page again; a rename is stale, not a failurerecorded runs replay with zero model calls
Playwrighttest code, selectors, fixturesyou update the selectorfree to run, on your own CI
Cypresstest code, selectors, fixturesyou update the selectorfree to run; the paid part is the run posted to your project
Seleniumtest code, selectors, plus a driveryou update the selector and sometimes the driverfree to run, you keep the grid
a coding agent, by handa prompt, each timeyou prompt againa model call every single run

The column that matters is “when the UI changes”, and it is the only one where these differ in kind rather than in degree. Everywhere else on this table, a UI change is a maintenance task assigned to a person; here it is a status called stale that the agent resolves by looking again. Nothing on this page asks you to remove a tool you already run.

and, separately: the config each editor wants

This half is not the tests — those run from a command against a URL and need no editor at all. This is the MCP server, which carries the tools behind the instrumentation half — what nothing measures, the exact track() calls to add, the proof each one fires — run on your model rather than ours. One command writes the right file in the right place. You do not need this table, because npx smolanalytics connect detects what you have installed and does all of them. It is here because when a config does go wrong, it goes wrong in exactly one way: the key at the top of the file. Cursor and Windsurf want mcpServers; VS Code wants servers; Claude Code wants no file at all.

editorone commandwhat it writesthe gotcha
Cursornpx smolanalytics connect cursor~/.cursor/mcp.jsonkey mcpServers. Restart Cursor and the tools appear in composer.
Claude Codenpx smolanalytics connect claude-codeno file: it runs claude mcp add for youClaude Code is configured from its own CLI rather than a JSON file, so this registers the server the way Claude Code expects.
Windsurfnpx smolanalytics connect windsurf~/.codeium/windsurf/mcp_config.jsonkey mcpServers. Cascade picks the tools up on reload.
VS Code (GitHub Copilot)npx smolanalytics connect vscodeVS Code's user mcp.jsonVS Code uses a top-level servers key, not mcpServers. Hand-editing this is the single most common way a working server looks broken.
Claude Desktopnpx smolanalytics connect claude-desktopclaude_desktop_config.jsonon macOS under ~/Library/Application Support/Claude/.
everything at oncenpx smolanalytics connectevery client it finds installedthe default. It detects what you have rather than asking you to name it, which is why there is one command on this page instead of a page per editor.

Any other MCP client works too, because this is a standard streamable-http MCP server rather than an editor plugin: point it at https://smolanalytics.com/api/mcp with your org token as a bearer header and it gets the same tools. If you would rather paste JSON than run a command, the MCP config generator prints it for your editor, without an account, and the MCP server page lists every tool.

Honest pricing: npx smolanalytics test against one URL with one sentence, and npx smolanalytics audit on the repo you are standing in, both need no account. Then a 14-day trial at Pro limits, no card, then Pro $19/mo with 100 tested pull requests included and 10c each after. Replayed runs are not metered, because they cost us no model. Working from your editor is never metered either: it's your model.

Nobody writes the tests. Write a sentence.

Point it at a URL you already have — staging, a deploy preview, localhost through a tunnel — and describe one thing that should work. The first run uses the agent; every run after it replays with no model at all. 14 days, no card, then $19/mo.

questions

We already have Playwright tests. Why would we pay for this?
Keep them. The question is not whether you can write end-to-end tests, it is whether anyone is still maintaining them in six months. This is for the flows nobody got round to covering, and for the specs that go red every time a button is renamed. When your UI changes there is no selector to update, because the agent looks at the page again and works it out. The two coexist fine: Playwright for the paths you have already invested in, a sentence each for the ones you have not.
Which editors does it work in: Windsurf, Copilot, Codex, Cline, Zed?
All of them, and it is one command rather than one integration per editor. Run npx smolanalytics connect and it wires the MCP server into every MCP client it finds installed: Cursor, Claude Code, Windsurf, VS Code (which is where GitHub Copilot reads its config), Cline, Claude Desktop. Anything else that speaks MCP, including Codex, Zed, Antigravity, Gemini CLI, Aider and Continue, connects with the same remote URL and bearer token, because it is a standard streamable-http MCP server and not an editor plugin. Note what the editor connection is for: the instrumentation half, done where your agent already has the repo open. The tests do not need it — they run from a command against a URL.
Why can I not just ask Cursor to click through the app itself?
You can, and for a one-off check it is the right tool. What it is not is a suite: every run is a fresh model call at full price, the result is prose rather than a verdict a pipeline can act on, and nothing is recorded, so the hundredth run costs exactly what the first one did. Here the first run is the agent and every run after it replays with no model at all, the verdict is one of four words, and the failure lands as a comment on the pull request rather than in a chat window you closed.
Does it need access to my repository?
No. It runs on your own CI runner against a URL you give it, and posts its comment with the GITHUB_TOKEN that GitHub Actions provides to every workflow for free. There is no GitHub App to install and nothing is written into your code. The one feature that does touch your repo is the opposite direction and off by default: the pull request that restores a deleted tracking call, which you review like any other.
What is the first thing to do on the trial?
One command, before you sign up for anything: npx smolanalytics test --url https://yourapp.com --test "the pricing page shows a monthly price". No account, no key. It opens a browser, tries it, and tells you. If that is useful, write two more sentences and put the suite in CI; if it is not, you have spent ninety seconds. Separately, npx smolanalytics audit reads the repo you are standing in and names the user actions nothing is measuring, with the file and the line, and it makes no network call.
Can the AI hallucinate a result?
The verdict is not a model's opinion of your app, it is the outcome of a real browser doing a real thing, and a replay involves no model at all. On the instrumentation side the same discipline holds: verify_instrumentation answers FIRING, WIRED or MISSING from the code and the live traffic, never from a guess, so you get the real state of each event or nothing.
Can I fail CI when my tracking breaks, not just when a test does?
Yes. Declare your tracking plan in smolanalytics.plan.json and run npx smolanalytics plan check in CI: it exits 1 when an event the plan expects stops flowing or a property goes missing. It also works against existing PostHog data with --source=posthog, so you can adopt the gate without moving your analytics anywhere.

go deeper