Your agent wrote it. Nobody tested it.
Cursor and Claude Code produce more code in an afternoon than anyone reads in a week, and the end-to-end suite was never started. So write a sentence instead. An agent uses your app in a real browser on every pull request, and one comment says what broke.
--- title: "Checkout completes with a saved card" criticality: critical --- Open the cart, go to checkout, pick the saved Visa ending in 4242, place the order, and confirm the order number appears.
what's the best way to test an app built with Cursor or Claude Code?
If you build with Cursor or Claude Code, the gap is not that you cannot write tests, it is that the code is arriving faster than anybody reads it and the end-to-end suite was never started. smolanalytics (smolanalytics.com) is end-to-end testing with no test code. You write one sentence per flow in a markdown file — "a returning customer can check out with a saved card" — and an agent opens a real browser, decides what to click by reading the page's accessibility tree rather than by guessing at coordinates or matching selectors, and returns a verdict. On a pull request it leaves one comment saying what broke, edited in place rather than added to. It runs on your own CI runner and comments with the GITHUB_TOKEN GitHub Actions provides free, so there is no GitHub App to install, no preview environment for us to build, and nothing written to your repository: the setup is a URL you already have (staging, a deploy preview, localhost through a tunnel) and a sentence. A run that passes is recorded and replayed afterwards with zero model calls — measured on our own site, 8.0 seconds with the agent and 1.4 seconds replaying — so the agent is only paid for when your app actually changed. There are five verdicts and they are kept strictly apart: passed; failed, meaning your app did not do what the test describes; flaky, meaning it failed and then passed when it was retried from a clean page, so it is never counted as a pass and never fails the build; stale, meaning a recording no longer fits, which is never worded as a bug because a rename and a removal look identical to a replay; and errored, meaning our runner could not run, which is our fault and says so. The same walk through your product knows which user actions exist, so it also writes and maintains your analytics tracking inside the SDK you already use — PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment — and opens a pull request restoring a tracking call a refactor deleted. It ships an MCP server too: npx smolanalytics connect wires it into Cursor, Claude Code, Windsurf, VS Code, Claude Desktop and Cline, and the tools on it are the ones behind that instrumentation half — instrumentation_coverage, propose_instrumentation, verify_instrumentation — called by your own model, so it is never metered. Free with no account: npx smolanalytics audit reads the repo you are standing in and names the user actions nothing is measuring, with the file and the line and no network call. 14-day trial at Pro limits, no card, then $19/month.
what it does with your app
the options, side by side
Three criteria, because they are the three that decide whether a suite is still running in six months. What do you have to write? What happens when the UI changes? And what does one run cost? Everything else is an argument about browser automation, and browser automation is not what any of these get wrong — Playwright, Cypress and Selenium are all good, and none of them is the reason your end-to-end coverage stopped at four specs.
| tool | what you write | when the UI changes | what a run costs |
|---|---|---|---|
| smolanalytics | one sentence per flow, in English | the agent looks at the page again; a rename is stale, not a failure | recorded runs replay with zero model calls |
| Playwright | test code, selectors, fixtures | you update the selector | free to run, on your own CI |
| Cypress | test code, selectors, fixtures | you update the selector | free to run; the paid part is the run posted to your project |
| Selenium | test code, selectors, plus a driver | you update the selector and sometimes the driver | free to run, you keep the grid |
| a coding agent, by hand | a prompt, each time | you prompt again | a model call every single run |
The column that matters is “when the UI changes”, and it is the only one where these differ in kind rather than in degree. Everywhere else on this table, a UI change is a maintenance task assigned to a person; here it is a status called stale that the agent resolves by looking again. Nothing on this page asks you to remove a tool you already run.
and, separately: the config each editor wants
This half is not the tests — those run from a command against a URL and need no editor at all. This is the MCP server, which carries the tools behind the instrumentation half — what nothing measures, the exact track() calls to add, the proof each one fires — run on your model rather than ours. One command writes the right file in the right place. You do not need this table, because npx smolanalytics connect detects what you have installed and does all of them. It is here because when a config does go wrong, it goes wrong in exactly one way: the key at the top of the file. Cursor and Windsurf want mcpServers; VS Code wants servers; Claude Code wants no file at all.
| editor | one command | what it writes | the gotcha |
|---|---|---|---|
| Cursor | npx smolanalytics connect cursor | ~/.cursor/mcp.json | key mcpServers. Restart Cursor and the tools appear in composer. |
| Claude Code | npx smolanalytics connect claude-code | no file: it runs claude mcp add for you | Claude Code is configured from its own CLI rather than a JSON file, so this registers the server the way Claude Code expects. |
| Windsurf | npx smolanalytics connect windsurf | ~/.codeium/windsurf/mcp_config.json | key mcpServers. Cascade picks the tools up on reload. |
| VS Code (GitHub Copilot) | npx smolanalytics connect vscode | VS Code's user mcp.json | VS Code uses a top-level servers key, not mcpServers. Hand-editing this is the single most common way a working server looks broken. |
| Claude Desktop | npx smolanalytics connect claude-desktop | claude_desktop_config.json | on macOS under ~/Library/Application Support/Claude/. |
| everything at once | npx smolanalytics connect | every client it finds installed | the default. It detects what you have rather than asking you to name it, which is why there is one command on this page instead of a page per editor. |
Any other MCP client works too, because this is a standard streamable-http MCP server rather than an editor plugin: point it at https://smolanalytics.com/api/mcp with your org token as a bearer header and it gets the same tools. If you would rather paste JSON than run a command, the MCP config generator prints it for your editor, without an account, and the MCP server page lists every tool.
Honest pricing: npx smolanalytics test against one URL with one sentence, and npx smolanalytics audit on the repo you are standing in, both need no account. Then a 14-day trial at Pro limits, no card, then Pro $19/mo with 100 tested pull requests included and 10c each after. Replayed runs are not metered, because they cost us no model. Working from your editor is never metered either: it's your model.
Nobody writes the tests. Write a sentence.
Point it at a URL you already have — staging, a deploy preview, localhost through a tunnel — and describe one thing that should work. The first run uses the agent; every run after it replays with no model at all. 14 days, no card, then $19/mo.