smolanalytics
menu
log inStart trial
the mcp server

Your agent reads what broke, and fixes the tracking.

One setup for Cursor, Claude Code, VS Code or Windsurf. Your agent runs the end-to-end suite against the app it just changed and reads back a verdict per test, in the same English the test was written in. Then it writes the tracking your repo is missing in the SDK it already uses, and proves it fires. Your own model does the work, so that part costs nothing.

14 days, no card. Then $19/mo with 100 tested pull requests, 10c each after. Keep the analytics you already have.
smolanalytics runs end-to-end tests that have no test code, and an MCP server is how the agent in your editor works with it, for Cursor, Claude Code, VS Code or Windsurf. There are three jobs. Running the suite is a local server, npx smolanalytics mcp, speaking MCP over stdio because the app under test is running on your machine: run_tests drives a real browser through the accessibility tree and comes back with a verdict per test — passed, failed (the app did not do what your sentence describes), stale (a recording stopped fitting, which is never a failure), errored (our runner, never your app) or flaky (it failed and then passed on a retry, so it is never counted as a pass); can_i_ship gives the same run with the gaps foregrounded, the flows that were not actually verified; list_tests says what is already covered. Explaining a failure is the same reply: failures that share a cause are grouped so twelve red tests read as one thing, and on a pull request the comment names the changed files it can connect to the test with the evidence, or says nothing. Writing and keeping the tracking is the remote server, the org connection at smolanalytics.com/api/mcp over HTTP: instrumentation_coverage names the user actions nothing measures, propose_instrumentation hands back the exact track() call for each site in whichever SDK the repo already runs, verify_instrumentation answers FIRING, WIRED or MISSING per event, suggest_instrumentation_fix finds where an event that went quiet was tracked, and record_deploy with deploy_impact says whether a ship moved an event. It is bring-your-own-model: the AI in your editor does the work, so that part is free and unmetered. Keep PostHog, GA4, Mixpanel or Plausible exactly where they are. A passing run is recorded and replayed with zero model calls, the agent wakes only when the recording stops fitting, and the local server never comments, shares or posts anywhere — the CLI does that when a person runs it. 14 days, no card. Then $19/mo with 100 tested pull requests, 10c each after. Keep the analytics you already have.

the three jobs, and which server each one is

An agent connected to this does three things, and it is worth being exact about where each of them runs, because two of them run on your machine and one of them cannot.

  • 1It runs the suite from the editor. This is the local server, npx smolanalytics mcp, MCP over stdio. run_tests takes the URL your app is running on, drives a real browser through the accessibility tree, and comes back with a verdict per test: passed, failed (the app did not do what your sentence describes), stale (a recording stopped fitting — never a failure), errored (our runner, never your app) or flaky (failed then passed on a retry, so never counted as a pass). can_i_ship is the same run with the gaps foregrounded — the flows that were not actually verified — and list_tests says what is already covered. A passing run is recorded and replayed with zero model calls. In CI nobody runs it: a workflow file does, on every pull request, and posts one comment.
  • 2It explains a failure instead of listing twelve. Same reply. Failures that share a cause — the same changed file, the same control, the same path — are grouped, so an agent reading the result fixes one thing rather than eleven copies of it. On a pull request the comment goes one step further and names the changed files it can connect to the failing test, with the string or path that connects them; a file it cannot connect is never mentioned, and zero matches means no blame at all.
  • 3It writes and keeps the tracking. This is the remote server, the org connection at smolanalytics.com/api/mcp over HTTP, because proving an event fired needs the events. instrumentation_coverage names the user actions nothing measures; propose_instrumentation hands back the exact call for each site in whichever SDK the repo already runs (Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment), and your agent applies it with its own editor; verify_instrumentation answers FIRING, WIRED or MISSING per event; and suggest_instrumentation_fix finds where an event that went quiet was tracked. When an event in your tracking plan stops arriving but the traffic behind it doesn't, smolanalytics finds the commit that deleted the track() call and opens a pull request putting it back.
in your editor — the shape of it, not a captured run
> run the checkout tests against localhost and tell me what broke
  ran 5 tests, 4 passed. one failed:
  "an expired card is refused with a message that says so" — the page said
  "something went wrong" instead of naming the card.
> anything in this repo that nothing is measuring?
  instrumentation_coverage: 3 actions with no event — payment, invite, delete

why the tests are a local server

A test drives a browser against an app that is running on your laptop, on your branch, with your model key. No hosted endpoint can reach that, so the testing half speaks MCP over stdio and runs in the same place the code does. Four rules shape it, and they are the reason it is a small server:

  • 1It never softens a verdict. Every tool returns the real status word, stale, errored and flaky included. An agent that reads "passed" for a stale recording tells its user the feature works, and that is worse than returning nothing.
  • 2It never writes anywhere public. No pull request comment, no share link, no post to a project. An agent calling a tool is exploring; publishing that to a team channel is a side effect nobody asked for. The CLI still does all of it when a person runs it.
  • 3It never spends without saying so. A run costs model calls only when a recording is missing or stale, every reply carries the cost line, and a ceiling on model calls is honoured per run.
  • 4It only runs what the change could have broken, if you ask. Pass a git ref as since and the run is limited to the tests that change could have touched, with what was skipped written out. Anything it cannot rule out still runs.
in your editor
> can I ship this?
  6 passed, 1 stale, 1 skipped by since.
  not verified: "a returning customer can check out with a saved card" — the
  recording stopped fitting and was not re-run. run it before you merge.

what you keep

Everything. The agent writes the tracking calls inside whichever SDK your repo already runs — Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment — and never adds a second one beside a working one, because two SDKs on one action double-count it. PostHog's collector keeps collecting, GA4 stays on your page, Plausible keeps counting. The tests are markdown files in your repo and the recordings sit beside them, so if this has not earned its place in a fortnight you delete a folder and you are exactly where you started.

To be exact about the boundary, because it matters: we write tracking for more vendors than we can read. Telling you an event stopped firing needs the numbers, and which connection can be read is written out on integrations.

what it is

smolanalytics runs end-to-end tests that have no test code: you write a sentence describing what should work, and an agent does it in a real browser against a URL you already have. The same walk through your product is what knows which user actions exist, which is why the same tool keeps your analytics tracking correct.

Both halves speak the Model Context Protocol, the open standard for connecting AI assistants to external tools — so the model in your editor runs the suite, reads a verdict, writes a track() call or checks whether it is firing the same way it calls any other MCP tool. The instrumentation tools come with two named prompts, whole routines under one name: instrument-my-app and did-my-deploy-break-anything.

Two things need no server and no account. One test — npx smolanalytics test --url https://yourapp.com --test "the pricing page shows a monthly price" — against a URL you already have, and npx smolanalytics audit, which reads the repo you are standing in and names the user actions nothing measures, with no network call on that path at all. What each surface is for is spelled out on how it works. The setup for both halves, including the workflow file that runs the suite on every pull request, is in the docs.

why a verdict here is worth reading

An agent that just edited your app has no way to find out whether the app still works, and an agent that guesses is worse than one that does not know. Three things make the answer trustworthy:

  • 1It reads the page, it does not squint at it. The agent works from the accessibility tree — the role, name and state of everything on the page — and picks an element, so the click that follows is a real locator with actionability checks. Nothing guesses a coordinate off a screenshot, which is how a run clicks the wrong thing and then blames the wrong feature.
  • 2Five verdicts, kept apart on purpose. A failure is a bug report: the page, the control, and what the page said instead. Stale, errored and flaky each get their own word and are never counted as a pass or dressed up as a failure, because a suite that cries wolf about a renamed button is a suite nobody reads by month two.
  • 3The tracking answer is checked, not guessed. verify_instrumentation says FIRING only when the event is arriving, WIRED when the call exists in the code but nothing has fired yet, MISSING when neither, per event. Your own model phrases the reply; the status comes from the code and the events.

the tracking tools, by name

Nine instrumentation tools and four deploy tools, and every one is written as the question that invokes it on every feature. The ones worth knowing before you connect:

  • 1regenerate_plan_from_code rebuilds the tracking plan from the track() calls the repo actually contains, so the plan describes the code and not somebody's memory of it. define_event, list_defined_events and delete_defined_event manage the business events named from clicks already captured.
  • 2event_source names the file and line an event fires from, and whether it is arriving. A count that dropped is a symptom; a missing call site is a cause.
  • 3record_deploy marks that you shipped, list_deploys and delete_deploy keep that list honest, and deploy_impact answers whether last night's deploy moved an event — the mean daily count in the days after against the days before, from your own events, and it says correlation, not proof, every time. A deploy that silently kills your signup event is a tracking bug, and this is where it shows up first.
in your editor
> is every event in my plan actually arriving?
  verify_instrumentation: 11 FIRING · 1 WIRED (invite_sent, no traffic yet) · 1 MISSING (plan_upgraded)
> signup stopped arriving — where was it tracked?
  suggest_instrumentation_fix: src/app/signup/page.tsx:41 — the submit handler, no track() call.
  insert: track("signup", { plan })

how to connect it

Two entries in your editor, because the two halves run in different places. The tests are a process your editor starts in the repo you have open:

the testing half, local (Claude Code shown)
claude mcp add smolanalytics-tests -- npx smolanalytics mcp
# then: "run the tests against http://localhost:3000 and tell me what broke"
#       "can I ship this?"

The tracking is one connection, not one per project. Settings hands you an organization API token pointed at smolanalytics.com/api/mcp, and npx smolanalytics connect writes it into whichever of Cursor, Claude Code, VS Code, Windsurf, Claude Desktop and Cline it finds installed. From there your agent works on any project by passing project="<name>" to any tool, which routes to that project's own instance for you. The secret read key stays server-side, so you never handle a key, and the token persists, set it up once:

the tracking half, one org connection (Claude Code shown)
claude mcp add --transport http smolanalytics https://smolanalytics.com/api/mcp \
  --header "Authorization: Bearer <your-org-token>"
# then: "write the tracking this repo is missing, in the SDK it already uses" or
#       "is every event in my plan actually arriving?"

Scope the token to a single project with ?project=<id> or to reads only with ?read_only=true.

The tools are identical in every editor. Exact per-editor config is in the docs, and the Cursor / Claude walkthrough is on the Cursor page.

what you can type

The work arrives while you write code, not while you stare at a dashboard. So it happens where you already are — the first three lines run against the app you have open, the rest change your repo:

run the checkout tests against localhost and tell me what brokerun_tests
can I ship this?can_i_ship
what does this project already verify?list_tests
write the tracking this repo is missing, in the SDK it already usespropose_instrumentation
is every event in my plan actually arriving?verify_instrumentation
signup stopped arriving — where was it tracked?suggest_instrumentation_fix
did last night's deploy move signups?deploy_impact

Your agent picks the tool by reading its description, the way it reads any other, which is why the descriptions are written for an agent and not for a listing. The full set, each one as the question that invokes it, is on every feature.

trying it without an account

Three things need no account at all, and one of them is the local server. One test — npx smolanalytics test --url https://yourapp.com --test "the pricing page shows a monthly price" — against a URL you already have; npx smolanalytics audit, which reads the repo you are standing in and names the user actions nothing measures, with no network call on that path at all; and npx smolanalytics mcp, which needs a tests/ folder and your own model key and nothing of ours. npx smolanalytics guard is free in the same way: it names the env vars this repo reads that nothing declares, so a deploy does not start and then fail on the line that reads one.

The org connection is the part that needs an account, because the tracking tools need a project to verify events against. When you want one, the 14-day trial provisions it in about a minute, no card.

questions

We already have Playwright tests. Why would we pay for this?
Keep them. The question is not whether you can write end-to-end tests, it is whether anyone is still maintaining them in six months. This is for the flows nobody got round to covering, and for the ones that go red every time a button is renamed. When your UI changes, there is no selector to update — the agent looks at the page again and works it out.
Can my agent run the tests through MCP?
Yes, and it runs them locally. npx smolanalytics mcp is a server that speaks MCP over stdio, so it runs in the same place the code does: run_tests takes the URL your app is running on and returns the verdict for every test in the suite, and can_i_ship returns the same run with what was not checked written out — recordings that went stale, tests that were flaky and prove nothing, tests skipped by since, runs that errored on our side. It costs model calls only when a recording is missing or stale, every reply carries the cost line, and it never comments on a pull request, publishes a share link or posts to a project. In CI nobody runs anything: the workflow file does it on every pull request and posts one comment.
Why is that server local, when the other one is not?
Because a test drives a browser against an app that is running on your laptop, on your branch, with your model key, and no hosted endpoint can reach that. The instrumentation half is the opposite: proving an event fired needs the events, and those live on a server. So the tests are a stdio server your editor starts as a process, and the tracking is the org connection at smolanalytics.com/api/mcp over HTTP. Two entries in your editor config, and each one is the only kind that could do its job.
Do I have to replace the analytics I already run?
No. Keep Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment exactly where they are; the agent writes the tracking calls inside whichever of them your repo already runs, and never adds a second SDK beside a working one. If you have no analytics at all, the same org connection can provision our own ingest for a project, but that is optional and not what this page is about.
PostHog ships an MCP server too, and it is free. Why this one?
Theirs is real and it is good, and if you are all-in on PostHog you should use it. The distinction is not quality, it is what the connection is for: theirs reads PostHog. This one is attached to a test runner, so the agent on the other end of it is doing work on your product — running the suite, reading what broke, putting back a tracking call a refactor deleted — and the SDK it writes into is whichever one your repo already runs, PostHog included. A server owned by a collector will never make your Amplitude or your GA4 better. If you are not already on PostHog, free is not a price you can pay, because the price is a migration.
What should I run first on the trial?
One test. npx smolanalytics test --url <your staging URL> --test "<one sentence>" needs no account at all, so run it before you sign up. If the tests/ folder is empty, npx smolanalytics suggest --url <your app> walks the running app in a real browser and drafts the flows worth testing as markdown files, quoting what it saw; delete the ones you disagree with. Then npx smolanalytics audit on the repo, free and with no network call, to see what your product does that nothing is measuring. After that the editor path: connect, and ask your agent to write the tracking the audit named.
What is an MCP server, in this product?
A backend that speaks the Model Context Protocol, the open standard for connecting AI assistants to external tools. Here it is two of them. The local one turns the test runner into tools your agent can call — run the suite, tell me whether I can ship, what is already covered — so an agent that just edited the app can find out whether the app still works, in the same English the test was written in. The remote one carries the instrumentation tools, so the same agent can write and repair the tracking calls for the actions the tests walk through.
Whose AI runs it, and what does it cost?
Your own. smolanalytics is bring-your-own-model: the AI already in your editor (your Cursor, Claude, or VS Code subscription) does the asking and the reading, and the test runner uses your own Claude API key. There is no separate AI bill from us. You pay for the plan, and calling the tools is never metered.
How do I connect my editor to it?
Two entries. The tests: add npx smolanalytics mcp to your editor as a stdio server, and it runs in the repo you have open. The tracking: npx smolanalytics connect writes the org connection into whichever of Cursor, Claude Code, VS Code, Windsurf, Claude Desktop and Cline it finds installed, pointed at smolanalytics.com/api/mcp with the organization API token from Settings. That one connection operates every project — pass project="<name>" to any tool and it routes to that project's own instance, with the read key kept server-side so you never handle it. Scope it to one project with ?project=<id> or to reads only with ?read_only=true. Set both up once; the token persists.
Am I locked in?
No. The tests are markdown files in your repo and the recordings sit beside them in .smolanalytics/recordings, so there is nothing of ours to export. Every project's instance is isolated, and nothing phones home. The trial is 14 days with no card; after that the meter is the tested pull request ($19/mo with 100 included, then 10c each), and the MCP surface is in every plan, never metered separately.
What can I actually ask it?
Two kinds of thing. About the app you just changed: "run the checkout tests against localhost and tell me what broke", "can I ship this?", "what does this project already verify?". About the tracking: "write the tracking this repo is missing, in the SDK it already uses", "is every event in my plan actually arriving?", "signup stopped arriving — where was it tracked?", "rebuild the tracking plan from the code that implements it", "did last night's deploy move signups?". The first group is the local server, the second is the org connection, and your agent picks the tool by reading its description the way it reads any other.
keep reading
Connect it to your editor
14 days, no card. Then $19/mo with 100 tested pull requests, 10c each after. Keep the analytics you already have.