smolanalytics
menu
log inStart trial
docs · setup

One sentence, then every pull request.

Start here and you have a test running in about a minute. There is no test code to write: you describe what should work in a sentence, an agent does it in a real browser against a URL you already have, and you get a verdict. Everything after that is the same thing on a schedule you did not have to remember.

The second half of this page is the instrumentation the same agent maintains, inside the analytics you already run. Our own tracking half — the SDK install, the ingest endpoint, the desk — is behind one closed door at the end, still exact, optional, for people who have no analytics at all.

the short answers

Run end-to-end tests for me?
Yes. That is the product
Try it without an account?
Yes — npx smolanalytics test
Write any test code?
No. One sentence per test
Need a GitHub App on my repos?
No
Write anything to my repository?
No
Name the file that broke it?
Yes, when the diff touches what the test used
File the bug in Linear or Jira?
Yes. Failed only, one issue per test
Build a preview environment for me?
No. You give it a URL
Write my first tests by reading the repo?
No. You write the sentences
Test a native mobile app?
No. Web browsers only
Seed test data?
No
Keep the analytics I already have?
Yes. It maintains their tracking
Replace PostHog, GA4 or Amplitude?
No. It runs beside them
Session replay, the video kind?
No

If a No is the thing you needed, say so at contact. Knowing what people are blocked on is how this list gets shorter.

smolanalytics runs end-to-end tests that have no test code. You write one sentence describing what should work, point the runner at a URL you already have — staging, a deploy preview, localhost through a tunnel — and an agent does it in a real browser: npx smolanalytics test --url https://yourapp.com --test "the pricing page shows a monthly price". It reads the page through its accessibility tree, so it picks a button rather than a pixel coordinate. A run that passes is recorded, and the recording replays with no model calls at all; measured against this site, the first run took 8.0s with the agent and the second 1.4s with no model. A folder of markdown files is a suite (one heading is one test, the prose under it is the test), and npx smolanalytics test --suite tests/ --url "$URL" --comment posts one comment on the pull request, edited in place on every push. There are five verdicts and they are never blurred: passed, failed (the app did not do what the sentence describes — a bug report), flaky (it failed and then passed on a retry from a clean page, so it is never counted as a pass and never fails the build), stale (a recording no longer fits, which a replay cannot tell apart from a rename, so it is never red) and errored (our runner could not run — never your app). It needs no GitHub App, no preview environment built for you, and writes nothing to your repository: it runs on your own Actions runner and comments with the GITHUB_TOKEN Actions gives every job. When a test fails on a pull request, the comment names the changed files most likely responsible, each with the string or path that ties it to the test, and says nothing when the diff touches nothing the test used; twelve red tests that share a cause are grouped as one. npx smolanalytics mcp puts the runner inside your editor over MCP, with run_tests, can_i_ship and list_tests. The other half is instrumentation: the same walk through your product knows which user actions exist, so it writes and maintains tracking calls inside the SDK you already use — Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment. When an event in your tracking plan stops arriving but the traffic behind it doesn't, smolanalytics finds the commit that deleted the track() call and opens a pull request putting it back. npx smolanalytics audit reads a repo for free, with no account and no network call, and names the user actions nothing measures; npx smolanalytics plan check fails CI when an event your tracking plan declares stops firing. For anyone with no analytics at all, the project can also hold the events itself: one script tag, or POST /v1/events from anything else. The tests need none of that.

your first test, in about a minute

Two things you already have: a URL that is reachable, and a sentence describing something that should work. Staging, a deploy preview, a localhost through a tunnel — any of them.

terminal
export ANTHROPIC_API_KEY=sk-ant-…

npx smolanalytics test \
  --url https://yourapp.com \
  --test "a visitor can sign up with an email and land on the dashboard"

The agent opens a real browser, reads the page the way a screen reader does — the actual buttons, fields and their state — decides what to click, and checks what it sees against your sentence. It never guesses at pixel coordinates, so it does not click the wrong thing and then blame the wrong feature. Add --headed to watch it happen.

The key is yours and stays yours: the agent runs on your machine, on your own Anthropic key, and replaying a recording needs no key at all. The first run also needs Playwright and Chromium (~50MB), which go in ~/.cache/smolanalytics, once. In a terminal it stops and tells you that before fetching anything — re-run with --yes to let it — and in CI, where there is nobody to answer, it installs. Nothing is written to your project either way.

Say what you expect to see. "checkout works" cannot fail usefully. Name the page, the control and the evidence, the way a careful person would check it by hand.

if you have no tests yet

npx smolanalytics suggest --url https://your-staging-url.com

A real browser walks your app — same-origin pages only, reading, never clicking or submitting — and writes the flows worth testing into tests/*.md, in the format --suite already runs. At most six proposals by default (--max <n>; a small app honestly yields fewer), into tests/ or wherever --out points, and a file that already exists is never overwritten. Delete the ones you disagree with.

It only proposes what it actually saw. Every proposal has to quote text that appears on a page the crawl read; one that cannot is dropped, out loud, with the reason. A model asked “what should an app like this test?” answers from every app it has ever read about — password resets, wishlists, coupon codes — and one such file is worse than an empty folder, because it fails forever against a feature you never built. It reads the running app, not your codebase: the sentences are still yours to keep or throw away.

a folder of sentences

A suite is a directory of markdown files. One heading is one test; the prose under the heading is the whole test. That is the entire format, and it is deliberately readable by someone who does not write code.

tests/checkout.md
# Checkout

## A shopper can add an item to the cart

From the storefront, open the first product, add it to the cart, and check that
the cart shows one line for that product at the price the product page listed.

## An expired card is refused with a message that says so

At checkout, pay with card number 4000 0000 0000 0069, any future-looking name,
CVC 123, and check that the page stays on checkout and says the card is expired
— not a generic "something went wrong".

## An empty cart says it is empty

Open the cart with nothing in it and check it says the cart is empty and offers
a way back to the storefront, rather than showing a total of 0 with a checkout
button.
run the folder
npx smolanalytics test --suite tests/ --url https://staging.example.com

Optional frontmatter at the top of a file (title, criticality) is read and kept out of the sentence. A file that opens with one # title and lists ## flows under it is understood as a title plus its tests, not one enormous test.

Two things worth knowing before you commit the folder. The heading is the test's identity and the recording's filename, so renaming a heading throws its recording away and the agent runs that test fresh. And never put a real password or card number in there — the file is committed. Point the tests at a seeded account on staging and your provider's test card.

on every pull request

One workflow file, and the verdict arrives on the pull request that changed the app without anyone remembering to run anything. Copy this to .github/workflows/e2e.yml, add ANTHROPIC_API_KEY to the repository's Actions secrets (Settings → Secrets and variables → Actions), and keep one of the three "the preview URL" steps.

That is the whole install. No GitHub App across your repositories, no agent pushing a Dockerfile into your code, no environment built before you are allowed to see anything. It runs on your own Actions runner, against a URL your host already builds for the pull request, and comments with the GITHUB_TOKEN Actions hands every job for free.

.github/workflows/e2e.yml — the same file ships in the package as templates/github-action.yml
# .github/workflows/e2e.yml
name: e2e

on:
  pull_request:

# THESE TWO, AND NOTHING ELSE.
#   contents: read        actions/checkout reads tests/ out of this repository.
#   pull-requests: write  the verdict is one comment, edited in place on every push. Without it
#                         everything runs and only the comment fails, with a 403.
# Naming them here also takes everything else away: a compromised step in this workflow cannot
# push code, cut a release, or reach another repository.
permissions:
  contents: read
  pull-requests: write

# One run per pull request. Push twice in a minute and two runs drive two browsers against the
# same preview and race to edit the same comment.
concurrency:
  group: e2e-${{ github.event.pull_request.number }}
  cancel-in-progress: true

jobs:
  e2e:
    runs-on: ubuntu-latest
    # A hung browser otherwise runs to the six-hour default and bills six hours of somebody's
    # minutes for a test that was never going to finish.
    timeout-minutes: 30

    # WHO DOES NOT GET SECRETS, and would therefore run every test against an empty key and then
    # 403 on the comment: a fork's pull request, and dependabot. Skipping says that plainly.
    if: >-
      github.event.pull_request.head.repo.full_name == github.repository
      && github.actor != 'dependabot[bot]'

    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22

      # THE PREVIEW URL — ALL THREE OF THESE ARE OFF, AND FOR MOST REPOSITORIES THAT IS CORRECT.
      # With no step called preview, the URL below resolves to empty, --url "" reads as no --url
      # at all, and the runner finds the deployment itself through the deployments API. Uncomment
      # one only if that does not work for you.

      # (a) VERCEL builds the preview itself; this waits for it and reads its URL. Netlify and
      #     Render have equivalent actions. A third-party action inside a job holding
      #     pull-requests: write is worth a moment's thought, which is the other reason this is
      #     not the default.
      # - name: the preview URL
      #   id: preview
      #   uses: patrickedqvist/wait-for-vercel-preview@v1
      #   with:
      #     token: ${{ secrets.GITHUB_TOKEN }}
      #     max_timeout: 600

      # (b) A URL YOU ALREADY HAVE — staging, or a preview at a predictable address.
      # - name: the preview URL
      #   id: preview
      #   run: echo "url=https://staging.example.com" >> "$GITHUB_OUTPUT"

      # (c) BUILD AND SERVE IT HERE, for a static site with no preview host. Wait for the server:
      #     a test against a port nothing is listening on yet reports a broken app when the app
      #     is fine.
      # - name: the preview URL
      #   id: preview
      #   run: |
      #     npm ci
      #     npm run build
      #     npx --yes serve --listen 4173 dist &
      #     npx --yes wait-on --timeout 60000 http://localhost:4173
      #     echo "url=http://localhost:4173" >> "$GITHUB_OUTPUT"

      # THE RECORDINGS. These two steps are the economics of the whole tool. A CI runner starts
      # empty every time, so without this cache every test on every pull request is a fresh agent
      # run: slower, and paid for per run, forever. The key is per-commit so each run saves a
      # fresh entry; restore-keys takes the newest earlier one.
      - name: restore the recordings
        uses: actions/cache/restore@v4
        with:
          path: .smolanalytics/recordings
          key: smolanalytics-recordings-${{ github.sha }}
          restore-keys: |
            smolanalytics-recordings-

      # Chromium is downloaded each run and deliberately not cached: it is hundreds of megabytes
      # against a repository's 10 GB allowance, so a handful of runs would evict the recordings —
      # the part worth keeping and the part that cannot be re-downloaded.
      - name: e2e
        # WEEK ONE: REPORT, DO NOT BLOCK. A new tool that puts a red X on somebody's pull request
        # before it has earned any trust gets uninstalled instead of read. The comment is still
        # posted and the step is still marked failed in the log. Delete this line once the suite
        # has been right often enough that a failure should stop a merge.
        continue-on-error: true
        env:
          # Referenced as "$URL" below rather than pasted into the shell command, because a value
          # that came from another action must never be interpolated into a script.
          URL: ${{ steps.preview.outputs.url }}
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
          # Provided automatically by Actions. This is the whole of the GitHub access this needs.
          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
        run: npx smolanalytics@latest test --suite tests/ --url "$URL" --comment

      # SAVED AS ITS OWN STEP, AND if: always(). The all-in-one actions/cache saves only on
      # success — invisible until the day you delete the continue-on-error above, and from that
      # day a run with one failing test throws away every recording the agent repaired in it.
      # The failing run is the run that re-recorded the most.
      - name: save the recordings
        if: always()
        uses: actions/cache/save@v4
        with:
          path: .smolanalytics/recordings
          key: smolanalytics-recordings-${{ github.sha }}

The comment is one comment, edited in place: a headline count, a table of every test with how it ran (agent, or replayed with no model calls) and how long it took, and the failure text in full underneath, because the person reading it has not been watching the browser. Twenty pushes do not mean twenty comments. When a test failed, the comment also names the changed files most likely responsible and groups the failures that share a cause — see when a test fails.

On a big folder, --since <ref> runs only the tests this change could have broken and says in the headline what was skipped. Anything it cannot rule out still runs, and a skipped test is never counted as a pass.

Exit codes, if you gate on this later: 0 every test passed, 1 a test failed — the app did not do what the sentence describes, 2 this runner could not finish (no browser, no key, no network, or a recording stopped fitting and there was no key to re-check it with). Never your application.

Set two more environment variables and every verdict also lands in your project, under "Did anything break?" on the project page: SMOLANALYTICS_PROJECT (the project id) and SMOLANALYTICS_WRITE_KEY. The write key is public and can only add, never read, which is why the runner can carry it in CI. If that post fails the verdict above it still stands: telemetry never changes an exit code.

passed, failed, flaky, stale, errored

Five verdicts, and they are never blurred into each other. Blur them and a copy change pages somebody at 2am, or an outage on our side reads to you as your checkout being broken.

verdictwhat it means
passedThe app did what the sentence describes.
failedThe app did not do what the sentence describes. This is a bug report, and the reason is written out in full.
flakyIt failed, then passed when it was retried from a clean page. Nothing about your app changed in between, so this is not a pass and is never counted as one — but it does not fail the build either, because the run did eventually work. The reason carries both halves. A test that keeps doing this is hiding an intermittent bug.
staleA recording no longer fits. A replay cannot tell a rename from a removal, so this is never coloured or worded as a failure. The agent re-checks the test against your sentence and rewrites the recording.
erroredOur runner could not run: no browser, no key, no network. Never your application, and the comment says so in those words.

the recording, and why the second run is nearly free

A run that passes is recorded, and the recording replays with no model calls at all. Measured against this site: the first run took 8.0s with the agent, the second 1.4s with no model. Those are one flow on one machine rather than a promise about yours, but the shape holds — the agent only re-engages when the recording stops fitting the app.

a single test, recorded then replayed
# first run: the agent, and it writes the recording
npx smolanalytics test --url https://yourapp.com \
  --test "the pricing page shows a monthly price" --plan pricing.json

# every run after: the recording, no model, no key needed
npx smolanalytics test --url https://yourapp.com \
  --test "the pricing page shows a monthly price" --plan pricing.json

For a suite, recordings live in .smolanalytics/recordings (change it with --plans <dir>), one JSON file per test, named from the file it lives in and its heading. In CI they must be cached or every pull request pays the agent for every test, forever — which is what the two cache steps in the workflow above are for.

when a test fails: the file to blame, and the bug already written

A failed verdict on a pull request comes with the changed files most likely responsible. No model call: it is the intersection of what the failing run observed — the controls it clicked, the paths it visited, the proof text that vanished — and what the pull request changed, read from git's own diff against the base branch.

a suspect line, as it appears in the comment
src/Checkout.tsx — this PR removed the string 'Proceed to checkout' this test clicks

Two rules shape it. No suspicion without named evidence: every line says which string or path connects the file to the test, and a file it cannot connect is never mentioned. Zero matches means nothing at all: “fourteen files changed, one of them probably did it” is the whole diff ranked as vaguely suspicious, so when no changed file touches anything the test used, no suspect lines appear. A blame hint never changes a verdict or an exit code, and if there is no git, no base ref or the repository is too big to diff, it degrades silently to no suspects.

Twelve red tests are usually one thing. Failures are grouped by what they share: the same suspect file, each blamed independently for its own reason; the same control in both recordings (a real click on the same accessible name, not anybody's description of the failure); or, one level weaker, the same path. Never by the failure prose — two failures that both say “the page showed an error” are not evidence of a shared cause, and grouping on it is how a real second bug gets filed inside a group somebody dismissed as all one thing.

Why it was flaky. A flaky verdict carries both runs — every step, every target, every step's duration — compared step for step, with no model call. When the two differ in a way that names a race in your app, the reason says so with the evidence; when they cannot be told apart it says exactly that, because a wrong diagnosis sends somebody hunting a race that does not exist, or tells them to ignore one that does.

The bug is already written, so it can be filed. A failure here is a sentence naming the page, the control, what was expected and what the page did instead, plus the suspect and its evidence. Set the variables for your tracker and every failed test becomes one issue:

Linear, or Jira — set one pair, in CI
# Linear
SMOLANALYTICS_LINEAR_API_KEY=lin_api_…
SMOLANALYTICS_LINEAR_TEAM_ID=…

# Jira
SMOLANALYTICS_JIRA_URL=https://yourcompany.atlassian.net
SMOLANALYTICS_JIRA_EMAIL=you@yourcompany.com
SMOLANALYTICS_JIRA_API_TOKEN=…
SMOLANALYTICS_JIRA_PROJECT=ENG

One issue per test, however many pushes it fails on: each carries a fingerprint of the test it came from, and an existing open issue with that fingerprint is updated rather than duplicated. Only failed is filed. stale is our recording aging, flaky is a ticket everybody closes as cannot-reproduce, and errored is our runner — filing any of them against your product is a lie about whose fault it is.

Telling somebody who is not looking at the pull request. Set SMOLANALYTICS_SLACK_WEBHOOK and the verdict lands in Slack (SMOLANALYTICS_WEBHOOK posts the same as JSON anywhere else). What is sent is the ship verdict, not a count — “Do not ship this” and the causes, rather than “3 failed” — and by default it speaks only when something failed or went unverified, because a channel that says “all good” on every push is a channel that gets muted. The URL is a bearer credential and is read from the environment only, never a flag. Like the tracker and the comment, a delivery that fails can never reach the exit code.

the configuration nothing declares

terminal, or a CI step
npx smolanalytics guard          # this repo; or: npx smolanalytics guard ./api --json

An environment variable the code reads with no fallback and names in no .env.example, README, Dockerfile, compose file or workflow is a deploy that starts and then fails on the line that reads it. This reads the repository it is standing in and lists those: no key, no account, no network, no config.

A read with a fallback is not a requirement. process.env.PORT || 3000 cannot break a deploy and is never reported as if it could, because flagging optional configuration is how a check becomes noise and then gets deleted. --json is for a CI step that wants to act on it.

tests behind a sign-in

Most tests worth writing are behind a login, so the sign-in is recorded and reused exactly the way everything else here is.

export SMOLANALYTICS_LOGIN_EMAIL=qa@yourcompany.com
export SMOLANALYTICS_LOGIN_PASSWORD=...

npx smolanalytics test --suite tests/ --url "$URL" \
  --login "sign in as {{email}} with {{password}}"

The agent signs in once for the whole suite, and every test after that starts already signed in — measured at fifty tests, one sign-in. If the session expires part way through, it signs in again, once, and carries on.

The password is read from the environment and never written down. The recording stores {{password}}, not your password, and resolves it at the moment of the keystroke — so a recording you commit, or a CI cache somebody restores, carries the shape of your login and nothing else. The saved session lands in .smolanalytics/auth/, which is given a .gitignore of its own the first time it is written, because that file holds a live session cookie.

If a sign-in does not work, that is reported as errored — our side — and never as a failed test. A wrong password says nothing about whether your product works, and a red X on a working login is worse than no test at all.

Already generate a Playwright storage state? --auth-file path/to/state.json uses it and skips all of this.

the page that passes while looking broken

A test’s proof is text on the page. That means a build whose CSS failed to load, that rendered blank, or that is showing a crash overlay with the text still in the DOM would otherwise pass — a green tick on a page nobody could use. So a passing run also checks that the page actually rendered:

The steps all worked and the page still says what it should — but a
full-viewport error surface is covering the page: <div#o> opens with
"Unhandled Runtime Error"

It fires only on catastrophe: a blank viewport, stylesheets that failed to load, a framework error surface. A canvas game with no text, an image-only gallery, a dark theme, a cookie banner covering the whole viewport and an app that paints half a second late are all left alone — a check that cried wolf would cost more trust than the bug it caught. --no-render-check turns it off.

iframes, shadow DOM, and what a test creates

Controls inside same-origin iframes and open shadow roots are read and can be acted on, and a step says which document it happened in, so a click inside an embedded checkout is not reported as a click on your page. A closed shadow root genuinely cannot be read from outside, and the run says so rather than reporting the page as empty — the failure worth avoiding is an agent confidently calling a working embedded checkout broken.

For flows that create data, write {{email}}, {{password}} or {{runid}} in the sentence. Each run gets its own identity on an address that cannot receive mail, prefixed smoltest so one LIKE 'smoltest%' finds everything a test ever made. Add --teardown <url> and that identity is POSTed to an endpoint of yours after every run — including the runs that failed — so your app can delete what the test created.

the other half: keeping your tracking correct

The same walk through your product knows which user actions exist. That is what makes the second half possible, and it is the half nobody else has: we write and maintain your analytics tracking inside the SDK you already use — Google Analytics, PostHog, Plausible, Mixpanel, Amplitude or Segment. We do not replace your analytics. We keep its instrumentation correct.

Start with the free one. It needs no account, no key, and makes no network call, so your code never leaves your machine:

terminal
npx smolanalytics audit

It reads the repo you are standing in and names the user actions nothing is measuring, with the file and the line — payments first, then signups, logins, invites. Then your coding agent writes the calls over MCP (propose_instrumentation hands back the exact call for each site, in whichever SDK the repo already runs), and verify_instrumentation answers FIRING, WIRED or MISSING per event rather than making you wait a day to find out.

When an event in your tracking plan stops arriving but the traffic behind it doesn't, smolanalytics finds the commit that deleted the track() call and opens a pull request putting it back. The line comes back verbatim out of your own git history, in your own SDK, and it is off by default twice over: a deployment-wide switch, and a per-project setting that starts off.

The gate for CI, once a plan exists:

a CI step, with SMOLANALYTICS_KEY in the environment
npx smolanalytics plan check --project <name> --window 24

It fails the step when an event your tracking plan declares has stopped firing, so a refactor that deleted a track() call is caught on the pull request rather than a week later in a report. --window <hours> counts only recent events; --key and --url stand in for the environment variables.

Already running PostHog? Connect it from your project's setup page: paste a read-scoped key, we prove it against a real query before storing it, and the instance reads it as bounded aggregate queries on a schedule. Writing a tracking call needs only your repo, so we write for all six SDKs above; telling you a call stopped firing needs the numbers, and today we can read PostHog or our own ingest. The rest of the connections are on the integrations page, and the whole instrumentation loop has a page of its own.

the runner inside your editor

The same runner is an MCP server over stdio, so the agent you already write code with can run the suite and read every verdict without leaving the editor:

the MCP entry, in whichever editor you use
{ "command": "npx", "args": ["smolanalytics", "mcp"] }

Three tools: run_tests runs the suite and hands back every verdict, can_i_ship is the verdict plus what was not checked, and list_tests is what the folder already covers. It is local only — it never comments on a pull request, never publishes a share link, and never records a run to a project — and it needs the same ANTHROPIC_API_KEY the terminal does.

run the checkout tests against the preview and tell me what broke
can I ship this? say what you did not check
write the tracking this repo is missing, in the SDK it already uses

The third of those is the hosted half, and it is one more command. It writes the entry for your organization's MCP endpoint (smolanalytics.com/api/mcp, the API token from Settings) into every assistant it finds installed, or just one:

the hosted tools: one command, then restart your editor
npx smolanalytics connect                 # cursor, claude-code, vscode, windsurf, claude-desktop, cline
npx smolanalytics connect cursor --key <org token>

One token operates every project: the agent provisions a new one with create_project and reaches any of them by passing project="<name>", with the read key kept server-side. What it does there is the next section. Every editor's exact config is on the MCP page.

the tracking half · optional · open for the install guide

if you have no analytics at all

Everything above is the tests, and they need none of this. Behind this door is the analytics instance that comes with every project: autocapture, funnels, retention, paths and cohorts from one snippet or one endpoint, the desk that reads it, revenue webhooks and deploy markers, environments, the control-plane API and the event contract. If you already run PostHog, GA4, Mixpanel, Amplitude, Plausible or Segment, you do not need any of it — see keeping your tracking correct instead.

one command

It works out what your project is, edits the one file that needs editing, and tells you which file before it touches it.

terminal
npx smolanalytics init --host https://YOUR_HOST --key YOUR_WRITE_KEY
output
  detected  Next.js (App Router)
  file      app/layout.tsx
  edited    app/layout.tsx
  edited    .env.local

It edits for you on Next.js (both routers), SvelteKit, Vite, Create React App and plain HTML. On Nuxt and Astro it prints the install and changes nothing, because neither installs as a script tag in an HTML file and a generic snippet there gives you a page that looks instrumented and sends nothing.

Running it twice is safe: the second run sees the tracker already there and leaves everything alone. If it can't find a safe place to insert, it prints the snippet rather than guessing.

the one endpoint

Getting data in is deliberately small. In a browser, a script gives you autocapture and helpers. Everywhere else, you send events with one HTTP call, no dependency to add or keep updated:

the universal ingestion contract
POST https://YOUR_HOST/v1/events
Authorization: Bearer YOUR_WRITE_KEY
Content-Type: application/json

{ "name": "checkout", "distinct_id": "user_123", "properties": { "amount": 29, "plan": "pro" } }

That is the entire API for sending data. A single event or an array of up to 10,000. The write key is write-only, so it is safe in client code. Keep the distinct_id stable and identical across web and server, and a person's events join into one funnel on their own.

a website or web app

Paste this into your <head>. It captures pageviews and clicks immediately; add track() for the moments you want funnels on.

index.html
<script src="https://YOUR_HOST/sdk.js"></script>
<script>
  smolanalytics.init("YOUR_WRITE_KEY", { host: "https://YOUR_HOST" });

  // the moments that matter (optional, but this is what powers funnels)
  smolanalytics.identify("user_123");             // ties a person's events together
  smolanalytics.track("signup", { plan: "pro" });
  smolanalytics.track("activate");                // your core aha moment
</script>

react or vue

Same script, loaded once from your entry file so it is bundled with the app. SPA route changes (React Router, TanStack, Vue Router) are captured automatically, no per-route wiring.

src/analytics.ts (React, Vite / CRA / any SPA)
export function initAnalytics() {
  if (document.getElementById("smol")) return;
  const s = document.createElement("script");
  s.id = "smol"; s.src = "https://YOUR_HOST/sdk.js";
  s.onload = () => (window as any).smolanalytics.init("YOUR_WRITE_KEY", { host: "https://YOUR_HOST" });
  document.head.appendChild(s);
}
// call initAnalytics() once in src/main.tsx, before render
Vue: same script in index.html, or Nuxt via nuxt.config.ts
export default defineNuxtConfig({
  app: { head: { script: [
    { src: "https://YOUR_HOST/sdk.js" },
    { children: `smolanalytics.init("YOUR_WRITE_KEY", { host: "https://YOUR_HOST" });` },
  ] } },
});

next.js, app router or pages

Load it with next/script in your root layout. Pageviews and clicks are captured on every route.

app/layout.tsx
import Script from "next/script";

export default function RootLayout({ children }: { children: React.ReactNode }) {
  return (
    <html lang="en">
      <body>{children}</body>
      <Script src="https://YOUR_HOST/sdk.js" strategy="afterInteractive" />
      <Script id="smol-init" strategy="afterInteractive">
        {`smolanalytics.init("YOUR_WRITE_KEY", { host: "https://YOUR_HOST" });`}
      </Script>
    </html>
  );
}

Server-side events (a Route Handler, a Server Action, a Stripe webhook) post straight to the endpoint with the same distinct_id you pass to identify():

app/api/checkout/route.ts
await fetch(`${process.env.SMOLANALYTICS_HOST}/v1/events`, {
  method: "POST",
  headers: { "Content-Type": "application/json", Authorization: `Bearer ${process.env.SMOLANALYTICS_KEY}` },
  body: JSON.stringify({ name: "checkout", distinct_id: userId, properties: { amount: 29 } }),
});

a multi-tenant saas

A real SaaS emits events from two places, and smolanalytics is built for exactly that split. The browser sends product usage; your server sends the things the browser never sees (payments, provisioning, webhooks, cron). Three rules make it clean:

  • 1One identity everywhere. Call identify(userId) in the browser and send that same distinct_id from the server. Client and server events fuse into one funnel per user.
  • 2Sites are never the meter. Every event is stamped with its site by the SDK, so all your surfaces (marketing site, app, docs) live on one instance and one bill. You are billed on events, not on how many sites or how many users you have.
  • 3Hard isolation when you need it. Each project is its own isolated instance (its own server and data). Keep a big customer's data fully separate by giving them their own instance.
server: the revenue events the browser can't see
// Stripe webhook, billing cron, or provisioning job, any backend language.
await fetch(`${SMOL_HOST}/v1/events`, {
  method: "POST",
  headers: { "Content-Type": "application/json", Authorization: `Bearer ${SMOL_KEY}` },
  body: JSON.stringify({
    name: "subscription_started",
    distinct_id: user.id,                 // the SAME id you identify() in the app
    properties: { plan: "pro", mrr: 29, seats: 3 },
  }),
});

Now the funnel from a marketing pageview, to signup in the app, to a payment on your server is one connected path, and you ask it directly: "what's the signup to paid conversion, and how long does it take?"

mobile: ios, android, react native, flutter

Analytics only — the test agent drives web browsers, not native apps. Native SDKs, one line to initialize. Each handles an offline-safe event queue, sessions, device context, and screen() tracking (screens power funnels and paths). Call identify() on login to tie a person's events together; the write key is public (send-only), safe to ship.

iOS (Swift · Swift Package Manager)
import SmolAnalytics

// once, in your App / AppDelegate
SmolAnalytics.initialize(writeKey: "WRITE_KEY", host: "https://YOUR-INSTANCE")

SmolAnalytics.track("signup", ["plan": "pro"])
SmolAnalytics.screen("Checkout")     // screens power funnels + paths
SmolAnalytics.identify("user-123")   // on login; reset() on logout
Android (Kotlin · JitPack)
import com.smolanalytics.SmolAnalytics

// once, in Application.onCreate()
SmolAnalytics.initialize(this, "WRITE_KEY", "https://YOUR-INSTANCE")

SmolAnalytics.track("signup", mapOf("plan" to "pro"))
SmolAnalytics.screen("Checkout")
SmolAnalytics.identify("user-123")
React Native / Expo (npm)
// npm install smolanalytics-react-native
import smol from "smolanalytics-react-native";

smol.init("WRITE_KEY", { host: "https://YOUR-INSTANCE" });
smol.track("signup", { plan: "pro" });
smol.screen("Checkout");
smol.identify("user-123");
Flutter (pub.dev)
// pubspec.yaml: smolanalytics: ^0.1.0
import 'package:smolanalytics/smolanalytics.dart';

Smolanalytics.init("WRITE_KEY", host: "https://YOUR-INSTANCE");
Smolanalytics.track("signup", {"plan": "pro"});
Smolanalytics.screen("Checkout");
Smolanalytics.identify("user-123");

No dependency? Every SDK is a thin wrapper over one call, so you can also POST JSON to /v1/events with the toolkit's own HTTP client (URLSession, OkHttp, fetch, http) and batch up to 10,000 events per array to save battery.

any backend language

The events the browser never sees (payments, webhooks, cron jobs, API usage) post from any language. Same distinct_id as the client so they join up.

Node
await fetch(`${process.env.SMOL_HOST}/v1/events`, {
  method: "POST",
  headers: { "Content-Type": "application/json", Authorization: `Bearer ${process.env.SMOL_KEY}` },
  body: JSON.stringify({ name: "checkout", distinct_id: userId, properties: { amount: 29 } }),
});
Python
import requests
requests.post(f"{HOST}/v1/events",
    headers={"Authorization": f"Bearer {KEY}"},
    json={"name": "signup", "distinct_id": user_id, "properties": {"plan": "pro"}})
Go
body, _ := json.Marshal(map[string]any{"name": "checkout", "distinct_id": userID, "properties": map[string]any{"amount": 29}})
req, _ := http.NewRequest("POST", host+"/v1/events", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
http.DefaultClient.Do(req)
Ruby
require "net/http"; require "json"
uri = URI("#{HOST}/v1/events")
Net::HTTP.post(uri, { name: "signup", distinct_id: user_id }.to_json,
  "Authorization" => "Bearer #{KEY}", "Content-Type" => "application/json")
PHP
$ch = curl_init("$host/v1/events");
curl_setopt_array($ch, [
  CURLOPT_POST => true,
  CURLOPT_HTTPHEADER => ["Authorization: Bearer $key", "Content-Type: application/json"],
  CURLOPT_POSTFIELDS => json_encode(["name" => "checkout", "distinct_id" => $userId]),
]);
curl_exec($ch);

built with an ai builder, no code at all

If you built your app with Lovable, Bolt, v0, Replit, or Base44 and you do not write code, you never open an editor. Sign up, and we hand you one prompt with your write key already in it. Paste it into your builder's chat and its AI drops the snippet in the right place for that platform.

getting your data back out

Every project runs as its own isolated instance: its own server and its own data, never a shard of a shared cluster. Nothing phones home, and everything you collect exports in one file any time, so leaving is a download, not a migration project. That portability is the deal: you pay for the product, and the data is always yours to walk away with.

Prefer to kick the tyres with real data first? The live demo is a populated instance, no install, and every account starts with a 14-day trial at Pro limits, no card.

where the results show up

The project page opens on the tests: "Did anything break?", every recent run with its verdict, whether it ran the agent or a recording, and how long it took. A wall of replay rows with one agent row among them is the economics of the tool made visible.

The analytics half returns the desk, and it is not a screen: it arrives as npx smolanalytics desk in your terminal, as GET /v1/investigate, as the investigate tool in your editor, and as the weekly email. Your instance serves no dashboard to go and look at. What the desk carries: the single most expensive finding first, a queue of everything else the investigation found (each with a status: needs you, watch, fix first, recovered, verified, acted, auto-reverted), the quarter's metric movements with multiple-comparison correction so a noisy quarter can't fake a story, and a note when your product is still below the detection floor rather than a pretend finding. The classic reports (funnels, retention, paths, cohorts and the rest) are the other /v1 endpoints and the MCP tools that wrap them.

The three report tools worth knowing by name from the editor: investigate (the whole desk in one call), backtest (replay your history, each finding dated by when it would first have surfaced), and mark_finding_acted. Every answer comes from the same deterministic reports the HTTP API serves, so the model phrases the reply and never invents the figure.

The queue is not just a list, it closes. When you fix something, mark the finding acted: the mark_finding_acted tool from your editor, or one call from anywhere:

the outcome ledger
POST https://YOUR_HOST/v1/findings/acted
Authorization: Bearer YOUR_READ_KEY

{ "fingerprint": "<from the finding>", "note": "shipped fix in a1b2c3d" }

When the metric then recovers, the finding upgrades to verified: you acted on a date, and the metric recovered within N days. If you acted and it is still down, it says that instead. A regression that recovers on its own retires itself as recovered (never "fixed", no causality claimed), and a guardrailed flag that gets auto-reverted shows the receipt in the same queue. The weekly brief lands in Slack or Discord with the same tags: [verified], [acted], [recovered], [needs you].

revenue webhooks and deploy markers

Both optional, both self-serve, both change what a finding says. Point your payment provider's webhook at your instance and a finding whose metric carries an amount is priced in dollars instead of people:

revenue: paste one URL into your processor's dashboard
POST https://YOUR_HOST/v1/revenue/stripe         # or lemonsqueezy · polar · dodo

Pass your distinct_id as the checkout reference so payments join the same person as their product events. Only findings on a revenue-bearing metric get a dollar figure; nothing else is dressed up in money.

Deploy markers tie a metric change to the ship that correlates with it (correlation, not proof, and the copy says so). One line in CI, a GitHub Actions step, the GitHub App on the cloud, or nothing at all: flipping a feature flag is recorded as a ship automatically.

This is also the feed the pull-request comment about metrics runs on — a different comment from the test verdict, and a slower one. With the GitHub App installed, a cron waits 20 hours after the marker so there is a real after-window, scores every candidate metric in your own events (significance beats magnitude, so a proven 8% outranks an unproven 40%), and posts before and after per day onto the PR that shipped it. It stops considering a deploy after four days, because a comment on a stale pull request is noise, and it is claimed per commit before it posts, so a double-fired cron cannot comment twice. No computable movement, no comment.

deploy marker: one curl in your build (write key, same as events)
curl -X POST https://YOUR_HOST/v1/deploys \
  -H "Authorization: Bearer YOUR_WRITE_KEY" \
  -d '{"sha": "'$(git rev-parse HEAD)'", "message": "'"$(git log -1 --pretty=%s)"'"}'

Every integration that exists (payment providers, deploy markers, Slack and Discord delivery, importers from PostHog/Mixpanel/Amplitude/Umami, Search Console) is on the integrations page.

keeping previews out of production numbers

Every event is stamped with an env, and every report hides anything that isn't production by default. You don't configure this: localhost, private network addresses, dev tunnels (ngrok, cloudflared), Netlify deploy previews and staging./preview./qa. subdomains are detected and kept out of your real numbers. This is also what keeps a test run against your preview URL out of them.

Detection is deliberately cautious, because hiding real traffic is far worse than showing a little preview traffic. There is no blanket *.vercel.app rule, since plenty of production sites live there with no custom domain. On Vercel, pass the environment through and it wins over any guess:

app/layout.tsx
smolanalytics.init("YOUR_WRITE_KEY", {
  host: "https://YOUR_HOST",
  env: process.env.NEXT_PUBLIC_VERCEL_ENV, // "production" | "preview" | "development"
});

To look at hidden traffic, add ?env=preview (or development) to the report request, or filter on env anywhere. Nothing is discarded at ingest; it is all stored, just scoped out of the default view.

Set anything you like with env. The values hidden by default are development, preview, staging, test and ci; an unrecognised value stays visible, so a typo can never make your production numbers disappear.

the api, and embedding it in your own product

Every core operation is available over plain HTTP, authenticated with your org API token from Settings. No SDK, no meeting, no partnership form. If you ship a boilerplate, a template, or an app builder, this is everything you need to put analytics in it.

control plane
GET    /api/v1/projects        list your projects, plan and trial state
POST   /api/v1/projects        create one (instant, instances are pre-warmed)
GET    /api/v1/projects/:id
DELETE /api/v1/projects/:id    tears down the instance, then the record

Responses never include the secret read key. Only the public write key, which is ingest-only and cannot read your data.

Set someone up before they have an account

This is the one worth knowing about. You can provision a working instance for a user who has never heard of us, hand them a live write key immediately, and let them claim ownership later. Signup stops being step one of using your product.

provision on someone's behalf
curl -X POST https://smolanalytics.com/api/v1/claimable \
  -H "Authorization: Bearer YOUR_ORG_TOKEN" \
  -d '{"name": "their-app"}'

{
  "project_id":   "prj_...",
  "instance_url": "https://....fly.dev",
  "write_key":    "sa_...",          // works immediately, put it in their app
  "claim_url":    "https://smolanalytics.com/claim?p=...&t=...",
  "expires_at":   "2026-07-29T..."
}

Send them the claim_url whenever you like. Events flow from the moment the write key is wired in, and whoever opens that link takes ownership of the project and everything already collected. Links are single-use and expire (24 hours by default, up to 7 days via hours). An agent that already has a token can skip the browser entirely with POST /api/v1/claimable/accept.

The claim link is shown once. We store only a hash of it, so if you lose one, mint another rather than looking the old one up.

the event contract

fieldrequiredwhat it is
nameyesThe event, e.g. signup, checkout. A $ prefix marks internal web events (pageviews); yours have no prefix.
distinct_idnoWho did it. Use one stable value across web and server so a person's events join. Omit for anonymous counts.
propertiesnoAny JSON object: plan, amount, source. What you break funnels and cohorts down by.
  • Auth: Authorization: Bearer YOUR_WRITE_KEY. Write-only, safe in client code.
  • Batch: POST an array of up to 10,000 events in one request (max 4MB). Over the cap returns a clean 413.
  • Endpoint: POST /v1/events on your instance host. That is the entire write API.
  • Full API and every tool: see every feature and the API doc.

questions

What do I actually have to write?
One sentence per test, in a markdown file. "A returning customer can check out with a saved card." No selectors, no page objects, no fixtures, no test code at all — there is nothing in tests/ that a non-programmer could not read. Say what you expect to see rather than "checkout works": name the page, the control and the evidence, because a test that cannot fail usefully is worse than no test.
What does it need access to?
A URL and a Claude API key. The URL is one you already have (staging, a deploy preview, localhost through a tunnel) and the key is yours, in your own environment as ANTHROPIC_API_KEY — the agent runs on your machine or your CI runner, not on ours. There is no GitHub App to install, no preview environment we build for you, and nothing written to your repository. In CI the comment is posted with the GITHUB_TOKEN that Actions hands every job for free.
Does every run cost a model call?
No, and this is the part worth understanding before you price it. The first run uses the agent and takes about as long as a person would. It records what worked, and every run after that replays the recording with no model calls at all. Measured against this site: 8.0s for the first run, 1.4s for the replay. The agent comes back only when the recording stops fitting your app, which is exactly when judgement is worth paying for. In CI that means caching .smolanalytics/recordings — the workflow above does it in two steps.
What happens when I rename a button?
The replay stops fitting and the test is reported stale, not failed. A replay cannot tell a rename from a removal, so calling it a failure would page someone over a copy change. The agent re-checks that test against the sentence you wrote, decides for itself whether the app still does what you described, and rewrites the recording. There is no selector to update, because there was never a selector.
Can it write the first tests for me by reading my codebase?
No. You write the sentences. It also does not build a preview environment for you (you give it a URL), does not drive mobile apps (web browsers only), and does not seed test data (it uses whatever state the environment is already in). Those are real limits of what is shipped, not a roadmap tease.
Where does the instrumentation half come in?
The same walk through your product is what knows which user actions exist, so the same tool writes and maintains your tracking calls — inside the SDK you already run, not ours. We do not replace your analytics; we keep its instrumentation correct. Start with npx smolanalytics audit, which reads the repo you are standing in and names the user actions nothing measures, with the file and the line, with no account and no network call.
Do I need to install an SDK or a package?
Not for the tests. For the analytics half: on the web, no, it is one script tag with no npm install and no build step. On mobile there are native SDKs (Swift, Kotlin, React Native/Expo, Flutter): one line to initialize, and they handle an offline-safe event queue, sessions, and screen() tracking for you (funnels and paths need screens). For servers, webhooks, and cron there is no SDK to install, you POST JSON to /v1/events with the HTTP client your language already has. Every path is the same one ingestion endpoint underneath.
How do browser events and server events end up in the same funnel?
The same distinct_id. Call identify(userId) in the browser and send the same value as distinct_id from your backend, and a user's pageview, their signup on the client, and their payment webhook on the server all join into one person's timeline and one funnel. Pick a stable id (your user id, or a generated id you store in Keychain / SharedPreferences on mobile).
Which key do I use, and is it safe in client code?
There are two, on purpose. The WRITE key is public: it can send events and nothing else, so shipping it in a web page, a mobile binary, or a public repo is safe, it exposes no data and grants no read access. The READ key is secret: it reads your reports and connects your AI over MCP, and it appears on your project page in the cloud, under “connect your coding agent”. Never put the read key in client code. The write key cannot read your data, so a scraped write key leaks nothing.
I built my app with AI and do not write code. Can I still use this?
Yes. You never open an editor or a terminal. Sign up, copy the one prompt we hand you (your write key already in it), paste it into your builder's chat, and its AI drops the snippet in the right place. There is a page with the exact prompt for Lovable, Bolt, v0, Replit, and Base44.
Start the 14-day trial
no credit card · then from $19/mo