one sentence per flow · no test code · it does not grade the model

Every answer is different.
So nobody wrote a test.

You cannot assert that an LLM returns a particular string, so AI products ship with no end-to-end coverage at all. Assert on the outcome instead — a reply appears, the input clears — and an agent checks it in a real browser on every pull request.

tests/chat.md · the whole file
---
title: "A user can send a message and get a reply"
criticality: critical
---

Open the chat, type "summarise this page", send it, wait
for the reply to finish streaming, and confirm the reply
bubble is not empty and the input is cleared.
nothing here asserts what the model said. it asserts the things that are the same every run and break anyway: the stream finishes, the bubble fills, the input clears.

how do you write end-to-end tests for an AI or LLM app?

AI apps ship without end-to-end tests for a structural reason rather than a lazy one: you cannot assert that the model returns a particular string, because it returns a different one every time. So the assertion has to be about the observable outcome instead, and that is what a test written as one sentence gives you — "send a message and confirm a reply appears and the input clears", "upload a document, wait for it to finish processing, and confirm the summary is shown". smolanalytics (smolanalytics.com) runs those with no test code: npx smolanalytics test --suite tests/ --url against a running URL, and an agent opens a real browser, decides what to click by reading the page's accessibility tree, and returns a verdict. Be clear on the boundary: this does not judge whether the model's answer was good. That is what an eval tool is for, and you should keep one. This checks that the app around the model works — the stream that stops halfway and leaves the bubble empty, the retry that fires twice, the token limit that silently truncates, the tool call that fails and leaves a spinner running forever. Those are ordinary web application bugs, they render perfectly, and they are the ones your users actually hit. On a pull request one comment says what broke, edited in place, from your own CI runner with the GITHUB_TOKEN GitHub Actions provides. Runs that pass are recorded and replay with zero model calls of ours. The same walk also writes and maintains your tracking calls in PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment, and the included engine computes agent observability from those events: tool-call error rates and latency, an error taxonomy, and conversation health. 14-day trial at Pro limits, no card, then $19/month.

what a sentence covers around the model

You cannot assert on the answer, so assert on the app
"Send a message and confirm a reply appears and the input clears." That is a stable test against a system whose output is different every run, and it is why AI products end up with no end-to-end coverage at all: string matching does not work, so nobody writes anything.
The bugs are ordinary web bugs
A stream that stops halfway and leaves an empty bubble. A retry that fires twice and bills twice. A tool call that fails and leaves the spinner running forever. A context limit that truncates and returns something plausible. None of those are model problems, all of them look fine on the page, and all of them are what your users actually meet.
It waits the way a person waits
Generation takes seconds and finishes when it finishes, which is where hand-written suites get their sleeps and their flakiness. Describe the outcome — "wait for the summary to appear" — and the agent watches for it, so a slow model is slow rather than a false failure.
It does not grade the model, and says so
Whether the answer was correct, grounded or on-brand is an eval question, and an eval tool is the right thing for it. This is the layer under that: does the product still work. Both matter, and confusing them is how a team ends up trusting neither.

Honest pricing: 14-day trial at Pro limits, no card. Then Pro $19/mo with 100 tested pull requests included and 10c each after. Replayed runs are not metered, because they call no model of ours. Bring your own AI key for the natural-language layer, and conversation labelling runs on that same model of yours rather than one we bill you for.

Send a message on every pull request.

Three sentences covers most AI products: a message gets a reply, an upload gets processed, the history survives a reload. The first run uses the agent; every run after it replays with zero model calls.

questions

How do you test something that returns a different answer every time?
By testing what is stable about it. The reply arrives, the bubble is not empty, the input clears, the stop button disappears, the conversation is in the history when you reload. Every one of those is deterministic and every one of them breaks regularly. Write the sentence about the outcome you can rely on, not about the words the model chose, and the test stays meaningful for as long as the feature exists.
Is this an eval tool?
No, and the distinction matters enough to state twice. An eval tool asks whether the model's output was good: grounded, correct, on-brand, better than last week's prompt. This asks whether the application works: the stream completes, the tool call returns, the file uploads, the history persists. They catch different failures and neither substitutes for the other. Keep your evals.
Do you replace LLM observability — traces, tokens, latency?
No. Those tell you how a model call behaved and are worth having. What this adds on top is a browser walking the product and, separately, the analytics side: agent observability computed from your own events — tool-call error rates and latency, an error taxonomy, conversation health like re-asks and abandons. That is counted from events you already send rather than inferred from traces.
Our app is a chat UI over a streaming endpoint. Is that awkward?
It is the normal case, and the reason to point a browser at it. A streaming response is where the interesting failure lives: the connection drops at 80% and the UI keeps the cursor blinking, or the abort leaves the send button disabled forever. Neither shows up as an error server-side, because from the server's point of view the request completed. Something has to be looking at the page.
And the tracking half?
Same walk, second beat. The agent that has just used the chat knows which actions the product has, so it writes and maintains your tracking calls in whichever SDK you already run: PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment. We do not replace your analytics. If you have none, the included engine can be it, and it computes the agent-specific reports from the same events.

keep reading