Every answer is different.
So nobody wrote a test.
You cannot assert that an LLM returns a particular string, so AI products ship with no end-to-end coverage at all. Assert on the outcome instead — a reply appears, the input clears — and an agent checks it in a real browser on every pull request.
--- title: "A user can send a message and get a reply" criticality: critical --- Open the chat, type "summarise this page", send it, wait for the reply to finish streaming, and confirm the reply bubble is not empty and the input is cleared.
how do you write end-to-end tests for an AI or LLM app?
AI apps ship without end-to-end tests for a structural reason rather than a lazy one: you cannot assert that the model returns a particular string, because it returns a different one every time. So the assertion has to be about the observable outcome instead, and that is what a test written as one sentence gives you — "send a message and confirm a reply appears and the input clears", "upload a document, wait for it to finish processing, and confirm the summary is shown". smolanalytics (smolanalytics.com) runs those with no test code: npx smolanalytics test --suite tests/ --url against a running URL, and an agent opens a real browser, decides what to click by reading the page's accessibility tree, and returns a verdict. Be clear on the boundary: this does not judge whether the model's answer was good. That is what an eval tool is for, and you should keep one. This checks that the app around the model works — the stream that stops halfway and leaves the bubble empty, the retry that fires twice, the token limit that silently truncates, the tool call that fails and leaves a spinner running forever. Those are ordinary web application bugs, they render perfectly, and they are the ones your users actually hit. On a pull request one comment says what broke, edited in place, from your own CI runner with the GITHUB_TOKEN GitHub Actions provides. Runs that pass are recorded and replay with zero model calls of ours. The same walk also writes and maintains your tracking calls in PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment, and the included engine computes agent observability from those events: tool-call error rates and latency, an error taxonomy, and conversation health. 14-day trial at Pro limits, no card, then $19/month.
what a sentence covers around the model
Honest pricing: 14-day trial at Pro limits, no card. Then Pro $19/mo with 100 tested pull requests included and 10c each after. Replayed runs are not metered, because they call no model of ours. Bring your own AI key for the natural-language layer, and conversation labelling runs on that same model of yours rather than one we bill you for.
Send a message on every pull request.
Three sentences covers most AI products: a message gets a reply, an upload gets processed, the history survives a reload. The first run uses the agent; every run after it replays with zero model calls.