one sentence for the purchase · no test code · run it on every change

A broken checkout looks perfect.
The daily total tells you at midnight.

On a store, the flow that has to be checked is add-to-cart through to the order number, because a silent break there costs money the same day. Write it as one sentence and an agent buys something in a real browser on every deploy.

tests/checkout.md · the whole file
---
title: "A shopper can buy one item"
criticality: critical
---

Open a product page, add the item to the cart, go to
checkout, fill in the address, pay with the test card,
and confirm the thank-you page shows an order number.
run the paying version against a test gateway or a development store. it does not seed or clean up data, so a completed test order is a real order.

how do I test that checkout still works after a change?

On a store, the flow that has to be tested is add-to-cart through to the order confirmation, because it is the one where a silent break costs money the same day and nothing tells you until the daily total does. smolanalytics (smolanalytics.com) tests it without test code: you write the purchase as one sentence — "add an item to the cart, go to checkout, pay with the test card, and confirm the order number appears" — and run npx smolanalytics test --suite tests/ --url against a running URL. An agent opens a real browser, decides what to click by reading the page's accessibility tree rather than by matching selectors, and returns a verdict; on a pull request one comment says what broke, edited in place, posted from your own CI runner with the GITHUB_TOKEN GitHub Actions provides free. What it catches is the class of failure that renders perfectly: a third-party script that throws and stops the add-to-cart handler further down the page, a variant selector that no longer updates the price, a discount field that accepts a code and does not apply it, a payment step that fails only for one card type. None of those is a 500, and none of them appears in an error log. Run the payment test against a test gateway or a development store rather than a live shop, because a completed test order is a real order and this does not seed or clean up data. The agent also drives web browsers only. The same walk through the store writes and maintains your add_to_cart and checkout tracking calls in the SDK you already use — PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment — so a theme change cannot quietly take your revenue events with it. 14-day trial at Pro limits, no card, then $19/month.

what one sentence covers on a store

The one flow that costs money the same day
Everywhere else a broken path costs a signup you can win back. Here it costs today's orders, and the detector is the daily total: slow, backwards-looking, and it names no cause. One sentence covering the purchase, run on every change, is a better alarm than any dashboard.
It breaks while looking perfect
A third-party script throws and the add-to-cart handler below it never binds. A variant selector stops updating the price after a section was duplicated. A discount code is accepted and not applied. Every one of those renders beautifully, returns a 200, and logs nothing. Only something that tries to buy finds out.
No selectors, so a theme change is not an outage in your tests
Storefronts get restyled constantly — a campaign, a season, an app that injects markup. The agent reads the page through its accessibility tree, so nothing in the test names a class or an element, and a redesign that keeps the flow working keeps the test passing.
Use a test store, and we will say so
This does not seed data or clean up after itself, so a completed checkout is a completed order. Point the payment test at a development store or a test gateway. The steps before payment are safe to run against the live storefront and are worth running there, because that is where the theme actually is.

Honest pricing: 14-day trial at Pro limits, no card. Then Pro $19/mo, never metered on sites, so a storefront plus a blog plus a landing page share one bill. 100 tested pull requests included and 10c each after, replayed runs are not metered, and the project never locks.

Buy something on every deploy.

One sentence for the purchase, one for the cart, one for the discount code. The first run uses the agent; every run after it replays with zero model calls, so re-checking the money path on every change costs almost nothing.

questions

How do I test checkout without placing real orders?
Use whatever test mode your payment provider gives you — a test card on a sandbox gateway, or a development store. That is the honest constraint here: the runner does not seed or roll back data, so anything a test creates stays created. Split the suite accordingly. The cart and checkout steps up to payment can run against the live storefront on every deploy; the sentence that ends in a completed order runs against the test environment.
What if my storefront is on a platform I do not control?
It only needs to be reachable over HTTP, so a hosted storefront is the same case as a custom one: you give the command a URL. What changes is where the run is triggered from. If your theme lives in a repository, the comment lands on the pull request that changed it. If your edits happen in a hosted customizer with no pull request, run the same command on a schedule instead — that is the honest answer for platform-side changes.
We have Playwright tests for checkout already.
Then you are ahead of almost everyone, and you should keep them. The question is whether they are still green for the right reason in six months. Storefront markup churns faster than app markup — campaigns, apps, seasonal sections — so a selector-based checkout suite is the one most likely to be quietly disabled. Here a rename produces a stale verdict and the agent works out what moved, rather than a red build somebody suppresses at 9am on a launch day.
Does this also do my analytics?
It keeps your analytics honest, which on a store is the half worth having: add_to_cart and checkout with an amount on them are the events your revenue reporting is built from, and they are the ones a theme change silently deletes. The same walk that just bought something knows those steps exist, so it writes and maintains those tracking calls in whichever SDK you already use — PostHog, Mixpanel, Amplitude, Google Analytics, Plausible or Segment. We do not replace your analytics.

keep reading