for product managers

You cannot read a test suite. You can write a sentence.

"A returning customer can check out with a saved card." That line is a complete end-to-end test. An agent uses the app in a real browser on every pull request and one comment says what broke.

It is also, word for word, the acceptance criterion you already wrote in the ticket. There is no translation step into test code, which is the step where the intent goes missing, and nothing for anyone to maintain when the button moves.

the arithmetic, before the pitch

This does not give you your week back, and the site should keep saying so. Against how PMs actually spend a 48-hour week, the work it removes — noticing something looks wrong, asking an engineer to confirm it, waiting, plus the amortised chasing of instrumentation — is somewhere between 1.8 and 4.9 hours. That is about 1.1x. Meetings, customers, specs, launches and saying no are the other thirty-six hours, and none of them are touched here.

The honest claim is zero to one on a different question. "Is the flow I specced in March still working?" has never had a cheap answer: you ask an engineer, or you open the app and click it yourself, and either way you get the answer once, on the day you remembered to ask. Pendo measured what that costs at scale — 80% of features rarely or never used, which is partly a discovery problem and partly nobody ever going back. This moves the checking from rarely to every pull request, in a sentence you wrote, with nobody assigned to it.

And it does not replace QA or an engineer, which is worth stating plainly because the opposite claim is the one that would cost you the room. It does not explore. It does not notice that something feels wrong. It checks the things somebody wrote down, forever, which is precisely the work that people are worst at and least motivated to repeat.

what you actually get

Your acceptance criteria, as the test

the sentence you already wrote in the ticket

"A returning customer can check out with a saved card." That line, in a markdown file, is a complete end-to-end test. There is no translation step into code, which is where intent usually gets lost, and no page object for anyone to maintain. If you can specify it, it can be checked — and it is checked on every pull request rather than on the day somebody remembers.

One comment on the pull request

the regression, named, in the review

When something you specced stops working, one comment appears on the pull request that caused it, edited in place rather than stacked, saying which sentence failed and what the agent saw instead of what you described. That is early enough to be a review note. The alternative timing is a customer email, and everything between those two points is the value.

Five words, kept strictly apart

and only one of them is a bug report

Passed. Failed, meaning the app did not do what the test describes — a bug report. Flaky, meaning it failed once and passed on retry: unreliable, never counted as a pass, and it does not fail the build. Stale, meaning a recording no longer fits the page, which is what a rename produces, and it is never worded or coloured as a failure because a renamed button and a deleted one look identical to a replay. Errored, meaning our runner could not run at all, which is ours and says so. A tool that reports a copy change as a bug is a tool your engineers will teach you to ignore.

The flows nobody walks

signup, billing, invites, cancellation

The paths your team exercises daily are the ones least likely to be broken. The dangerous set is the ones only strangers walk: registration, the upgrade screen, the second seat, the password reset, cancellation. Nobody internally has done any of them in months. One sentence each is the cheapest coverage in the product and it is coverage of exactly the flows that convert or churn.

The measurement half, once the tests are running

and it is genuinely second

The same walk through the product knows which user actions exist, so it also keeps the tracking correct in whatever analytics you already run, which is the reason a funnel quietly loses a step after a refactor. From there the analytics side still does what it did: every product event tested for step changes, findings ranked by what they cost, the deploy that correlates named, and a finding you mark acted on upgraded to verified only if the metric recovered after that date.

questions

Does this replace QA, or an engineer?

No, and a page that implied it would be lying to the person most likely to champion it. It does not explore, it does not use judgement about whether something feels wrong, and it only ever checks the things somebody wrote down. What it changes is coverage of the written-down things: a flow you specced in March is checked on every pull request forever, rather than the day somebody remembers to click it. That is a job that currently happens rarely, not a job that currently takes a person hours.

So what does it actually save me?

Less time than a vendor would tell you, and the arithmetic is on this page. It does not touch the thirty-six hours of meetings, customers, specs, launches and saying no. What it removes is the loop where you notice something looks wrong, ask an engineer to confirm it, and wait — plus the meeting where nobody can say whether the thing shipped last month still works. What it adds is an answer you did not previously have at any price: a verdict, on every change, in words you wrote.

Do I write these, or does engineering?

Either, and in practice both. The sentence is the acceptance criterion, so if you already write "a returning customer can check out with a saved card" in the ticket, you have written the test. Engineers tend to add the ones you would not think of, and nobody has to translate anything into code, which is the step where the intent usually gets lost.

What do I do when something fails?

Read it. The failure names the sentence and what the agent saw instead: "expected the order number, the page still showed the payment form with an error under the card field." That is a bug report a person can act on and a stakeholder can understand, on the pull request that caused it, which is the difference between a regression found in the review and a regression found by a customer.

We already have PostHog. Why add this?

Keep it. This runs beside whatever you have and answers a different question. Analytics tells you a number moved; it cannot tell you the flow is broken, because a flow that is broken for a subset of people looks like a slightly worse conversion rate for a week. A test that walks the flow says which step failed, on the day it started failing, with the change that caused it in the same view.

keep reading