backtest

What would it have told you, and on which day?

Run the same investigator over your own history with the clock moved back, one sweep at a time. Every finding comes back stamped with the day it would first have reached you, and the detection lag in days.

This is the optional tracking half of smolanalytics, and it needs a quarter of your own events to say anything. The product itself starts somewhere else: one sentence describing what should work, an agent that uses your app in a real browser on every pull request, and a comment saying what broke.

An analytics backtest replays your own event history through the same detection logic that runs today, with the clock moved back one step at a time, and reports each finding dated by when it would first have been surfaced along with the detection lag in days. smolanalytics does this with the backtest MCP tool in your editor or GET /v1/backtest?days=90&step=1 over HTTP, both the same code path as the live investigation rather than a reporting mode built for demos. It is possible because every report is a pure function of the raw event log (no pre-aggregated rollups, no sampling), so any question can be recomputed as of any past date. Each sweep sees only events that had already happened, and deploy markers are filtered to the sweep's own date, so a finding can never cite a release that had not shipped yet.

running on the demo right now

live replay · 31 sweeps · 2026-07-03 → 2026-10-01
90 days of a real log. Six things in it nobody read at the time.
computed just now off the demo's real history by /v1/backtest, not a screenshot
  1. 14 augsignup fell 26% on 2026-08-12
    ~240 people/mo2d to become conclusive
    100% of the loss is Safari.
  2. 17 augactivate fell 25% on 2026-08-14
    ~105 people/mo3d to become conclusive
    100% of the loss is Firefox.
  3. 17 augcheckout fell 33% on 2026-08-13
    ~60 people/mo4d to become conclusive
    Ship f6a7b8c "new pricing page" landed the same day (correlation, not proof). 67% of the loss is macOS.
  4. 20 augopen rose 47% on 2026-08-10
    ~660 people/mo10d to become conclusivecause not narrowed
  5. 20 augsignup fell 28% on 2026-08-10
    ~300 people/mo10d to become conclusive
    92% of the loss is US.
  6. 23 augcheckout rose 67% on 2026-08-19
    ~120 people/mo4d to become conclusivecause not narrowed
2 of these could not be narrowed to a ship or a segment, because the demo records no deploys. wire yours up and that line becomes which ship did it.
15 more in the same replay, in the order they'd have arrived. nothing here is picked: these are the first 6, in arrival order.

your last quarter already contains this. bring your PostHog or Mixpanel history across with migrate_from, then ask your agent for backtest. both are MCP tools in your editor, both live during the trial · sweeps every 3 days here, daily on your instance

why a replay over a real quarter finds something

Not because the tool is clever. Because of a measured prior about how software gets built.

01

Most shipped work does not move the metric it was built to move.

Microsoft's published measurement across thousands of controlled experiments found roughly a third of tested ideas improve their target metric; Google and Netflix have put it closer to one in ten. Pendo's telemetry across 615 subscriptions found 80% of features are rarely or never used. A replay over any real product's last quarter will find something, because the base rate guarantees it.
02

The date is the part no other tool can produce.

A dashboard says “here is your product”. A replay says “on 14 June this would have told you checkout was broken”, and you know exactly when you actually noticed. That second date is the one nobody else can produce, and it is the honest measure of what the tool is worth to you specifically.
03

The detection lag is printed.

Findings need enough data to become conclusive, so most arrive a few days after the change itself. We print that number rather than hide it, because “we'd have told you the same day” is usually false and a sceptical reader checks the least believable claim first. Four days late still beats the six weeks it usually takes someone to notice on their own.
04

Nothing is cherry-picked, and the same finding is reported once.

Findings render in the order they would have arrived, and the total is stated. A regression visible on forty consecutive daily sweeps is one thing you would have been told about once, not forty: deduplicated on the metric and the change day, keeping the earliest sweep that saw it.
05

It cannot read forward, and that is tested in both directions.

Each sweep receives only events strictly before its own date; deploys and experiments are filtered the same way. Testing only that a future ship is never named would pass just as well if attribution were broken outright, so the control case is tested too: a ship that landed on the change day must be named. Writing that control is how we found that ship attribution had never fired for anyone, because a threshold measured in fractions had been passed a whole number.

running it on your own quarter

your editor
> using smolanalytics, backtest my last 90 days
> ...and again at 180 days, stepping every 2

The backtest MCP tool, running on the model you already pay for, so it works from the first hour of the trial. Ends with the only question that matters: for each line, did you know, did you not know, or is it wrong?

http
GET /v1/backtest?days=90&step=1

Same code path, same defaults. Memoized for an hour, because it is an expensive exact answer rather than a cheap approximate one.

14 days, no card. Then $19/mo with 100 analysed deploys, 10c each after.

questions

What is an analytics backtest?

Running your analytics against history you have already lived through, to see what it would have surfaced and on which day. smolanalytics re-runs the same investigator that runs today, once per step across a window, with the clock moved back each time, so each sweep sees only the events that had happened by that date. Every finding comes back stamped with the day it would first have been reported and the detection lag: how long the data took to become conclusive after the change itself.

How is this different from just looking at a chart of the last 90 days?

A chart shows you the past with today's knowledge. A backtest shows you what you would have KNOWN at the time, which is a claim about a specific day you remember. "On 14 June this would have told you checkout was broken, and you found out on 2 July" is checkable against your own memory. "Checkout dipped in June" is not a claim about anything.

Can it see the future while replaying?

No, and this is the part that decides whether the whole artefact is worth anything. Each sweep is handed only the events that occurred strictly before its own date, deploy markers are filtered the same way, and experiments that had not started yet are dropped. A finding dated before the change it describes, or a June regression attributed to a July release, would make every lag number a fabrication, so both are tested, in both directions: the future ship must not be named, and the ship that landed on the drop must be.

Why can smolanalytics replay history when other analytics tools cannot?

Because every report here is a pure function of the raw event log. There are no pre-aggregated rollup tables and no sampling, so any question can be recomputed as of any past date. Tools built on rollups cannot recompute a past day's answer independently; they can only read the rollups that already exist, which were written with whatever definitions were in force at the time.

Will the replay find something on my product?

Almost certainly, and that is not a boast about the tool. Microsoft's published measurement across thousands of controlled experiments found roughly a third of tested ideas improve the metric they were built to improve; Google and Netflix have reported closer to one in ten. Pendo's telemetry across 615 subscriptions found 80% of features are rarely or never used. Any real product's last quarter contains waste. The replay's job is to date it.

What if I have no history here yet?

Then the replay has nothing to replay, and it will say so rather than manufacture a quarter. This is the one honest weakness of an event-based product on day one: a step change needs a baseline. Two routes out. Bring your existing log across with migrate_from, a keyed pull of your PostHog or Mixpanel events over a date range you name, dry run first, and the replay reads that. Or start with the checks that read code instead of events, which answer immediately and need no account at all: npx smolanalytics audit names the user actions your repo does not measure, and npx smolanalytics plan check fails CI when a planned event stops firing.

Does it report nothing if nothing happened?

Yes, and it says what it looked at. An empty replay reads "across the whole window there was no step change big enough to be worth telling you about", followed by the list of metrics it swept, so an honest empty is distinguishable from a broken run. A backtest that manufactures findings to look impressive is worth less than one that comes back empty, because the first thing a sceptic does is check a line.

How do I run it on my own data?

Two ways, over whatever history you already have. From your editor it is the backtest MCP tool, which runs on the model you already pay for and therefore works throughout the trial. Over HTTP it is GET /v1/backtest?days=90&step=1, where days defaults to 90 and step defaults to 1, meaning it re-checks daily. Both are the same code path as the live investigation, not a separate reporting mode built for demos. If your log here is new, bring history across first: migrate_from is a keyed pull of your PostHog or Mixpanel events over a date range you name.

keep reading