Experiment

What we paid for AI in 30 days, and what we'd pay for again

A pre-registered audit of Ovriko's 30-day AI spend, abandoned on 2026-08-27 under its own published abandon condition after logging lapsed for more than three consecutive days.

Published

The record

Abandoned
Hypothesis
At least 20% of Ovriko's total in-scope AI spend for the window can be identified as removable without losing any output classified as used under the definition registered below. This single threshold decides PASS or FAIL on its own.
Period
2026-08-22 — 2026-08-27
Headline result
No result. Abandoned on 2026-08-27 under the pre-registered abandon condition for a logging lapse of more than three consecutive days. No PASS, FAIL, KEEP or FIRE is claimed.

The question

Of the AI tools and usage paid for over one month to build and run Ovriko, which produce output that is actually used - and what would survive if every AI bill had to be justified line by line?

The honest version of this question is uncomfortable, which is why it is worth asking first. AI spend accumulates quietly: a subscription here, a credit top-up there, three tools doing approximately the same job because each was convenient on the day it was added.

The hypothesis

At least 20% of total in-scope AI spend for the window can be identified as removable without losing any output classified as used.

That sentence, and nothing else on this page, decides PASS or FAIL.

A separate prediction that does not decide anything

There is also a prediction worth recording: that a minority of line items will account for the majority of used output. It is reported descriptively when the window closes, whether it holds or not.

It is not threshold-bearing. It cannot cause a PASS, it cannot cause a FAIL, and no combination of it with the spend threshold changes the outcome. Recording it here, separated from the hypothesis, is what stops it being quietly promoted into evidence later.

How the removable-spend percentage is calculated

Fixed now, so it cannot be reinterpreted once the numbers are visible:

removable-spend % =
    ( sum of window cost of line items qualifying under FIRE rule 1 or FIRE rule 2 )
  / ( total in-scope AI spend for the window )
  x 100

Numerator. The window cost of every line item that qualifies under FIRE rule 1 (no use) or FIRE rule 2 (substitutable), and nothing else. A line item qualifying under rule 1 or rule 2 still counts even if rule 3 also happens to apply to it.

Denominator. Total in-scope AI spend for the window. In-scope spend that could not be allocated stays in the denominator — it is never removed to make the percentage easier. It simply cannot appear in the numerator, because spend that could not be attributed cannot be shown to be removable.

Window cost means the cash cost recognised for the window under the cost rules below: usage as incurred, annual plans at one twelfth, free credits at zero, ex-tax, net of refunds received in the window.

Spend fired under rule 3 alone never enters the numerator. Nothing enters the numerator if removing it would lose a used output.

PASS if removable-spend % >= 20. FAIL if it is below 20. If the insufficient-evidence rule fires, there is no PASS or FAIL at all.

Scope: what belongs to Ovriko

The venue is Ovriko's own online operating activity: building, writing, designing, publishing and running this site and its media workflow.

Included: AI charges incurred for Ovriko work during the window.

Excluded, entirely and without exception:

  • unrelated personal AI use
  • AI use belonging to any other business or venture
  • anything that cannot be attributed to Ovriko work under the allocation rule below

Where a single subscription covers both Ovriko and non-Ovriko work, only the Ovriko share is counted, using the allocation rule below. Where a charge cannot be attributed at all, it is recorded as unallocable rather than guessed at, and unallocable spend has its own consequence - see the insufficient-evidence rule.

What counts as AI spend

Subscriptions, per-token and usage charges, AI add-ons to existing tools, and paid credits consumed during the window, where the primary function is AI inference or AI-assisted work.

Not counted: general infrastructure without an AI function, hardware, and non-AI software. Borderline calls are logged with a reason at the time the call is made, never reconstructed afterwards.

What counts as used output

Deliberately not counted:

  • output that was generated
  • output that was reviewed
  • output saved for possible later use but not used during the window
  • output described as impressive, promising or useful, but never actually used

Rewritten drafts. An AI draft that was substantially rewritten before it shipped can still count, as exactly one used output, but only where the log records contemporaneously that the AI output materially formed the basis of the final artifact. It does not count automatically because an AI was involved at some point. There is no percentage-of-text test - inventing one would imply a precision that is not being measured.

Multiple candidates. If ten alternatives are generated and one of them materially forms the basis of what is used, that is one used output, not ten. An artifact is never counted twice because several revisions were generated on the way to it.

The gap between "generated a lot" and "used any of it" is the entire subject of this experiment, so the definition is drawn tightly on purpose.

Cost and time accounting

Annual subscriptions are amortised at one twelfth of the annual price for the window. No annual plan is counted at its full price.

Free trials and credits are recorded at zero cash cost, but the usage and review time they consume is counted in full, and they are listed separately so they cannot flatter the totals. Any trial that would auto-convert to a paid plan is noted.

Tax is recorded separately; line items are compared ex-tax for comparability. Refunds and credits received within the window are netted off and disclosed.

Shared or multi-purpose tools are allocated by logged AI-session minutes as a share of total logged minutes for that tool. The allocation is published even where it is unflattering.

Time is logged per session in minutes - prompting, waiting, reviewing, correcting, re-running - recorded contemporaneously, never reconstructed. Operator time is valued at USD $100/hour, a declared internal decision rule for this experiment only. It is not a claim about market compensation or opportunity cost.

How the logs are kept

The used-output log and the time log are the evidence this experiment rests on, so the rule for keeping them is stated before it starts. This is a procedural commitment, not a technical guarantee: nothing here is enforced by software, and it is not claimed to be.

  • Entries are recorded the same day, before the result is known.
  • Existing entries are not silently rewritten.
  • Where an entry needs correcting, the correction is added as a dated note alongside the original, so both the original entry and the change remain visible.

The KEEP / FIRE rule

A line item is FIRE if any of the following hold, judged only from contemporaneous logs after the window closes:

  1. No use. It produced zero used outputs during the window.
  2. Substitutable. Every used output it produced could have been produced by another retained line item, where the specific substitute was named in the log at the time the output was created.
  3. Negative time economics. Its allocated cost for the window exceeded the value of the net time it saved, where net time saved is the sum, across its used outputs, of the pre-logged unaided estimate minus the logged actual time, valued at $100/hour.

Otherwise it is KEEP.

Rule 3 depends on a pre-logged estimate of how long a task would have taken unaided. That estimate is an operator judgement recorded before the task is performed, not a measurement, and it is reported as such. It deliberately has no influence on whether the hypothesis passes.

Evidence retained

Invoices and receipts, usage dashboard exports, the daily used-output log, the session time log, and the dated pre-registration itself.

Personal data, account identifiers, private correspondence and commercially sensitive material are redacted, and the fact of redaction is disclosed. Credentials, API keys and secrets are never published as evidence in any form. Vendor names and prices are published - they are not confidential.

Known confounders

Published here rather than discovered later:

  • A single operator, a single business, a single month. This is a case, not a general law.
  • Workload is not constant. A month of heavy building and a month of heavy writing consume AI differently.
  • Novelty effect may inflate early usage of anything recently added.
  • No blinding is possible. The operator knows the experiment is running, which can itself change purchasing and usage behaviour.
  • Tool prices and plans may change mid-window.
  • "Used" is a self-assessed judgement. The only mitigation is same-day logging, before the outcome is known.

What the used-output metric cannot see. The definition is narrow on purpose, and narrowness has a price. It will undercount at least two kinds of real value:

  • Preventive value - catching an error before a bad change ships. It produces no artifact, so it does not count as a used output unless it directly produced a qualifying logged decision under the definition above.
  • Learning and explanation - understanding something better with an AI's help, where no qualifying artifact or directly logged decision results.

Neither is quietly promoted into a used output to fix the shortfall. Both are disclosed as a limitation in the published result, so a line item that scored badly is not read as having produced nothing.

Abandon conditions

The experiment is abandoned, and published as abandoned with its reason, if:

  • logging lapses for more than three consecutive days
  • a major workload disruption makes the month unrepresentative
  • an accounting change makes spend unattributable

Insufficient evidence

The experiment completes but declines both an outcome and a decision if either:

  • fewer than 10 used outputs are recorded across the window, or
  • more than 20% of total in-scope AI spend cannot be allocated under the rules above

In that case the result is published as complete with insufficient evidence and the reason attached. It is not rounded to FAIL, and it is not rounded to FIRE.

Status

Abandoned on 2026-08-27. The measurement window opened on 2026-08-22 and closed early, on day 6 of 30, when the abandon condition published above fired: logging lapsed for more than three consecutive days.

The four ledgers this experiment rests on - spend, outputs, time and corrections - hold no entries for any day of the window. Under the rules published above an unlogged day cannot be reconstructed afterwards, and nothing has been backfilled, estimated or inferred to close the gap. There is therefore no data, no removable-spend percentage and no verdict: no PASS, no FAIL, no KEEP, no FIRE, and no insufficient-evidence ruling either, which is reserved for an experiment that completes.

This is the outcome the abandon condition was written to produce. It was published before the window opened so that a lapse would end the experiment in the open rather than be quietly repaired later, and applying it costs less than the credibility of a result assembled from memory would.

The method above is unchanged and stands as published on 2026-08-21. A future audit of Ovriko's AI spend would be a new experiment with its own pre-registration, not a continuation of this one.

Related

Research

How Ovriko tests, and what it will not publish

The operating rules behind every Ovriko experiment - how a hypothesis is fixed before any money is spent, how cost and time are counted, what PASS, FAIL, KEEP and FIRE mean, and the things Ovriko will not publish at any price.