What we paid for AI in 30 days, and what we'd pay for again
A pre-registered audit of Ovriko's 30-day AI spend, abandoned on 2026-08-27 under its own published abandon condition after logging lapsed for more than three consecutive days.
Published
The record
Abandoned
Hypothesis
At least 20% of Ovriko's total in-scope AI spend for the window can be identified as removable without losing any output classified as used under the definition registered below. This single threshold decides PASS or FAIL on its own.
Period
2026-08-22 — 2026-08-27
Headline result
No result. Abandoned on 2026-08-27 under the pre-registered abandon condition for a logging lapse of more than three consecutive days. No PASS, FAIL, KEEP or FIRE is claimed.
The question
Of the AI tools and usage paid for over one month to build and run Ovriko,
which produce output that is actually used - and what would survive if every
AI bill had to be justified line by line?
The honest version of this question is uncomfortable, which is why it is worth
asking first. AI spend accumulates quietly: a subscription here, a credit
top-up there, three tools doing approximately the same job because each was
convenient on the day it was added.
The hypothesis
At least 20% of total in-scope AI spend for the window can be identified as
removable without losing any output classified as used.
That sentence, and nothing else on this page, decides PASS or FAIL.
A separate prediction that does not decide anything
There is also a prediction worth recording: that a minority of line items will
account for the majority of used output. It is reported descriptively when the
window closes, whether it holds or not.
It is not threshold-bearing. It cannot cause a PASS, it cannot cause a FAIL,
and no combination of it with the spend threshold changes the outcome. Recording
it here, separated from the hypothesis, is what stops it being quietly promoted
into evidence later.
How the removable-spend percentage is calculated
Fixed now, so it cannot be reinterpreted once the numbers are visible:
removable-spend % =
( sum of window cost of line items qualifying under FIRE rule 1 or FIRE rule 2 )
/ ( total in-scope AI spend for the window )
x 100
Numerator. The window cost of every line item that qualifies under FIRE
rule 1 (no use) or FIRE rule 2 (substitutable), and nothing else. A line item
qualifying under rule 1 or rule 2 still counts even if rule 3 also happens to
apply to it.
Denominator. Total in-scope AI spend for the window. In-scope spend that
could not be allocated stays in the denominator — it is never removed to make
the percentage easier. It simply cannot appear in the numerator, because spend
that could not be attributed cannot be shown to be removable.
Window cost means the cash cost recognised for the window under the cost
rules below: usage as incurred, annual plans at one twelfth, free credits at
zero, ex-tax, net of refunds received in the window.
Spend fired under rule 3 alone never enters the numerator. Nothing enters the
numerator if removing it would lose a used output.
PASS if removable-spend % >= 20. FAIL if it is below 20. If the
insufficient-evidence rule fires, there is no PASS or FAIL at all.
Scope: what belongs to Ovriko
The venue is Ovriko's own online operating activity: building, writing,
designing, publishing and running this site and its media workflow.
Included: AI charges incurred for Ovriko work during the window.
Excluded, entirely and without exception:
unrelated personal AI use
AI use belonging to any other business or venture
anything that cannot be attributed to Ovriko work under the allocation rule
below
Where a single subscription covers both Ovriko and non-Ovriko work, only the
Ovriko share is counted, using the allocation rule below. Where a charge cannot
be attributed at all, it is recorded as unallocable rather than guessed at, and
unallocable spend has its own consequence - see the insufficient-evidence rule.
What counts as AI spend
Subscriptions, per-token and usage charges, AI add-ons to existing tools, and
paid credits consumed during the window, where the primary function is AI
inference or AI-assisted work.
Not counted: general infrastructure without an AI function, hardware, and
non-AI software. Borderline calls are logged with a reason at the time the
call is made, never reconstructed afterwards.
What counts as used output
Deliberately not counted:
output that was generated
output that was reviewed
output saved for possible later use but not used during the window
output described as impressive, promising or useful, but never actually used
Rewritten drafts. An AI draft that was substantially rewritten before it
shipped can still count, as exactly one used output, but only where the log
records contemporaneously that the AI output materially formed the basis of the
final artifact. It does not count automatically because an AI was involved at
some point. There is no percentage-of-text test - inventing one would imply a
precision that is not being measured.
Multiple candidates. If ten alternatives are generated and one of them
materially forms the basis of what is used, that is one used output, not ten.
An artifact is never counted twice because several revisions were generated on
the way to it.
The gap between "generated a lot" and "used any of it" is the entire subject of
this experiment, so the definition is drawn tightly on purpose.
Cost and time accounting
Annual subscriptions are amortised at one twelfth of the annual price for
the window. No annual plan is counted at its full price.
Free trials and credits are recorded at zero cash cost, but the usage and
review time they consume is counted in full, and they are listed separately so
they cannot flatter the totals. Any trial that would auto-convert to a paid plan
is noted.
Tax is recorded separately; line items are compared ex-tax for
comparability. Refunds and credits received within the window are netted
off and disclosed.
Shared or multi-purpose tools are allocated by logged AI-session minutes as
a share of total logged minutes for that tool. The allocation is published even
where it is unflattering.
Time is logged per session in minutes - prompting, waiting, reviewing,
correcting, re-running - recorded contemporaneously, never reconstructed.
Operator time is valued at USD $100/hour, a declared internal decision rule
for this experiment only. It is not a claim about market compensation or
opportunity cost.
How the logs are kept
The used-output log and the time log are the evidence this experiment rests on,
so the rule for keeping them is stated before it starts. This is a procedural
commitment, not a technical guarantee: nothing here is enforced by software, and
it is not claimed to be.
Entries are recorded the same day, before the result is known.
Existing entries are not silently rewritten.
Where an entry needs correcting, the correction is added as a dated note
alongside the original, so both the original entry and the change remain
visible.
The KEEP / FIRE rule
A line item is FIRE if any of the following hold, judged only from
contemporaneous logs after the window closes:
No use. It produced zero used outputs during the window.
Substitutable. Every used output it produced could have been produced by
another retained line item, where the specific substitute was named in the
log at the time the output was created.
Negative time economics. Its allocated cost for the window exceeded the
value of the net time it saved, where net time saved is the sum, across its
used outputs, of the pre-logged unaided estimate minus the logged actual
time, valued at $100/hour.
Otherwise it is KEEP.
Rule 3 depends on a pre-logged estimate of how long a task would have taken
unaided. That estimate is an operator judgement recorded before the task is
performed, not a measurement, and it is reported as such. It deliberately has
no influence on whether the hypothesis passes.
Evidence retained
Invoices and receipts, usage dashboard exports, the daily used-output log, the
session time log, and the dated pre-registration itself.
Personal data, account identifiers, private correspondence and commercially
sensitive material are redacted, and the fact of redaction is disclosed.
Credentials, API keys and secrets are never published as evidence in any form.
Vendor names and prices are published - they are not confidential.
Known confounders
Published here rather than discovered later:
A single operator, a single business, a single month. This is a case, not a
general law.
Workload is not constant. A month of heavy building and a month of heavy
writing consume AI differently.
Novelty effect may inflate early usage of anything recently added.
No blinding is possible. The operator knows the experiment is running, which
can itself change purchasing and usage behaviour.
Tool prices and plans may change mid-window.
"Used" is a self-assessed judgement. The only mitigation is same-day logging,
before the outcome is known.
What the used-output metric cannot see. The definition is narrow on purpose,
and narrowness has a price. It will undercount at least two kinds of real value:
Preventive value - catching an error before a bad change ships. It produces
no artifact, so it does not count as a used output unless it directly produced
a qualifying logged decision under the definition above.
Learning and explanation - understanding something better with an AI's
help, where no qualifying artifact or directly logged decision results.
Neither is quietly promoted into a used output to fix the shortfall. Both are
disclosed as a limitation in the published result, so a line item that scored
badly is not read as having produced nothing.
Abandon conditions
The experiment is abandoned, and published as abandoned with its reason, if:
logging lapses for more than three consecutive days
a major workload disruption makes the month unrepresentative
an accounting change makes spend unattributable
Insufficient evidence
The experiment completes but declines both an outcome and a decision if either:
fewer than 10 used outputs are recorded across the window, or
more than 20% of total in-scope AI spend cannot be allocated under the
rules above
In that case the result is published as complete with insufficient evidence and
the reason attached. It is not rounded to FAIL, and it is not rounded to FIRE.
Status
Abandoned on 2026-08-27. The measurement window opened on 2026-08-22 and
closed early, on day 6 of 30, when the abandon condition published above fired:
logging lapsed for more than three consecutive days.
The four ledgers this experiment rests on - spend, outputs, time and
corrections - hold no entries for any day of the window. Under the rules
published above an unlogged day cannot be reconstructed afterwards, and nothing
has been backfilled, estimated or inferred to close the gap. There is therefore
no data, no removable-spend percentage and no verdict: no PASS, no FAIL, no
KEEP, no FIRE, and no insufficient-evidence ruling either, which is reserved
for an experiment that completes.
This is the outcome the abandon condition was written to produce. It was
published before the window opened so that a lapse would end the experiment in
the open rather than be quietly repaired later, and applying it costs less than
the credibility of a result assembled from memory would.
The method above is unchanged and stands as published on 2026-08-21. A future
audit of Ovriko's AI spend would be a new experiment with its own
pre-registration, not a continuation of this one.
The operating rules behind every Ovriko experiment - how a hypothesis is fixed before any money is spent, how cost and time are counted, what PASS, FAIL, KEEP and FIRE mean, and the things Ovriko will not publish at any price.