What we paid for AI in 30 days, and what we'd pay for again
A pre-registered audit of Ovriko's 30-day AI spend, abandoned on 2026-08-27 under its own published abandon condition after logging lapsed for more than three consecutive days.
Research
The operating rules behind every Ovriko experiment - how a hypothesis is fixed before any money is spent, how cost and time are counted, what PASS, FAIL, KEEP and FIRE mean, and the things Ovriko will not publish at any price.
Published
An observation, offered as an observation and not a measurement: a great deal of the advice written about AI and business appears never to have been tested by the person giving it. It reads as though it were assembled from announcements, from other people's threads, and from the reasonable-sounding assumption that a tool which demonstrates well also works on a Tuesday afternoon under real conditions with real money attached.
That is a suspicion, not a finding. It is the reason this site exists, and it is evidence of nothing.
Ovriko is an attempt to do the boring version instead: run the thing, count what it cost, and write down what happened.
This document is the contract that work runs under, published before there are any results - the only point at which rules like these are credible, since rules written afterwards can be shaped to fit the numbers.
This is a description of method, not a claim of achievement. At the time of writing, Ovriko has not completed an experiment and has no findings to point at.
What it has is a set of commitments specific enough to be broken visibly. A threshold published before an experiment starts can be checked against the one reported after it ends. If those two numbers ever differ without an explanation, this document is the evidence.
Four things get published here, and they do not get to look the same.
A measured fact is a number that came from an instrument or a record: an invoice, a log, a dashboard export, a stopwatch. It has a source, and the source is named.
An observation is something noticed while working. It is often the reason an experiment gets designed in the first place. It proves nothing on its own, and it is labelled as commentary.
An interpretation is what Ovriko thinks a measured fact means. It is the most fragile category - where a real number gets turned into a story - and it is marked as interpretation rather than smuggled in beside the data.
An opinion is an opinion. It is allowed, and it is labelled.
Collapsing these four into one confident voice is how business writing misleads people who are about to spend money, and it is usually done without meaning to.
Every experiment starts with a claim that could turn out to be wrong, and a threshold that decides the matter. Both are written down before anything is bought, built or measured.
The threshold is a number. Not "meaningfully faster" or "a clear improvement" - a number, chosen in advance, with the reasoning for choosing it stated. A threshold is a business decision made under uncertainty, not a fact derived from evidence, and it is presented that way.
What changes and what deliberately stays the same are defined up front, so a result can be attributed rather than assumed. Metrics and the cost model are fixed before the first measurement.
The published cost of an experiment is not the subscription price. It is everything the experiment consumed:
Time is a cost. Where operator time is priced, the hourly figure used is declared inside the experiment, and it is a declared internal decision rule - not a claim about market rates or what anyone's hour is worth in general.
Failed spend counts. Money burned on a path that went nowhere is reported rather than quietly dropped from the total, which is the most flattering and least honest form of cost accounting available.
Ovriko's experiments are small. One business, one operator, one period, one set of conditions. That is a real limitation and it does not get buried in a closing paragraph.
A single run answers what happened here. It does not answer what will happen for you, and no amount of confident writing changes that. Where a result rests on a small sample, a short window, or a single operator's judgement, the write-up says so in the same place it reports the number.
A finished experiment is judged on two independent axes. They answer different questions, and they are allowed to disagree.
PASS and FAIL answer one question only: did the hypothesis meet the threshold declared before the test? Nothing about preference, satisfaction or whether the thing felt good to use enters into it.
KEEP and FIRE answer a different question: given the full cost, including time, does this stay in the business or get removed?
All four combinations are real, and the interesting ones are the disagreements.
A tool can clear its threshold and still be removed - it did what was predicted, but the supervision it needed cost more than the work it returned. That is PASS + FIRE, and it is the result most likely to be useful to someone about to buy the same thing.
A tool can miss its threshold and still be kept, because it turned out to earn its cost for a reason the hypothesis never anticipated. That is FAIL + KEEP, and publishing it requires saying plainly that the original claim was wrong.
Not every experiment produces an answer, and the ones that do not are worth publishing anyway.
An experiment designed but not started is planned. One in progress is running, and carries no verdict, because a verdict before the end of a measurement window is just a guess with formatting. One stopped without a usable result is abandoned, published with the reason.
And an experiment can finish, be measured, and still fail to support a verdict - too little data, too much of the cost unattributable, conditions that changed underneath it. That is recorded as insufficient evidence, with the reason. It is not rounded down to FAIL or FIRE. Neither of those is what "we could not tell" means, and the difference matters to anyone deciding something on the back of it.
Any relationship that could bend a conclusion is disclosed inside the experiment it touches.
Buying a tool is not one of those relationships. Ovriko pays for the software it tests, as an ordinary customer, and being a vendor's customer is expected rather than disclosable - the first experiment is an audit of exactly those bills.
What gets disclosed is money or influence moving in the other direction. As of publication: no vendor has paid Ovriko, no vendor sponsors it, it receives no affiliate or referral compensation from any vendor it tests, and no vendor has been given any say in a conclusion.
Any of that can change, and if it does it is disclosed in the experiment it affects rather than in a page nobody reads. So is anything else received because Ovriko is testing or publishing about a vendor: free access, a discount, credits, gifts, sponsorship, or any other consideration.
Experiments need real businesses to run inside. Where a partner business cannot be named, it is anonymised, and the fact that it has been anonymised is stated.
Some evidence cannot be published. Personal data, partner records and commercially sensitive material stay unpublished even when including them would make a result more convincing. Where that happens the redaction is disclosed rather than hidden, so a reader can see the shape of what is missing and discount the conclusion accordingly.
Vendor names and prices are not confidential and are published.
Some of the numbers published here will turn out to be wrong. When that happens, the correction is added to the piece as a dated note explaining what changed and why. Numbers are not silently edited into being right.
A publication that quietly fixes its mistakes is indistinguishable from one that never made any.
Nothing has been published here yet beyond this document and the pre-registered design of the first experiment. Its hypothesis, its threshold and its cost rules are public before its measurement window opens, so they can be checked against whatever gets reported at the end.
Anyone can claim to work this way. The only version of the claim worth reading is the one where the rules were published first and are still there afterwards, to be held against the result.
Topics
Related
A pre-registered audit of Ovriko's 30-day AI spend, abandoned on 2026-08-27 under its own published abandon condition after logging lapsed for more than three consecutive days.