Research

How Ovriko tests, and what it will not publish

The operating rules behind every Ovriko experiment - how a hypothesis is fixed before any money is spent, how cost and time are counted, what PASS, FAIL, KEEP and FIRE mean, and the things Ovriko will not publish at any price.

Published

An observation, offered as an observation and not a measurement: a great deal of the advice written about AI and business appears never to have been tested by the person giving it. It reads as though it were assembled from announcements, from other people's threads, and from the reasonable-sounding assumption that a tool which demonstrates well also works on a Tuesday afternoon under real conditions with real money attached.

That is a suspicion, not a finding. It is the reason this site exists, and it is evidence of nothing.

Ovriko is an attempt to do the boring version instead: run the thing, count what it cost, and write down what happened.

This document is the contract that work runs under, published before there are any results - the only point at which rules like these are credible, since rules written afterwards can be shaped to fit the numbers.

What this is, and what it is not

This is a description of method, not a claim of achievement. At the time of writing, Ovriko has not completed an experiment and has no findings to point at.

What it has is a set of commitments specific enough to be broken visibly. A threshold published before an experiment starts can be checked against the one reported after it ends. If those two numbers ever differ without an explanation, this document is the evidence.

Observation is not evidence

Four things get published here, and they do not get to look the same.

A measured fact is a number that came from an instrument or a record: an invoice, a log, a dashboard export, a stopwatch. It has a source, and the source is named.

An observation is something noticed while working. It is often the reason an experiment gets designed in the first place. It proves nothing on its own, and it is labelled as commentary.

An interpretation is what Ovriko thinks a measured fact means. It is the most fragile category - where a real number gets turned into a story - and it is marked as interpretation rather than smuggled in beside the data.

An opinion is an opinion. It is allowed, and it is labelled.

Collapsing these four into one confident voice is how business writing misleads people who are about to spend money, and it is usually done without meaning to.

The hypothesis comes first

Every experiment starts with a claim that could turn out to be wrong, and a threshold that decides the matter. Both are written down before anything is bought, built or measured.

The threshold is a number. Not "meaningfully faster" or "a clear improvement" - a number, chosen in advance, with the reasoning for choosing it stated. A threshold is a business decision made under uncertainty, not a fact derived from evidence, and it is presented that way.

What changes and what deliberately stays the same are defined up front, so a result can be attributed rather than assumed. Metrics and the cost model are fixed before the first measurement.

What a cost actually is

The published cost of an experiment is not the subscription price. It is everything the experiment consumed:

  • licences, usage charges and credits
  • setup time
  • supervision while it ran
  • review of whatever it produced
  • rework when the output was wrong
  • money spent on approaches that were abandoned

Time is a cost. Where operator time is priced, the hourly figure used is declared inside the experiment, and it is a declared internal decision rule - not a claim about market rates or what anyone's hour is worth in general.

Failed spend counts. Money burned on a path that went nowhere is reported rather than quietly dropped from the total, which is the most flattering and least honest form of cost accounting available.

One run is a case, not a law

Ovriko's experiments are small. One business, one operator, one period, one set of conditions. That is a real limitation and it does not get buried in a closing paragraph.

A single run answers what happened here. It does not answer what will happen for you, and no amount of confident writing changes that. Where a result rests on a small sample, a short window, or a single operator's judgement, the write-up says so in the same place it reports the number.

The four verdicts

A finished experiment is judged on two independent axes. They answer different questions, and they are allowed to disagree.

PASS and FAIL answer one question only: did the hypothesis meet the threshold declared before the test? Nothing about preference, satisfaction or whether the thing felt good to use enters into it.

KEEP and FIRE answer a different question: given the full cost, including time, does this stay in the business or get removed?

All four combinations are real, and the interesting ones are the disagreements.

A tool can clear its threshold and still be removed - it did what was predicted, but the supervision it needed cost more than the work it returned. That is PASS + FIRE, and it is the result most likely to be useful to someone about to buy the same thing.

A tool can miss its threshold and still be kept, because it turned out to earn its cost for a reason the hypothesis never anticipated. That is FAIL + KEEP, and publishing it requires saying plainly that the original claim was wrong.

When there is no verdict

Not every experiment produces an answer, and the ones that do not are worth publishing anyway.

An experiment designed but not started is planned. One in progress is running, and carries no verdict, because a verdict before the end of a measurement window is just a guess with formatting. One stopped without a usable result is abandoned, published with the reason.

And an experiment can finish, be measured, and still fail to support a verdict - too little data, too much of the cost unattributable, conditions that changed underneath it. That is recorded as insufficient evidence, with the reason. It is not rounded down to FAIL or FIRE. Neither of those is what "we could not tell" means, and the difference matters to anyone deciding something on the back of it.

What Ovriko will not publish

  • Invented results, measurements, costs, company names, testimonials or audience numbers. Not in articles, not in drafts, not as placeholder text in an unfinished interface.
  • Numbers whose source cannot be shown.
  • A single unrepeated run presented as proof.
  • Conclusions shaped by whoever paid for them.
  • Advice that has not been tested here first.

Money, incentives and disclosure

Any relationship that could bend a conclusion is disclosed inside the experiment it touches.

Buying a tool is not one of those relationships. Ovriko pays for the software it tests, as an ordinary customer, and being a vendor's customer is expected rather than disclosable - the first experiment is an audit of exactly those bills.

What gets disclosed is money or influence moving in the other direction. As of publication: no vendor has paid Ovriko, no vendor sponsors it, it receives no affiliate or referral compensation from any vendor it tests, and no vendor has been given any say in a conclusion.

Any of that can change, and if it does it is disclosed in the experiment it affects rather than in a page nobody reads. So is anything else received because Ovriko is testing or publishing about a vendor: free access, a discount, credits, gifts, sponsorship, or any other consideration.

Experiments need real businesses to run inside. Where a partner business cannot be named, it is anonymised, and the fact that it has been anonymised is stated.

Privacy limits the evidence on purpose

Some evidence cannot be published. Personal data, partner records and commercially sensitive material stay unpublished even when including them would make a result more convincing. Where that happens the redaction is disclosed rather than hidden, so a reader can see the shape of what is missing and discount the conclusion accordingly.

Vendor names and prices are not confidential and are published.

Corrections

Some of the numbers published here will turn out to be wrong. When that happens, the correction is added to the piece as a dated note explaining what changed and why. Numbers are not silently edited into being right.

A publication that quietly fixes its mistakes is indistinguishable from one that never made any.

What happens next

Nothing has been published here yet beyond this document and the pre-registered design of the first experiment. Its hypothesis, its threshold and its cost rules are public before its measurement window opens, so they can be checked against whatever gets reported at the end.

Anyone can claim to work this way. The only version of the claim worth reading is the one where the rules were published first and are still there afterwards, to be held against the result.

Related