The method

How a verdict is decided

The method is published before any record is issued, and it does not change to fit a result. This page is the version a reader can hold a record against. If a record and this page disagree, the record is wrong.

Everything here is written to be argued with. A method that cannot be attacked cannot be trusted, so the object being graded, the tolerances, the sources and the verdict set are all fixed in public before anything is run.

What a record is about

A record is never about a project. It is about one exact thing, observed once, and that thing is identified completely enough that someone else could set it up and disagree.

Those seven fields are the whole subject of a record. Change any one of them and it is a different subject, which needs a different record.

This is not bookkeeping. It is the line between a finding and an accusation. A statement about a named version under a named configuration at a named time is something that was observed. A statement about a project is a claim about everything it will ever do, and this register does not make those.

How the observation is made

A run happens from outside the tool, over the tool's own interface, the same way anything else would call it. No test hooks are installed, no internals are reached into, and nothing about the tool is modified to make it observable. If a failure can only be seen by instrumenting the tool from the inside, it is not the failure this register grades, because that is not the failure a caller experiences.

Every tolerance is written down before the run, not after it. What counts as a match, how much drift is acceptable, how many places a number is compared to, what staleness is allowed: all of that is fixed in the suite in advance, so a result cannot be rescued by loosening the definition once it is known.

Findings are compared against a source published independently of the test and named in the record. Nothing is graded against an expectation this register invented, and nothing is graded against a second tool's opinion.

The five verdicts

There are five, they are the whole set, and a record carries exactly one. There is no average, no aggregate, and no score.

VerdictWhat it means
PASSThe tool returned the correct answer within the tolerances fixed before the run, and the tolerances are stated in the record
FAIL_SAFEThe tool did not return a correct answer, and it said so. It raised, refused, returned an explicit error, or otherwise signalled that the caller should not proceed
FAIL_UNSAFEThe tool returned something wrong and presented it as if it were right. Nothing in the response tells a caller not to use it
INDETERMINATEThe observation could not decide. Upstream was unavailable, the environment was not reproducible, or the evidence does not distinguish between explanations
OUT_OF_SCOPEThe case falls outside what the tool undertakes to do, so there is nothing here to grade

Two rules about that set are not negotiable.

FAIL_SAFE beats FAIL_UNSAFE, and it is not PASS. A tool that stops when it cannot answer is doing the correct thing and is graded well above one that guesses. It still did not answer, so it is not a pass. Collapsing those two into one word is the thing that makes existing reliability language useless.

INDETERMINATE is a real verdict, not an absence. It gets published like any other. A register that quietly drops the runs it could not decide is reporting a filtered view of its own evidence, and its published results stop meaning anything.

The loss cause taxonomy

Every deviation is classified before it is recorded, so a record says what kind of failure it found and not only that it found one. The taxonomy is versioned from the first record it is applied to, because a classification scheme that changes silently makes its own history unreadable. Version v0 covers eight causes.

CodeFailure
silent_emptyReturned an empty result where data existed, with no signal
stale_valueReturned cached or outdated data presented as current
wrong_entityReturned data for a different entity than the one requested
fabricated_fieldReturned a field populated with a value that has no upstream basis
partial_truncationReturned a subset of the result with no truncation signal
unsignaled_fallbackServed from a fallback source without disclosing the substitution
auth_degradationSilently downgraded to a lower privilege tier and returned reduced data as complete
schema_driftUpstream shape changed and the adapter absorbed it into a wrong but valid response

Each code is meant to be decidable from a run artifact without a judgment call, which is what keeps the classifications consistent as the corpus grows. A cause that needs an opinion to apply does not belong in the taxonomy, and if one turns out to need an opinion it gets replaced in a new version rather than reinterpreted in the old one.

Three kinds of record, never merged

A record answers one question, and which question it answers is part of what it is.

They are never combined into a single result. A tool can be correct in isolation and wrong in a sequence, and a caller who needs to know which one broke is not helped by a number that averaged them.

What a record says, and what it refuses to say

A record does say that these specific inputs were sent to this exact subject at this time, that these outputs came back, that they were compared against this named ground truth source, that deviations were classified under this published taxonomy, and that anyone holding the record can rerun the suite and get the same result.

A record does not say that the subject is safe, secure, or free of vulnerabilities. It does not say the subject is fit for your use case. It does not say the subject will behave the same way tomorrow, or on a different commit, or against a different upstream. It does not say the maintainer is trustworthy, and it does not say they are untrustworthy. It says nothing in a currency: there is no coverage limit, no trust figure, and no number that looks like one.

This register is an evidence provider, not a custodian. It holds no funds, insures nothing, and settles nothing. Where a remedy is warranted, that is a licensed counterparty's product, and a record is an input to it.

Records expire by construction

Every record carries a validity window. Queried outside that window it returns EXPIRED, whatever the verdict was. Not "last checked a while ago", not a stale date printed next to a green tick. The schema cannot represent a pass that is out of date, because the ability to represent one is the thing that gets abused.

There will never be a badge image. An image a maintainer copies into a readme is a claim that outlives its evidence. It keeps rendering after the record expires, keeps rendering after the record is withdrawn, and cannot be revoked by the party that issued it. That exact mechanism is how this category of business has failed before. A mark here is resolved against the register, live, or it does not resolve at all.

Who pays

The buyer pays. Never the graded party, for its own first verdict. A maintainer cannot commission a record on its own package, cannot preview one before publication, and cannot negotiate one. The moment revenue depends on the subject's satisfaction, every verdict becomes a negotiation and the output is worth nothing. This constraint is expensive early. It is also the entire thing.

One exception exists and it is narrow. A maintainer may fund a re-test of a subject already in the register, after a fix. Three conditions are mandatory and all three are published on the record itself: the funding source is disclosed on the record, the methodology is the already-published suite and cannot be altered for the run, and the result publishes regardless of outcome. A vendor who pays for a re-test and fails gets a FAIL_UNSAFE record with their name on the invoice line.

Right of reply, and what is withheld while it runs

No adverse record publishes on the day it is written. The party responsible for the subject receives the complete evidence and a fixed window to answer before anything appears.

If they show the finding is wrong, the record is withdrawn, and the withdrawal publishes in its place rather than the record quietly disappearing. If they fix the subject, the record still stands against the version that was tested, because that version was observed and the observation was accurate. A new record can then be issued against the fix. If they say nothing, the window closes and the record publishes as written.

Before the window opens, the finished record is hashed and the hash is committed. When the record publishes, anyone can verify that what was published is what was committed to, so a subject cannot negotiate a finding softer during their own reply window.

The disclosure rule

The register may say that a record is held and how many are held. It may not say what a held record found, not even its verdict, and it may not name who the record is about.

The second half is the part that is easy to get wrong. Announcing that a finding is held against a named tool publishes the accusation and withholds only the evidence, which is worse than publishing both. So while anything is held, the register names no held subject at all.

That gate is enforced in the code that builds the register rather than by remembering to apply it. The page is compiled from the records on every build, and the build refuses to write the page if a held record's identifier or its subject's name reaches the output by any path.

The method has a version

Every record names the version of this method it was decided under, and this page is that version in readable form. When the method changes, the change is published as a new version and old records keep pointing at the version they were actually decided under. They are not silently re-interpreted under new rules, and they are not quietly regraded.

A method that can be edited after a result is a method that can be edited to fit a result.

Where the method actually is

This page is the method in readable form. The method in runnable form is at github.com/arian-gogani/nobulex-registry, MIT licensed: the harness, the probes, the loss cause taxonomy, the record generator, the renderer, and the publication gate described above. Read it, run it, and attack it. A verdict produced by a suite nobody can inspect is not evidence. It is an opinion with a procedure attached to it.

Every claim on this page is a promise about behavior, and a promise you cannot check is worth what its author is worth. One of them you can settle right now, from a terminal, without trusting anything said here. The register page is generated by that code and never edited by hand, which means the file brand/register.html in the repository and the bytes served at nobulex.com/register are not two versions of a page. They are one file.

curl -sS https://nobulex.com/register | shasum -a 256

That prints 4b8aa822c5883b1022ef1aa1459768a2c024a64674ca823eae5ce24acfcf9051, and so does shasum -a 256 brand/register.html in a clone. If the two ever disagree, a person edited the register after the code produced it, and every other claim on this page should be read in that light. Send the mismatch and it publishes.


The register is the only page here that states anything about a record, and it is the only page here that is not written by hand. If a record contradicts this method, the record is wrong and the correction publishes beside it.

Disputes and corrections: nobulex.dev@gmail.com. If you maintain something recorded here and believe a finding is wrong, send the evidence. It publishes beside the record, unedited.

Read the register →