PASS
Correct within a tolerance that was pinned before execution, against a named ground truth source.
N O B U L E X
The independent reliability registry for agent tools Buyer funded, never vendor funded
Payment rails prove money moved. Nobulex proves what happened on the other side.
Financial data first, because it is the one place the right answer is checkable.
Package registries check what a tool is · Scanners check what it contains Nobulex checks whether the answer it returned was true.
THE FAILURE
Every agent that calls a tool inherits that tool's failures, and the failures that matter are not the loud ones. A tool that raises an error is workable, because you can retry it or route around it. The dangerous case is the tool that returns something well formed, plausible and materially wrong: empty where data existed, stale while claiming to be current, scoped to a different entity, truncated with no signal, served from an undisclosed fallback. Nothing raises. The schema validates. The agent proceeds.
The failure has a name here: silent semantic corruption. It passes every schema check and every vulnerability scanner, because nothing about it is malformed. It is simply not true. A person reading a stock price that is four days stale may notice. An agent will trade on it.
Uptime does not detect it. Stars do not detect it. A green test suite written by the same people who wrote the tool does not detect it, because the fixture and the bug were authored by the same hand.
One test governs this register: does it fail loud, or does it lie quiet?
THE FIVE VERDICTS
Never an average, never a percentage, never a grade out of ten. A single number invites the exact behaviour the verdicts exist to prevent, which is glancing at a figure instead of reading what was actually tested.
Correct within a tolerance that was pinned before execution, against a named ground truth source.
Could not answer, and said so. The tool refused, errored, or returned an explicit null. Still a failure, and not equivalent to a pass.
Returned something plausible and materially wrong, with no error signal. This is the verdict the whole register exists to detect.
No ground truth was available to decide the question. Not a pass, and never reported as one.
Outside the conditions pinned for this run. Not a pass either, and recorded rather than quietly dropped.
A tool that stops when it cannot be sure has left you something to route around. A tool that guesses confidently has not. How each verdict is decided →
THE SUBJECT
An assay office does not certify a silversmith. It tests one object and records what was tested, what standard was met, who tested it, and when. The mark travels with the object, not the maker. Nobulex grades a seven part tuple, and changing any one part makes it a different subject with no inherited history.
There is no badge image and there never will be. A graphic sitting on someone else's site is a claim about the past that keeps asserting itself in the present. Status is resolved against the register, live, or it is not a status.
HOW A RECORD IS MADE
The suite speaks to a subject the way an agent would, over its own protocol, and never imports its code. A test that imports the thing it is testing has already agreed to the thing's own idea of what happened. Tolerances are fixed before the run, so a result cannot be reinterpreted into a pass after the fact.
WHAT A VERDICT WARRANTS
| A record does say | A record does not say | |
|---|---|---|
| Scope | This exact build, in this configuration, was given these inputs at this time | That a different version, config or upstream will behave the same way |
| Finding | Under the pinned conditions that behaviour was correct, safely refused, or materially wrong | That the tool is secure, or free of vulnerabilities |
| Subject | What the object did when it was observed | That the vendor is trustworthy. Objects are tested, not companies |
| Time | The observation time and the window the record is valid for | That the result still holds today. Check the window |
| Liability | That the evidence is reproducible by anyone running the same suite | That any loss is indemnified. This is an evidence provider, never a custodian |
WHERE THIS STARTS
Most claims an agent tool makes cannot be graded. Ask a summariser whether its summary was faithful and there is no authority to appeal to. Ask a market data tool what a stock closed at on a given day and there is exactly one right answer, published by a named authority, in writing, before anyone went looking for it. That property is rare, and it is the reason the register starts here rather than somewhere larger.
A register that opened on the hardest domain would be unfalsifiable on day one. Starting where the answer is checkable means every early record can be argued with, and that is the only thing that makes a later one worth anything.
RIGHT OF REPLY
When a record finds something, the party responsible for the object receives the complete evidence and a fixed window to answer before anything is published. If they show the finding is wrong, the record is withdrawn and the withdrawal is published in its place. If they fix the object, the record still stands against the version that was tested, and a new record can be issued against the fix. If they say nothing, the window closes and the record publishes as written.
The rule this register holds itself to is deliberately narrow. It may say that a record is held and how many are held. It may not say what a held record found, not even its verdict, and it may not name who the record is about.
The second half of that is the part that is easy to get wrong. A register announcing that a finding is held against a named tool has published the accusation and withheld only the evidence, which is the worst of the two options. So while anything is held, no held subject is named at all.
Before the window opens, the finished record is hashed and the hash is written down. When the record publishes, anyone can check that what was published is what was committed to, and that nothing softened while the subject was replying.
The register is not a party to a dispute it records. It publishes what it observed, and what the observed party said back.
WHERE THIS STANDS
A register's first duty is to describe itself as accurately as it describes anything else.
01
How many records exist, how many are held, and what any of them found lives on the register, which is compiled from the records every time this site is built. Nothing about a record is retyped here. A second hand-kept copy of a register is how a name under embargo eventually gets published by accident.
02
Tolerances are pinned before a run and the methodology is written down. The suite that produces the verdicts has not been released, so no independent party has re-run one and reached the same answer. Until that happens, every verdict rests on one party's word. That is a weakness, and it is named here rather than left to be found.
03
The party being graded does not pay, is not offered a way to pay, and cannot buy a re-run on friendlier conditions. What a subject is owed is the evidence and a window to answer it. Nothing else is for sale to them, because a register the graded party funds is a brochure with a serif typeface.
04
This is not a certification. No regulator recognises it, no insurer prices off it, and it carries no liability for anyone's loss. It is a record of what a specific object did when it was observed, under conditions anyone can read. That is the whole of the claim.
05
Nobulex is built and run by one person. Stated here rather than discovered later, because a register that overstates itself has already failed the test it applies to everything else.
The register is the only page here that says anything about a record, and it is the only page here that is not written by hand. Read the register →
FAQ
Nobulex is an independent reliability registry for agent tools. It tests exact package versions against a named ground truth source and publishes a reproducible record of whether each one failed loud or lied quiet.
A response that is well formed and plausible but materially wrong: stale, mis-scoped, truncated, empty where data existed, or served from an undisclosed fallback. It passes every schema check and every scanner, because nothing about it is malformed. It is simply not true.
A subject tuple, never a project name: the package, an exact version or commit, a stated configuration, a named upstream source, the execution environment, a versioned test suite, and an observation time. Change any one of those and it is a different subject.
No. A record carries one of five verdicts: PASS, FAIL_SAFE, FAIL_UNSAFE, INDETERMINATE, or OUT_OF_SCOPE. There is no average, no percentage and no grade out of ten, because a single number invites a buyer to glance at a figure instead of reading what was actually tested.
The buyer, never the graded party. The moment revenue depends on the subject's satisfaction, the verdicts are worth nothing.
The maintainer receives the full run artifact and has seven days to answer. The answer publishes beside the record, unedited. The verdict is fixed before the window opens and the window cannot change it; only a new run under newly pinned conditions can.
THE REGISTER
Every count and every subject on the register is compiled from the records on each build. This page states none of it, on purpose.
Corrections and disputes: nobulex.dev@gmail.com. If you maintain something recorded here and believe a finding is wrong, send the evidence. It publishes beside the record, unedited.