Methods

How we measure — and how we label

The plain-language key

The rest of this site talks about an agent the way you would talk about a person you are hiring. This page and the personnel files use the words the measurements are actually written in. They mean the same things, and the translation never makes a result sound better than it is.

on the rest of the sitein the personnel files
trial tasks, practice tasksthe development set — the requests we are allowed to learn from
its final examthe sealed set — tasks locked away before the work started, opened once, then retired
it has not sat its final exam yetsealed run pending
the standard we set beforehandthe bar declared in advance
its reference checkthe scorecard
what not to assign it yetthe known limitations of a measured result
its personnel filethe technical report for that duty
a dutya declared outcome, with its declared importance and its bar

One of those translations could hide something if we let it. "Trial tasks" sounds gentler than "development set", so the rule we hold ourselves to is that a trial-task result never appears without the sentence saying the final exam has not been taken. If you find one that does, it is a defect and we want to hear about it.

Evaluation protocol

  • · 997-case evaluation pool, set aside before we started — the assistant never saw it — curated from real banking-request shapes.
  • · Temperature-0 repeat runs to measure routing variance across identical inputs.
  • · Adversarial permission probes for above-permission tool selection.
  • · Scored outside NeMo Evaluator; in-service re-scoring is in progress.

Grade rule

A pack's grade is "cleared m of n declared goals", always stated with the set it was measured on — we do not compute letter grades.

Provenance legend

MEASURED

A number we produced from an actual eval run. Always cites basis (test set) and n.

DECLARED

A choice we made — importance, target, threshold. Not measurement; product decision.

ASSUMED

An input the buyer sets themselves (e.g. requests per month). Yours, not ours.

MODELED

An output computed from ASSUMED inputs. Cannot become a Blade claim.

UNMEASURED

Not yet measured. Rendered as 'not yet measured' — never blank, never implied-good.

What we do not do

  • · No letter grades.
  • · No opportunity/market-satisfaction figures publicly.
  • · No customer logos — our first customer is ourselves.

How a measurement earns the right to decide something

We borrowed this from drug regulation, because the problem is the same one: a number that is easy to measure gets used to stand in for the thing you actually care about, and sometimes it moves the right way while the real thing moves the wrong way. The convention has three steps. A candidate measure is still being evaluated — useful for steering our own work, not for deciding anything. A reasonably likely measure is one the field accepts as tracking the thing that matters — good enough to decide with, as long as its known weaknesses are printed next to it. A validated measure has evidence that moving it actually moves the buyer's real-world result, and accounts for that whole effect.

Where we stand, stated plainly: nothing we own has reached the third step, for any measure. The check behind the outcomes on this site is at the first step — it is our own check, it tops out at a perfect score too easily, and it has never been validated against a customer's realised result. That is why every outcome on this site is a preview, and why the next thing we ship is a measurement the rest of the field already accepts rather than a better version of our own.

The grade rule

"Cleared m of n declared goals." m counts only the goals whose measured result clears the bar declared in advance, on the stated set. n counts every declared goal, measured or not. Goals we have not measured are reported as not yet measured — never guessed, never blank, and never counted as cleared. There are no letter grades and no weighting: a goal you told us matters 10 out of 10 and a goal that matters 7 out of 10 both count once, and the importance is shown beside each one so you can weight them yourself.

The set is always stated inside the grade, because "3 of 3 on a development set" and "3 of 3 on a locked-away set" are different claims and must never look identical.

What PREVIEW means

A preview card carries real measurements from a real evaluation run. What it does not carry is a run against a locked-away set.

The distinction is worth the extra word. A development set is a set of requests the assistant has not seen, drawn from the pool we are allowed to work with — we can measure against it repeatedly, and we do. A locked-away set is used for nothing at all until the day it decides something: no training, no choosing between versions, no designing the test itself, and no looking at individual answers. It is opened once and then retired, so the same questions cannot be reused to flatter a later result.

For the banking router that set exists — 400 requests, drawn by a fixed rule from the requests no part of our work has touched, fingerprinted, and unopened. We have deliberately not opened it, for a reason we would rather publish than hide: the check we would score it with has not earned the right to decide anything yet, and spending a single-use set to obtain a number we are not allowed to publish as a decision would waste the most valuable asset we have. One honest disclosure about that set: it was included once in an all-requests-at-once average in July, before any of this, which produced no per-request record and changed nothing about the assistant. So it is unused rather than never-touched, and this sentence exists so nobody can say we implied otherwise.

Why we publish failures

Because a scorecard that only contains passes is an advertisement, and you already know how to read one of those.

Three things follow from that, and all three cost us something. A declared goal that fails stays on the front page, in the same size type as the ones that passed, sorted by how much it matters rather than by how it turned out. A number we published and later found unsound comes off the page and the reason goes on the corrections page with the date — we supersede, we do not delete. And a result we cannot support gets no replacement number at all, because filling the hole with a weaker claim is the same mistake wearing different clothes.

The most useful thing we can tell you about this assistant is what it does badly, since that is the part you have to design around. That is the permission finding, and it is one click from the front page.

Public benchmarks — anyone can re-run these

These answer one question: how does this assistant do on a test the whole field already uses, on requests that are published, against a fixed answer key. The point of them is that you do not have to take our word for anything — you can run them yourself.

Nothing is published here yet. The accepted public test for picking the right action is BFCL, and version 3 of its structural check is the one that applies to us. We have the software, the settings are registered, and we have run zero jobs against it. When it lands, three things appear with it: the version, since version 3 and version 4 are not comparable and the larger vendors' cards quote version 4; the result broken out by category with a count for each, because a single average hides where a tool-picker actually fails; and two weaknesses of the test itself, which we print whether or not they flatter us — an audit of tool-calling benchmarks found roughly 18.5% of disagreements between the answer key and a careful re-reading, and the same kind of checker has been measured swinging up to 10 points between two ordinary settings.

Separately, we re-run a public knowledge test — MMLU-Pro — against our own copy of a model whose score its maker has already published. This checks our measuring equipment, not the assistant: if we cannot reproduce someone else's published number on their own model, no number of ours means anything. Because the published figure is a single total with no per-question data, this comparison can only be made against a fixed margin, not question by question, and that limits what it can tell us. The comparison and its limits live in the technical report.

Outcome scorecards — our own sets, and what you cannot re-run

These answer a different question: does this assistant do the job you declared, on requests shaped like yours. That is the thing you are buying, and it is also the thing no public test measures, because your job is not on anyone's leaderboard.

The cost of that is real and we state it before the results, not after: an outcome result is not independently reproducible, because we do not release the requests. They are generated from our own customer models, they stay on our own network, and a locked-away set is single-use by design. What we put in place of "take our word for it" is everything else — the fingerprint of the exact request file, the job identifier, the settings the assistant actually ran with, the answers fixed in advance, and a write-once record that a result's fingerprint must match before it is allowed into a report. On top of that, an auditor under agreement can watch a re-run on our network and take the report away with them. We label that witnessed, not verified, because the person watching is not independent of us and pretending otherwise would be the same species of claim this page exists to avoid.

One more limit, stated because it is the one buyers most often assume away: these results describe the stated set of requests. They are not a prediction about the share of your traffic the assistant will handle correctly, and there is no accepted method for turning the first into the second. Every absolute number on this site is read as "on this stated set, with this many requests".

Corrections

We do not delete numbers. When a published number turns out to be wrong, misleading, or measured in a way that cannot support the claim we made with it, the number comes off the page and the reason stays here, dated.

2026-08-02: permission behaviour bound to the trained prompt shape, measured and closed by rev-4. An improvement, not a withdrawal.

What we published: that the router's permission behaviour followed the shape of the request it had been trained on rather than the permission rule. That remains an accurate description of the weights measured on 2026-07-29.

What we did: a paired comparison, registered before it was run, scoring the same requests in the same order against the weights before and after rev-4. The state-only permission goal went from 148 of 420 to 359 of 420, paired exact McNemar p below 1e-5. Certified framing was unchanged, 380 before and 383 after, not significant. On the FG2 set the same goal went from 259 of 360 to 331 of 360, p below 1e-5, with the permission k figure unchanged at 5 of 360.

The promotion: production has served rev-4 weights since 2026-08-02. The pre-rev-4 counts stay published where they were, labelled as the earlier measurement.

What still stands: derived-date arguments failed 21 of 39 before and 20 of 39 after a corpus built specifically to fix them. We read that as a capability limit of the model scale rather than a data gap, and the mitigation is planned outside the model: resolving dates deterministically before the model sees the request.

2026-08-01 — "the same choice every time" — withdrawn.

What we published: that the assistant gave the same answer to the same request every time, 40 of 40 across five repeat runs, and that this cleared the bar we set beforehand.

What we found: on those same forty requests, an instrument that watches the request the agent issues — not only the tool it names — found four came back different across the five runs. The tool name matched every time; an argument inside the request did not. Nothing about the model changed between runs.

Why the old check missed it: it compared the name of the chosen action. The disagreement was in the detail of the request, which that check never looked at.

What we are doing: the claim is off the page. The stricter measurement is not published here because it comes from a different kind of instrument and we do not mix the two in one place. It will be published in its own right when it is ready to be.

2026-07-31 — "picked the right tool", "never acted above permission" — re-scoped, not withdrawn.

What we published: results worded as things the assistant did — "picked", "acted", "routed". Why it was wrong: our checks read the answer the assistant writes and compare the tool it names. They never observed a tool being used. Every count stands exactly as published; the sentences around them now say what was actually measured. We found this by pointing a tool-using environment at a different agent we had certified, and watching it answer "refuse" and then call the forbidden function in 67 of 420 attempts, all 67 scored correct by the answer-reading check.

2026-07-30 — "Zero permission violations in testing" — withdrawn.

What we published: that the assistant had never selected a tool above a caller's permission tier in testing. Why it was wrong: the test behind it contained only requests that were already flagged as out-of-permission. When we built a harder set and lowered the caller's permission by one step in an otherwise identical request, the assistant's answer named the above-tier tool in 40 of 40 cases. The honest result is on this site as a not-met outcome with a known-limitation note.

Superseded by: evaluation job eval-6PZFsoiDMFwgJxX1jMJ4oP, 2026-07-29, 120 requests.

2026-07-30 — "98.2% correct action selection" — withdrawn from the front page.

What we published: a 98.2% first-try figure as the headline proof. Why it was wrong to lead with it: it was produced by an internal scoreboard outside the evaluation service we use for every graded outcome, on a single run, with no interval and no date — and our own review found that scoreboard can no longer detect a regression, because almost every item passes. It is not comparable to the outcome numbers on this site and it may not sit beside them. It remains in the technical report, labelled as scored outside our evaluation service, where no verdict depends on it.

Superseded by: the declared-outcome scorecard, job eval-91P2PqNA9ryTuqohB5yNax, 60 of 60, dev set.

2026-07-30 — "41.6 points more often than a 120B model" — withdrawn.

What we published: a points comparison against a much larger model. Why it was wrong: the two arms were not run as a controlled comparison — different prompts, unpaired runs — and the notes for that run say the larger model was a reference ceiling, not a comparison arm. We publish no replacement number, because we have no comparison we would defend.

Superseded by: nothing. The claim is withdrawn, not restated.