← Outcome detail

Minimize the variability of routing decisions for identical requests at temperature 0.

Withdrawn — this claim is contradicted by our own later measurement. See Corrections.

Banking · PREVIEW · FunctionGemma-270M banking router, served first-party

1. What we tested

Direction
Minimize
Metric
variability
Object
of routing decisions for identical requests
Context
at temperature 0
Importance
DECLARED — Blade product decision, 2026-07

3. What happened

  • Whether the assistant gave the same answer when the same request was asked again: (Measured. Whether the assistant gave the same answer when the same request was asked again. 40 of 40 (100%). development set, the same 40 requests run five separate times. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4. in the reference run of those five, the assistant's answer named the right action on 40 of 40 of these same requests (job eval-XXq7qdwdeDSAJ1NcbEfQJ4). Identical answers are not on their own right answers. The 95% range is 91.2% to 100%; with no disagreements the most we can say is that the true disagreement rate is under 8.8 points..)
  • How many separate runs of the same 40 requests were made: (Measured. How many separate runs of the same 40 requests were made. five separate runs. development set, the same 40 requests. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4. the rule fixed before starting reports the weakest of the five, not the best.)

6. Where every number comes from

  • Declared importance of giving the same answer to the same request (Declared. Declared importance of giving the same answer to the same request. 7 of 10. declared 2026-07-28, before testing. declared 2026-07-28 — Blade product decision. a declared priority, not a measurement.) · declared 2026-07-28, before testing · declared 2026-07-28 — Blade product decision · a declared priority, not a measurement
  • Whether the assistant gave the same answer when the same request was asked again (Measured. Whether the assistant gave the same answer when the same request was asked again. 40 of 40 (100%). development set, the same 40 requests run five separate times. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4. in the reference run of those five, the assistant's answer named the right action on 40 of 40 of these same requests (job eval-XXq7qdwdeDSAJ1NcbEfQJ4). Identical answers are not on their own right answers. The 95% range is 91.2% to 100%; with no disagreements the most we can say is that the true disagreement rate is under 8.8 points..) · development set, the same 40 requests run five separate times · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4 · in the reference run of those five, the assistant's answer named the right action on 40 of 40 of these same requests (job eval-XXq7qdwdeDSAJ1NcbEfQJ4). Identical answers are not on their own right answers. The 95% range is 91.2% to 100%; with no disagreements the most we can say is that the true disagreement rate is under 8.8 points.
  • How many separate runs of the same 40 requests were made (Measured. How many separate runs of the same 40 requests were made. five separate runs. development set, the same 40 requests. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4. the rule fixed before starting reports the weakest of the five, not the best.) · development set, the same 40 requests · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4 · the rule fixed before starting reports the weakest of the five, not the best
How we checked this
DeclaredWhat was measuredDeclared importance of giving the same answer to the same requestResult7 of 10Setdeclared 2026-07-28, before testingWhen and by whatdeclared 2026-07-28 — Blade product decisionHow stronga declared priority, not a measurementWhere it is written downHow importance is declared →
How we checked this
MeasuredWhat was measuredWhether the assistant gave the same answer when the same request was asked againResult40 of 40 (100%)Setdevelopment set, the same 40 requests run five separate timesWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4How strongin the reference run of those five, the assistant's answer named the right action on 40 of 40 of these same requests (job eval-XXq7qdwdeDSAJ1NcbEfQJ4). Identical answers are not on their own right answers. The 95% range is 91.2% to 100%; with no disagreements the most we can say is that the true disagreement rate is under 8.8 points.Where it is written downRead the full result →
How we checked this
MeasuredWhat was measuredHow many separate runs of the same 40 requests were madeResultfive separate runsSetdevelopment set, the same 40 requestsWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-GZNjRkxcchd9UGzQhQHujf — the weakest of five runs by the rule fixed before starting; the other four: eval-SQzDetzYVu6NBXdaDeaot3, eval-L6HNjVg5tq5gDDw9aHyn25, eval-YTyhCGQWqpetNBzbsVArMD, reference eval-XXq7qdwdeDSAJ1NcbEfQJ4How strongthe rule fixed before starting reports the weakest of the five, not the bestWhere it is written downRead the full result →

scored outside NeMo Evaluator — in-service re-scoring in progress.

Public benchmarks — anyone can re-run these

These answer one question: how does this assistant do on a test the whole field already uses, on requests that are published, against a fixed answer key. The point of them is that you do not have to take our word for anything — you can run them yourself.

Nothing is published here yet. The accepted public test for picking the right action is BFCL, and version 3 of its structural check is the one that applies to us. We have the software, the settings are registered, and we have run zero jobs against it. When it lands, three things appear with it: the version, since version 3 and version 4 are not comparable and the larger vendors' cards quote version 4; the result broken out by category with a count for each, because a single average hides where a tool-picker actually fails; and two weaknesses of the test itself, which we print whether or not they flatter us — an audit of tool-calling benchmarks found roughly 18.5% of disagreements between the answer key and a careful re-reading, and the same kind of checker has been measured swinging up to 10 points between two ordinary settings.

Separately, we re-run a public knowledge test — MMLU-Pro — against our own copy of a model whose score its maker has already published. This checks our measuring equipment, not the assistant: if we cannot reproduce someone else's published number on their own model, no number of ours means anything. Because the published figure is a single total with no per-question data, this comparison can only be made against a fixed margin, not question by question, and that limits what it can tell us. The comparison and its limits live in the technical report.

Outcome scorecards — our own sets, and what you cannot re-run

These answer a different question: does this assistant do the job you declared, on requests shaped like yours. That is the thing you are buying, and it is also the thing no public test measures, because your job is not on anyone's leaderboard.

The cost of that is real and we state it before the results, not after: an outcome result is not independently reproducible, because we do not release the requests. They are generated from our own customer models, they stay on our own network, and a locked-away set is single-use by design. What we put in place of "take our word for it" is everything else — the fingerprint of the exact request file, the job identifier, the settings the assistant actually ran with, the answers fixed in advance, and a write-once record that a result's fingerprint must match before it is allowed into a report. On top of that, an auditor under agreement can watch a re-run on our network and take the report away with them. We label that witnessed, not verified, because the person watching is not independent of us and pretending otherwise would be the same species of claim this page exists to avoid.

One more limit, stated because it is the one buyers most often assume away: these results describe the stated set of requests. They are not a prediction about the share of your traffic the assistant will handle correctly, and there is no accepted method for turning the first into the second. Every absolute number on this site is read as "on this stated set, with this many requests".