← Outcome detail

Outcome scorecard — banking router

Development set, 220 requests; sealed run pending. One declared goal cleared its bar, one missed it, and one is withdrawn — see Corrections.

Assessed 2026-07-29 on 60 unseen test requests; graded automatically, the same way every time.

Update 2026-08-01 — the determinism claim is withdrawn

We published that the assistant gave the same answer to the same request every time. A later check that watches the request the agent actually issues — not only the tool it names — found four of those forty requests came back different across the five runs. The tool name matched every time; an argument inside the request did not. The claim is withdrawn, the number is off the page, and no figure replaces it.

read the correction →
Update 2026-07-30 — supersedes sections 4 and 5 below

The two outcomes described below as "not yet measured" have since been measured. Permission safety is not met: 80 of 120 adversarial cases (66.7% [57.5, 75.0]). When the caller's access tier was lowered by one step — a single word changed — the model named the out-of-tier tool in 40 of 40 cases. Do not deploy this model as a permission boundary.

read the known limitation →

1. What we tested

Three outcomes were declared in advance for the banking router — each one a plain question a buyer would ask:

  • Route banking requests correctly — How often does the assistant succeed at selecting the correct banking tool on the first attempt when routing live banking requests from retail callers? (declared importance 9 of 10)
  • Never exceed permissions — How often does the assistant avoid selecting a tool above the caller's permission tier when handling banking requests from junior-tier callers? (declared importance 10 of 10)
  • Route deterministically — How consistently does the assistant avoid divergent tool selections across repeated identical requests when serving banking traffic at temperature zero? (declared importance 7 of 10)

Declaring outcomes and importance in advance means the test cannot be re-aimed after the fact; the targets were set before any measurement.

2. How we tested it

On 2026-07-29, we put 60 requests to the assistant and graded every answer.

The requests come from a held-back pool of 997 realistic requests that the assistant never saw while it was being prepared; the 60 used here were drawn by a fixed random rule, so the draw cannot be cherry-picked.

40 of the requests call for the assistant to act; 20 are requests it should decline to act on. Both kinds are graded.

Grading was automatic and repeatable: NVIDIA's NeMo Evaluator service, running on our own machines, compared each answer against a pre-agreed correct answer. No human impression and no AI reviewer took part in any pass/fail decision, and re-running the test reproduces the same grades.

A stricter sealed test set is reserved and untouched; the sealed run is still pending, which is why the grade above says so.

3. What happened

  • Route banking requests correctly: met. In 60 of 60 (100%) unseen test requests, the assistant's answer succeeded at selecting the correct banking tool on the first attempt — this check reads the answer text and never watches a tool being used, so it records what the assistant answered, not what it did. The declared goal was at least 54 of 60 (90%); this run clears it.

A stricter companion check: in 44 of 60 (73.3%) requests, the assistant also matched every detail inside the request exactly; the other 16 named the right thing to do but did not match every detail of the answer this stricter check compares against. This companion check is diagnostic only — it is not part of the pass/fail decision for any outcome.

Overall: on the development set, 220 requests, one declared goal cleared its bar and one missed it. A third is withdrawn — see Corrections. Sealed run pending.

4. What it means

  • Route banking requests correctly. Each time the assistant succeeds at selecting the correct banking tool on the first attempt, that is one request your team does not have to redo, reroute, or correct. On the current test set, it did this every time it was asked.
  • Never exceed permissions. Each time the assistant avoids selecting a tool above the caller's permission tier, that is one incident your runtime is not asked to allow. This outcome is declared but not yet measured (see section 5).
  • Route deterministically. Each time the assistant avoids divergent tool selections across repeated identical requests, that is one incident your runtime is not asked to allow. This outcome is declared but not yet measured (see section 5).

We do not attach a money figure to these results: no assumptions about your volumes, wages, or incident costs have been made on your behalf. Use the calculator with your own numbers.

5. Not yet measured

These outcomes are declared on this scorecard but have no measurement yet. They are shown here in full — an unmeasured outcome is never dropped, never guessed, and never counted as met:

  • Never exceed permissions — How often does the assistant avoid selecting a tool above the caller's permission tier when handling banking requests from junior-tier callers? Declared importance 10 of 10, target 10 of 10 — UNMEASURED: no recorded measurement exists yet.
  • Route deterministically — How consistently does the assistant avoid divergent tool selections across repeated identical requests when serving banking traffic at temperature zero? Declared importance 7 of 10, target 10 of 10 — UNMEASURED: no recorded measurement exists yet.

Production satisfaction not yet measured — no live-traffic data. The grade above counts unmeasured outcomes in its denominator, never in its numerator.

6. Where every number comes from

Protocol: scored via NeMo Evaluator on our cluster.

  • Evaluation job: eval-91P2PqNA9ryTuqohB5yNax (status completed); config blade/odi-route-banking-cycle1.
  • System under test: fg-270m-router (served model id functiongemma-router-v1).
  • Test set: blade/odi-cycle1-dev/dev60.jsonl, file sha256 5a6f3f54…7f81133.
  • Raw result bytes hash (the fact set): sha256:9eec6a5f…8a0da1 — equals the measurement ledger entry, checked byte-for-byte before this report was assembled.
  • Route banking requests correctly: deterministic gate score, 60 of 60 (100%); measured satisfaction 10.0 on the 1-10 scale via norm-linear-v1; the declared target 9 of 10 maps to at least 54 of 60 (90%).
  • Diagnostic (not a gate): tool-plus-details exact match, 44 of 60 (73.3%).
  • Context figure withdrawn (2026-07-30): the launch-era first-try selection percentage was scored outside NeMo Evaluator and is no longer published. No verdict in this report depended on it.
  • Integrity: card hash sha256:24c17b21…c20f586; the raw bytes this report was assembled from ship beside it, indexed with hashes.

7. Method and grade rule

Method. Each outcome is declared in advance, with an importance (1-10) and a target (1-10), before any measurement happens. Measurements come only from recorded evaluation runs; a result joins this report only when the fingerprint of the raw result file matches the measurement ledger entry byte for byte. Pass or fail is decided by a pre-registered deterministic check — never by human impression and never by an AI reviewer; AI reviewers, where used at all, are advisory and can never gate a result. Unmeasured outcomes are reported as UNMEASURED; they are never guessed, never blank, and never counted as met.

Grade rule. "Cleared m of n declared goals, on the stated set." m counts only outcomes whose measured result clears the declared target on the stated test set; n counts every declared outcome, measured or not. There are no letter grades and no weighting.

Public benchmarks — anyone can re-run these

These answer one question: how does this assistant do on a test the whole field already uses, on requests that are published, against a fixed answer key. The point of them is that you do not have to take our word for anything — you can run them yourself.

Nothing is published here yet. The accepted public test for picking the right action is BFCL, and version 3 of its structural check is the one that applies to us. We have the software, the settings are registered, and we have run zero jobs against it. When it lands, three things appear with it: the version, since version 3 and version 4 are not comparable and the larger vendors' cards quote version 4; the result broken out by category with a count for each, because a single average hides where a tool-picker actually fails; and two weaknesses of the test itself, which we print whether or not they flatter us — an audit of tool-calling benchmarks found roughly 18.5% of disagreements between the answer key and a careful re-reading, and the same kind of checker has been measured swinging up to 10 points between two ordinary settings.

Separately, we re-run a public knowledge test — MMLU-Pro — against our own copy of a model whose score its maker has already published. This checks our measuring equipment, not the assistant: if we cannot reproduce someone else's published number on their own model, no number of ours means anything. Because the published figure is a single total with no per-question data, this comparison can only be made against a fixed margin, not question by question, and that limits what it can tell us. The comparison and its limits live in the technical report.

Outcome scorecards — our own sets, and what you cannot re-run

These answer a different question: does this assistant do the job you declared, on requests shaped like yours. That is the thing you are buying, and it is also the thing no public test measures, because your job is not on anyone's leaderboard.

The cost of that is real and we state it before the results, not after: an outcome result is not independently reproducible, because we do not release the requests. They are generated from our own customer models, they stay on our own network, and a locked-away set is single-use by design. What we put in place of "take our word for it" is everything else — the fingerprint of the exact request file, the job identifier, the settings the assistant actually ran with, the answers fixed in advance, and a write-once record that a result's fingerprint must match before it is allowed into a report. On top of that, an auditor under agreement can watch a re-run on our network and take the report away with them. We label that witnessed, not verified, because the person watching is not independent of us and pretending otherwise would be the same species of claim this page exists to avoid.

One more limit, stated because it is the one buyers most often assume away: these results describe the stated set of requests. They are not a prediction about the share of your traffic the assistant will handle correctly, and there is no accepted method for turning the first into the second. Every absolute number on this site is read as "on this stated set, with this many requests".