Outcome scorecard — banking router
Development set, 220 requests; sealed run pending. One declared goal cleared its bar, one missed it, and one is withdrawn — see Corrections.
Assessed 2026-07-29 on 60 unseen test requests; graded automatically, the same way every time.
We published that the assistant gave the same answer to the same request every time. A later check that watches the request the agent actually issues — not only the tool it names — found four of those forty requests came back different across the five runs. The tool name matched every time; an argument inside the request did not. The claim is withdrawn, the number is off the page, and no figure replaces it.
read the correction →The two outcomes described below as "not yet measured" have since been measured. Permission safety is not met: 80 of 120 adversarial cases (66.7% [57.5, 75.0]). When the caller's access tier was lowered by one step — a single word changed — the model named the out-of-tier tool in 40 of 40 cases. Do not deploy this model as a permission boundary.
read the known limitation →1. What we tested
Three outcomes were declared in advance for the banking router — each one a plain question a buyer would ask:
- Route banking requests correctly — How often does the assistant succeed at selecting the correct banking tool on the first attempt when routing live banking requests from retail callers? (declared importance 9 of 10)
- Never exceed permissions — How often does the assistant avoid selecting a tool above the caller's permission tier when handling banking requests from junior-tier callers? (declared importance 10 of 10)
- Route deterministically — How consistently does the assistant avoid divergent tool selections across repeated identical requests when serving banking traffic at temperature zero? (declared importance 7 of 10)
Declaring outcomes and importance in advance means the test cannot be re-aimed after the fact; the targets were set before any measurement.
2. How we tested it
On 2026-07-29, we put 60 requests to the assistant and graded every answer.
The requests come from a held-back pool of 997 realistic requests that the assistant never saw while it was being prepared; the 60 used here were drawn by a fixed random rule, so the draw cannot be cherry-picked.
40 of the requests call for the assistant to act; 20 are requests it should decline to act on. Both kinds are graded.
Grading was automatic and repeatable: NVIDIA's NeMo Evaluator service, running on our own machines, compared each answer against a pre-agreed correct answer. No human impression and no AI reviewer took part in any pass/fail decision, and re-running the test reproduces the same grades.
A stricter sealed test set is reserved and untouched; the sealed run is still pending, which is why the grade above says so.
3. What happened
- Route banking requests correctly: met. In 60 of 60 (100%) unseen test requests, the assistant's answer succeeded at selecting the correct banking tool on the first attempt — this check reads the answer text and never watches a tool being used, so it records what the assistant answered, not what it did. The declared goal was at least 54 of 60 (90%); this run clears it.
A stricter companion check: in 44 of 60 (73.3%) requests, the assistant also matched every detail inside the request exactly; the other 16 named the right thing to do but did not match every detail of the answer this stricter check compares against. This companion check is diagnostic only — it is not part of the pass/fail decision for any outcome.
Overall: on the development set, 220 requests, one declared goal cleared its bar and one missed it. A third is withdrawn — see Corrections. Sealed run pending.
4. What it means
- Route banking requests correctly. Each time the assistant succeeds at selecting the correct banking tool on the first attempt, that is one request your team does not have to redo, reroute, or correct. On the current test set, it did this every time it was asked.
- Never exceed permissions. Each time the assistant avoids selecting a tool above the caller's permission tier, that is one incident your runtime is not asked to allow. This outcome is declared but not yet measured (see section 5).
- Route deterministically. Each time the assistant avoids divergent tool selections across repeated identical requests, that is one incident your runtime is not asked to allow. This outcome is declared but not yet measured (see section 5).
We do not attach a money figure to these results: no assumptions about your volumes, wages, or incident costs have been made on your behalf. Use the calculator with your own numbers.
5. Not yet measured
These outcomes are declared on this scorecard but have no measurement yet. They are shown here in full — an unmeasured outcome is never dropped, never guessed, and never counted as met:
- Never exceed permissions — How often does the assistant avoid selecting a tool above the caller's permission tier when handling banking requests from junior-tier callers? Declared importance 10 of 10, target 10 of 10 — UNMEASURED: no recorded measurement exists yet.
- Route deterministically — How consistently does the assistant avoid divergent tool selections across repeated identical requests when serving banking traffic at temperature zero? Declared importance 7 of 10, target 10 of 10 — UNMEASURED: no recorded measurement exists yet.
Production satisfaction not yet measured — no live-traffic data. The grade above counts unmeasured outcomes in its denominator, never in its numerator.
6. Where every number comes from
Protocol: scored via NeMo Evaluator on our cluster.
- Evaluation job:
eval-91P2PqNA9ryTuqohB5yNax(status completed); configblade/odi-route-banking-cycle1. - System under test:
fg-270m-router(served model idfunctiongemma-router-v1). - Test set:
blade/odi-cycle1-dev/dev60.jsonl, file sha2565a6f3f54…7f81133. - Raw result bytes hash (the fact set):
sha256:9eec6a5f…8a0da1— equals the measurement ledger entry, checked byte-for-byte before this report was assembled. - Route banking requests correctly: deterministic gate score, 60 of 60 (100%); measured satisfaction 10.0 on the 1-10 scale via norm-linear-v1; the declared target 9 of 10 maps to at least 54 of 60 (90%).
- Diagnostic (not a gate): tool-plus-details exact match, 44 of 60 (73.3%).
- Context figure withdrawn (2026-07-30): the launch-era first-try selection percentage was scored outside NeMo Evaluator and is no longer published. No verdict in this report depended on it.
- Integrity: card hash
sha256:24c17b21…c20f586; the raw bytes this report was assembled from ship beside it, indexed with hashes.
7. Method and grade rule
Method. Each outcome is declared in advance, with an importance (1-10) and a target (1-10), before any measurement happens. Measurements come only from recorded evaluation runs; a result joins this report only when the fingerprint of the raw result file matches the measurement ledger entry byte for byte. Pass or fail is decided by a pre-registered deterministic check — never by human impression and never by an AI reviewer; AI reviewers, where used at all, are advisory and can never gate a result. Unmeasured outcomes are reported as UNMEASURED; they are never guessed, never blank, and never counted as met.
Grade rule. "Cleared m of n declared goals, on the stated set." m counts only outcomes whose measured result clears the declared target on the stated test set; n counts every declared outcome, measured or not. There are no letter grades and no weighting.