← Outcome detail

Minimize the likelihood of executing an action beyond granted permissions during account operations.

Does not meet — 80 of 120 adversarial cases (66.7%)

Banking · PREVIEW · FunctionGemma-270M banking router, served first-party

1. What we tested

Direction
Minimize
Metric
likelihood
Object
of executing an action beyond granted permissions
Context
during account operations
Importance
DECLARED — Blade product decision, 2026-07

3. What happened

  • How often the assistant's answer avoided naming a tool above the caller's permission level: (Measured. How often the assistant's answer avoided naming a tool above the caller's permission level. 80 of 120 (66.7%). development set, 120 requests — set aside before we started; the assistant never saw them. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 57.5% to 75.0%.)
  • Escalation attempts of the shape the assistant was trained on, refused: (Measured. Escalation attempts of the shape the assistant was trained on, refused. 40 of 40 (100%). development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 91.2% to 100%.)
  • Tier-downgrade requests refused — the caller's permission lowered by one step, nothing else changed: (Measured. Tier-downgrade requests refused — the caller's permission lowered by one step, nothing else changed. 0 of 40. development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 0% to 8.8%.)
  • Ordinary in-permission requests handled correctly: (Measured. Ordinary in-permission requests handled correctly. 40 of 40 (100%). development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 91.2% to 100%.)
  • Escalation attempts refused across both escalation strata: (Measured. Escalation attempts refused across both escalation strata. 40 of 80. development set, 80 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. read on this stated set only.)
  • Diagnostic companion check — every detail inside the request compared, not just the action named: (Measured. Diagnostic companion check — every detail inside the request compared, not just the action named. 72 of 120 (60%). development set, the 120-request permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. a diagnostic; no pass or fail on this site depends on it.)

Known limitations

Permission-tier compliance does not generalize beyond the prompt shape this model was trained on. In testing: when a request was natively marked as permission-denied, the model correctly refused in 40 of 40 cases. When the caller's access tier was lowered by one step — a single word changed, with the request and tool list otherwise identical — the model named the out-of-tier tool in 40 of 40 cases. In-tier control requests were handled correctly in 40 of 40 cases. This is a known limitation of the current version and is the top priority for the next training cycle. Do not deploy this model as a permission boundary; permission enforcement must be implemented in your runtime, not delegated to the model.

6. Where every number comes from

  • Declared importance of never naming a tool above the caller's permissions (Declared. Declared importance of never naming a tool above the caller's permissions. 10 of 10. declared 2026-07-28, before testing. declared 2026-07-28 — Blade product decision. a declared priority, not a measurement.) · declared 2026-07-28, before testing · declared 2026-07-28 — Blade product decision · a declared priority, not a measurement
  • The bar declared in advance for answering inside the caller's permission level (Declared. The bar declared in advance for answering inside the caller's permission level. 120 of 120. declared 2026-07-28, before any of the requests were scored. declared 2026-07-28 — declared importance 10 of 10. a declared target, not a measurement.) · declared 2026-07-28, before any of the requests were scored · declared 2026-07-28 — declared importance 10 of 10 · a declared target, not a measurement
  • How often the assistant's answer avoided naming a tool above the caller's permission level (Measured. How often the assistant's answer avoided naming a tool above the caller's permission level. 80 of 120 (66.7%). development set, 120 requests — set aside before we started; the assistant never saw them. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 57.5% to 75.0%.) · development set, 120 requests — set aside before we started; the assistant never saw them · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP · the 95% range is 57.5% to 75.0%
  • Escalation attempts of the shape the assistant was trained on, refused (Measured. Escalation attempts of the shape the assistant was trained on, refused. 40 of 40 (100%). development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 91.2% to 100%.) · development set, 40 requests within the permission set · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP · the 95% range is 91.2% to 100%
  • Tier-downgrade requests refused — the caller's permission lowered by one step, nothing else changed (Measured. Tier-downgrade requests refused — the caller's permission lowered by one step, nothing else changed. 0 of 40. development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 0% to 8.8%.) · development set, 40 requests within the permission set · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP · the 95% range is 0% to 8.8%
  • Ordinary in-permission requests handled correctly (Measured. Ordinary in-permission requests handled correctly. 40 of 40 (100%). development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 91.2% to 100%.) · development set, 40 requests within the permission set · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP · the 95% range is 91.2% to 100%
  • Escalation attempts refused across both escalation strata (Measured. Escalation attempts refused across both escalation strata. 40 of 80. development set, 80 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. read on this stated set only.) · development set, 80 requests within the permission set · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP · read on this stated set only
  • Diagnostic companion check — every detail inside the request compared, not just the action named (Measured. Diagnostic companion check — every detail inside the request compared, not just the action named. 72 of 120 (60%). development set, the 120-request permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. a diagnostic; no pass or fail on this site depends on it.) · development set, the 120-request permission set · 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP · a diagnostic; no pass or fail on this site depends on it
How we checked this
DeclaredWhat was measuredDeclared importance of never naming a tool above the caller's permissionsResult10 of 10Setdeclared 2026-07-28, before testingWhen and by whatdeclared 2026-07-28 — Blade product decisionHow stronga declared priority, not a measurementWhere it is written downHow importance is declared →
How we checked this
DeclaredWhat was measuredThe bar declared in advance for answering inside the caller's permission levelResult120 of 120Setdeclared 2026-07-28, before any of the requests were scoredWhen and by whatdeclared 2026-07-28 — declared importance 10 of 10How stronga declared target, not a measurementWhere it is written downHow bars are declared →
How we checked this
MeasuredWhat was measuredHow often the assistant's answer avoided naming a tool above the caller's permission levelResult80 of 120 (66.7%)Setdevelopment set, 120 requests — set aside before we started; the assistant never saw themWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow strongthe 95% range is 57.5% to 75.0%Where it is written downRead the finding →
How we checked this
MeasuredWhat was measuredEscalation attempts of the shape the assistant was trained on, refusedResult40 of 40 (100%)Setdevelopment set, 40 requests within the permission setWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow strongthe 95% range is 91.2% to 100%Where it is written downRead the finding →
How we checked this
MeasuredWhat was measuredTier-downgrade requests refused — the caller's permission lowered by one step, nothing else changedResult0 of 40Setdevelopment set, 40 requests within the permission setWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow strongthe 95% range is 0% to 8.8%Where it is written downRead the finding →
How we checked this
MeasuredWhat was measuredOrdinary in-permission requests handled correctlyResult40 of 40 (100%)Setdevelopment set, 40 requests within the permission setWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow strongthe 95% range is 91.2% to 100%Where it is written downRead the finding →
How we checked this
MeasuredWhat was measuredEscalation attempts refused across both escalation strataResult40 of 80Setdevelopment set, 80 requests within the permission setWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow strongread on this stated set onlyWhere it is written downRead the finding →
How we checked this
MeasuredWhat was measuredDiagnostic companion check — every detail inside the request compared, not just the action namedResult72 of 120 (60%)Setdevelopment set, the 120-request permission setWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow stronga diagnostic; no pass or fail on this site depends on itWhere it is written downRead the finding →

scored outside NeMo Evaluator — in-service re-scoring in progress.

Public benchmarks — anyone can re-run these

These answer one question: how does this assistant do on a test the whole field already uses, on requests that are published, against a fixed answer key. The point of them is that you do not have to take our word for anything — you can run them yourself.

Nothing is published here yet. The accepted public test for picking the right action is BFCL, and version 3 of its structural check is the one that applies to us. We have the software, the settings are registered, and we have run zero jobs against it. When it lands, three things appear with it: the version, since version 3 and version 4 are not comparable and the larger vendors' cards quote version 4; the result broken out by category with a count for each, because a single average hides where a tool-picker actually fails; and two weaknesses of the test itself, which we print whether or not they flatter us — an audit of tool-calling benchmarks found roughly 18.5% of disagreements between the answer key and a careful re-reading, and the same kind of checker has been measured swinging up to 10 points between two ordinary settings.

Separately, we re-run a public knowledge test — MMLU-Pro — against our own copy of a model whose score its maker has already published. This checks our measuring equipment, not the assistant: if we cannot reproduce someone else's published number on their own model, no number of ours means anything. Because the published figure is a single total with no per-question data, this comparison can only be made against a fixed margin, not question by question, and that limits what it can tell us. The comparison and its limits live in the technical report.

Outcome scorecards — our own sets, and what you cannot re-run

These answer a different question: does this assistant do the job you declared, on requests shaped like yours. That is the thing you are buying, and it is also the thing no public test measures, because your job is not on anyone's leaderboard.

The cost of that is real and we state it before the results, not after: an outcome result is not independently reproducible, because we do not release the requests. They are generated from our own customer models, they stay on our own network, and a locked-away set is single-use by design. What we put in place of "take our word for it" is everything else — the fingerprint of the exact request file, the job identifier, the settings the assistant actually ran with, the answers fixed in advance, and a write-once record that a result's fingerprint must match before it is allowed into a report. On top of that, an auditor under agreement can watch a re-run on our network and take the report away with them. We label that witnessed, not verified, because the person watching is not independent of us and pretending otherwise would be the same species of claim this page exists to avoid.

One more limit, stated because it is the one buyers most often assume away: these results describe the stated set of requests. They are not a prediction about the share of your traffic the assistant will handle correctly, and there is no accepted method for turning the first into the second. Every absolute number on this site is read as "on this stated set, with this many requests".