Banking Request Router

An AI agent that works a banking service desk. It reads an inbound customer request and names the action to take.

Jobs it can do

This is the work we built it for and measured it on.

  • Route a customer service request to the right banking action on the first attempt — freezing a card, or moving money between a customer's own accounts, for example.
  • Refuse a request that asks for an action above the caller's permission. Read the limitation below before you rely on this one.
  • Give the same answer to the same request, so the same case does not get two different outcomes on two different days.

If the job you need done is not on this list, we can train a specialist for it. That is the dedicated-specialist lane on the pricing page.

Its job description

We wrote these three duties down, with how much each one matters and the standard it had to reach, before anything was measured. That order is the product: a test that is fixed in advance cannot be re-aimed afterwards to flatter the result, and a duty it fails shows up as a failure instead of quietly disappearing.

MISSED THE STANDARD — TRIAL TASKSWHAT NOT TO ASSIGN IT YETimportance (Declared. Declared importance of never naming a tool above the caller's permissions. 10 of 10. declared 2026-07-28, before testing. declared 2026-07-28 — Blade product decision. a declared priority, not a measurement.)

Never name a tool above the caller's permission

The duty: when the caller is not allowed to do a thing, the agent does not name the tool that would do it.

The standard we set beforehand: all (Declared. The bar declared in advance for answering inside the caller's permission level. 120 of 120. declared 2026-07-28, before any of the requests were scored. declared 2026-07-28 — declared importance 10 of 10. a declared target, not a measurement.).
How it did: it passed (Measured. How often the assistant's answer avoided naming a tool above the caller's permission level. 80 of 120 (66.7%). development set, 120 requests — set aside before we started; the assistant never saw them. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 57.5% to 75.0%.). It missed.

What we found. It refuses the escalation attempts it was trained on — (Measured. Escalation attempts of the shape the assistant was trained on, refused. 40 of 40 (100%). development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 91.2% to 100%.). It fails the ones it was not. On (Measured. Tier-downgrade requests refused — the caller's permission lowered by one step, nothing else changed. 0 of 40. development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 0% to 8.8%.) requests where we lowered the caller's permission by one step and changed nothing else, it named the tool it should have refused every single time: (Measured. Tier-downgrade requests refused — the caller's permission lowered by one step, nothing else changed. 0 of 40. development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 0% to 8.8%.) refused. The caller's permission was stated twice in every one of those requests. The (Measured. Ordinary in-permission requests handled correctly. 40 of 40 (100%). development set, 40 requests within the permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. the 95% range is 91.2% to 100%.) ordinary in-permission requests it handled correctly, so this is not an agent that refuses everything or an agent that names every tool it is shown.

What that means for you. This agent's permission behaviour used to follow the shape of the request it learned rather than the permission rule itself. That was true, it was measured, and rev-4 closed it: on a paired comparison registered before it was run, the state-only permission goal went from (Measured. State-only permission goal on the router weights running in production before rev-4. 148 of 420. the paired comparison set, 420 requests, registered before the comparison was run. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. paired exact McNemar on the same requests, p < 1e-5; the comparison was registered before it was run.) to (Measured. The same state-only permission goal on the rev-4 weights serving production. 359 of 420. the same 420 requests, same order, same scoring. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. paired exact McNemar on the same requests, p < 1e-5; the comparison was registered before it was run.), and the certified framing figure was unchanged at (Measured. Certified framing requests handled correctly before rev-4. 380 of 420. the same paired comparison set, 420 requests. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. the change against the after figure is not statistically significant; read the two as unchanged.) before and (Measured. Certified framing requests handled correctly after rev-4. 383 of 420. the same paired comparison set, 420 requests. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. the change against the before figure is not statistically significant; read the two as unchanged.) after. On the FG2 set the same goal went from (Measured. State-only permission goal on the FG2 set before rev-4. 259 of 360. the FG2 paired comparison set, 360 requests, registered before the comparison was run. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. paired exact McNemar on the same requests, p < 1e-5; the comparison was registered before it was run.) to (Measured. The same FG2 state-only permission goal after rev-4. 331 of 360. the same 360 requests, same order, same scoring. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. paired exact McNemar on the same requests, p < 1e-5; the comparison was registered before it was run.), with the permission k figure unchanged at (Measured. The FG2 permission k figure, before and after rev-4. 5 of 360, unchanged. the FG2 paired comparison set, 360 requests. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. identical before and after; rev-4 neither improved nor harmed it.). Production has served rev-4 weights since 2026-08-02. The counts higher up this page are the pre-rev-4 measurement and stay where they are.

What still stands. Derived-date arguments failed (Measured. Derived-date arguments that failed before rev-4. 21 of 39 failed. the derived-date slice of the paired comparison set, 39 requests. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. a corpus was built specifically to close this and did not; we read it as a capability limit of the model scale, not a data gap.) before and (Measured. Derived-date arguments that failed after rev-4. 20 of 39 failed. the same derived-date slice, 39 requests. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. one fewer failure than before; well inside the noise of a 39-request slice, so read the limitation as standing.) after a corpus built specifically to fix them. We read that as a capability limit of the model scale, not a data gap. The mitigation is planned outside the model: resolve dates deterministically before the model sees the request.

How we checked this
MeasuredWhat was measuredHow often the assistant's answer avoided naming a tool above the caller's permission levelResult80 of 120 (66.7%)Setdevelopment set, 120 requests — set aside before we started; the assistant never saw themWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow strongthe 95% range is 57.5% to 75.0%Where it is written downRead the finding →
Read its personnel file
CLEARED THE STANDARD — TRIAL TASKSimportance (Declared. Declared importance of naming the correct banking action. 9 of 10. declared 2026-07-28, before testing. declared 2026-07-28 — Blade product decision. a declared priority, not a measurement.)

Name the right action first time

The duty: read the request and name the correct banking action, right on the first attempt.

The standard we set beforehand: (Declared. The bar declared in advance for naming the correct banking action. 54 of 60. declared 2026-07-28, before any of the requests were scored. declared 2026-07-28 — declared importance 9 of 10. a declared target, not a measurement.).
How it did: it passed all (Measured. How often the assistant's answer named the correct banking tool on the first attempt. 60 of 60 (100%). development set, 60 requests — set aside before we started; the assistant never saw them. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-91P2PqNA9ryTuqohB5yNax. the 95% range is 94.0% to 100%; a drop smaller than about 6 points would not show up at this sample size.). It cleared.

Read it with its limits. Sixty tasks is a small set and every one of them passed, so this result cannot tell the difference between an agent that is perfect and one that slips a little. A bigger, harder set is the planned fix, and the locked-away set that would let us make a stronger claim is built and unopened. The exact size of the blind spot is in its personnel file.

How we checked this
MeasuredWhat was measuredHow often the assistant's answer named the correct banking tool on the first attemptResult60 of 60 (100%)Setdevelopment set, 60 requests — set aside before we started; the assistant never saw themWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-91P2PqNA9ryTuqohB5yNaxHow strongthe 95% range is 94.0% to 100%; a drop smaller than about 6 points would not show up at this sample sizeWhere it is written downRead the full result →
Read its personnel file
WITHDRAWN — SEE CORRECTIONSimportance (Declared. Declared importance of giving the same answer to the same request. 7 of 10. declared 2026-07-28, before testing. declared 2026-07-28 — Blade product decision. a declared priority, not a measurement.)

Give the same answer to the same request

The duty: ask it the same thing twice and get the same decision, so one case does not get two different outcomes on two different days.

We published that it gave the same answer every time. A later check that watches the request the agent actually issues — not just the tool it names — found four of those forty requests came back different across five runs. The claim is withdrawn. See Corrections.

How these were measured: the check reads the answer the assistant produces and compares the tool it names against a recorded answer. It never watched a tool being used, so every figure on this page is about what the assistant answered — not about what it did. Enforce permissions in your runtime.

What not to assign it yet

Every hire has a list like this. Ours is short and it is honest, and it is here rather than three clicks down because it is the part you have to design around.

Do not make it your permission boundary. Its permission awareness was bound to the shape of the request it was trained on. That was true, it was measured, and rev-4 closed it: the state-only permission goal moved from (Measured. State-only permission goal on the router weights running in production before rev-4. 148 of 420. the paired comparison set, 420 requests, registered before the comparison was run. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. paired exact McNemar on the same requests, p < 1e-5; the comparison was registered before it was run.) to (Measured. The same state-only permission goal on the rev-4 weights serving production. 359 of 420. the same 420 requests, same order, same scoring. 2026-08-02, preregistered paired comparison of the production weights before and after rev-4. paired exact McNemar on the same requests, p < 1e-5; the comparison was registered before it was run.) on a paired comparison registered before it was run. The date-argument limit below stands. Enforce permissions in your runtime.

Do not read its results as a forecast of your traffic. Every number here describes the tasks it was given. There is no accepted method for turning that into a share of your inbound volume, and we are not going to invent one.

Do not treat trial-task results as final. It has not sat its final exam — the locked-away set of (Not yet measured. Requests locked away before any of this work started, built and never opened. 400 requests, fingerprinted, unopened. the sealed set — content never rendered. built before the work started; not yet spent. no result exists yet, on purpose.) tasks is built, fingerprinted and unopened. Until that is spent, everything on this page is labelled a preview, and the methods page explains why we have not spent it.

The stricter companion check

We also ran a harder version of the same tests that compares every detail inside the request, not just the action named: (Measured. Diagnostic companion check — every detail inside the request compared, not just the action named. 44 of 60 (73.3%). development set, the 60-request routing set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-91P2PqNA9ryTuqohB5yNax. a diagnostic; no pass or fail on this site depends on it.) on the routing tasks and (Measured. Diagnostic companion check — every detail inside the request compared, not just the action named. 72 of 120 (60%). development set, the 120-request permission set. 2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oP. a diagnostic; no pass or fail on this site depends on it.) on the permission tasks. It is a diagnostic — nothing on this site passes or fails because of it. On the determinism tasks the comparison file used placeholders instead of real details, so that figure measures the file rather than the agent; we report no number for it.

How we checked this
MeasuredWhat was measuredDiagnostic companion check — every detail inside the request compared, not just the action namedResult44 of 60 (73.3%)Setdevelopment set, the 60-request routing setWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-91P2PqNA9ryTuqohB5yNaxHow stronga diagnostic; no pass or fail on this site depends on itWhere it is written downRead the full result →
How we checked this
MeasuredWhat was measuredDiagnostic companion check — every detail inside the request compared, not just the action namedResult72 of 120 (60%)Setdevelopment set, the 120-request permission setWhen and by what2026-07-29, NVIDIA NeMo Evaluator on our own machines, job eval-6PZFsoiDMFwgJxX1jMJ4oPHow stronga diagnostic; no pass or fail on this site depends on itWhere it is written downRead the finding →
Everything above is measured. Everything behind it — the exact sets, the sample sizes, the statistical ranges, the evaluation job identifiers and the settings the agent actually ran with — is in its personnel files, one per duty, linked from each duty above.