01 · How to Read This Benchmark
Coverage Tiers — and What "World-Class" Means Here
World-class coverage is not scenario count. No credible buyer benchmarks an underwriting platform by counting demo scenarios; past ~18–20 exemplars a demo bloats without persuading anyone. World-class means four things: (1) a complete, examiner-shaped taxonomy of what "coverage" even is; (2) honest tiering of each cell against evidence; (3) computed numbers that update when the registry updates; and (4) coverage ultimately proven by an evaluation suite (Wave 3), not narrated by a demo. This document is that instrument.
Demo-Covered — interactive scenario exists (S01–S14)
Wave 1 — path scenarios S15–S18
Wave 2 — guardrail gauntlet + product adds
Eval Suite — golden-file cases, Wave 3 (audit P0-6)
Roadmap — dated, not yet scheduled into a wave
Out of Scope — excluded from denominators, rationale stated
Tier discipline mirrors the Capability Matrix's "Demo-Validated Design" standard from the claims-reconciliation pass: a Demo-Covered cell means the behavior is exercised in the interactive demo on synthetic data — production implementation is validated during the 90-day pilot. Eval-Covered (Wave 3) is the stronger claim: ≥3 golden-file cases per cell executed by the automated evaluation harness on every release. Out-of-scope items are excluded from every denominator so percentages cannot be inflated by scoping.
02 · Composite Scorecard
Coverage by Stage — Computed
Weighted composite across the three dimensions (weights shown below, declared in the registry). The progression is deliberately honest: the eval-stage target is below 100% because disparate-impact analytics remains a dated roadmap item (Phase 2 · Q4 2026) and two cells stay roadmap — a benchmark that reaches 100% before the product does is marketing, not measurement.
03 · Dimension A — Decision Paths
Decision-Path Coverage
The paths a credit decision can travel — the dimension buyers probe hardest, because it is where governance lives. The original fourteen scenarios all showed the system succeeding — a systemic gap. Wave 1 (shipped) closed it: S15–S18 show the system encountering trouble — abstaining, holding, escalating, and being overruled — which is what a CCO, an MRM officer, and an examiner actually ask to see.
04 · Dimension B — Loan Products
Product Coverage
Product coverage is claimed per-product, never as "every credit type you underwrite." Three products are declared out of scope with rationale and excluded from all denominators — stating a boundary is a strength in diligence; silence is the only wrong answer. Agricultural production credit shipped in Wave 2 (S20) — it was the highest-frequency community-bank objection.
05 · Dimension C — Regulatory Events
Regulatory-Event Coverage
The compliance moments a file must survive. Five of the missing events are deterministic policy checks, not scenarios — which is why Wave 2's centerpiece — now shipped — is a single Guardrail Gauntlet scenario (S19) that fires flood determination, legal lending limit, Reg O, HVCRE, and the SCRA screen in one run, wired to the deterministic Policy Guardrail Engine (P0-4, 9 citable rules, 44-test oracle suite) rather than scripted text.
06 · Progressive Build Plan
Four Waves to World-Class
Each wave has entry criteria, exit gates, and the audit P0 items it depends on. The sequencing rule: governance paths before products, products before proof, proof before production claims.
WAVE 1 ✓ SHIPPEDGovernance Path Scenarios — S15–S18
S15 Abstention / Degraded Mode (bureau feed fails mid-run → graceful degradation, agent abstains, partial packet routed to human with SLA) · S16 Verification Hold (4506-C transcript materially disagrees with CPA statements → hold + fraud-risk referral; showcases the Calculation Engine's double-count guard) · S17 Reg B Incompleteness → Withdrawal (the third Reg B path the demo currently skips) · S18 Officer Override with Telemetry (override + mandatory rationale, logged as MRM drift-monitoring data).
Entry Criteria
Calculation Engine v1.0 integrated (done — P0-2); claims reconciliation complete (done — P0-1); scenario evidence declared in CALC_EVIDENCE for S16.
Exit Gates
All four scenarios run full HITL choreography with hash-chained ledger entries; S15 demonstrates abstention (no fabricated values); S17 issues the 30-day incompleteness clock; S18 override telemetry appears in the audit export.
WAVE 2 ✓ SHIPPEDGuardrail Gauntlet + Product Adds — S19–S20
S19 Guardrail Gauntlet: one CRE construction credit that trips five deterministic regulatory checks in a single run — flood mandatory-purchase, legal lending limit, Reg O insider, HVCRE classification, SCRA/MLA — plus an SLA breach that actually fires and escalates. Built on the Policy Engine (audit P0-4), so each check is a versioned rule evaluated against Calculation Engine artifacts, not scripted text. S20 Agricultural Production Credit: FSA-guaranteed operating line with commodity-cycle stress — the most-named community-bank gap.
Entry Criteria
Policy Engine MVP (P0-4): versioned JSON rule set, deterministic evaluator, replay test against S01–S18 producing identical exception lists.
Exit Gates
All five S19 checks are rule-engine evaluations with rule IDs in the ledger; S20 stress uses engine formulas (price/yield haircuts); zero hand-typed ratios introduced.
WAVE 3 ✓ SHIPPEDEval-Proven Coverage — the Golden-File Suite
Coverage graduates from narrated to proven (audit P0-6): every non-roadmap registry cell gets ≥3 golden-file cases (input package → expected artifacts, exceptions, decision path, notices) executed by an automated harness on every release, with the pass rate published in this document. Demo scenarios remain the narrative layer — one instructive exemplar per archetype — while the eval suite carries the coverage claim.
Entry Criteria
Waves 1–2 shipped; Calculation Engine + Policy Engine artifacts stable; memo verifier (P0-7) available to score narrative outputs against artifacts.
Exit Gates
≥90% of eligible cells eval-covered at ≥3 cases; 100% pass rate as release gate; regression run required for any formula or rule version bump.
WAVE 4Production-Verified — Pilot Evidence
unlocks a tier this document cannot grant itself
The tier above Eval-Covered exists only after a 90-day institutional pilot: shadow-mode concordance on the partner's own book (≥90% decision alignment, ≥98% spread accuracy — the reconciled pilot gates), MRM validation sign-off on the 500-application replay, and fair-lending review. This document deliberately has no way to mark a cell "Production-Verified" — that column is added by pilot results, not by editing the registry.
Entry Criteria
Wave 3 exit gates; one certified LOS sandbox integration (P0-8); server-side evidence store & ledger (P0-5); security baseline (P0-10).
Exit Gates
Pilot exit criteria per the Pilot Structure document — the gates this suite's own documents already publish.
07 · Honest Limits
What This Benchmark Is Not
This taxonomy was curated by CAIBots from the structure of the OCC Comptroller's Handbook credit-risk booklets and FFIEC examination topics; it is a vendor-authored benchmark, not a certification, and no regulator has endorsed it. During a pilot, the partner institution should map this registry onto its own product set and exam findings and re-weight accordingly — the registry is in this document's source precisely so that is a five-minute edit, after which every number on this page recomputes. Percentages measure breadth of coverage, not depth or accuracy of any cell; depth is Wave 3's job. And a benchmark authored by the builder shares the limitation this platform's own SR 11-7 language now states about itself: independent validation is performed by the institution, not claimed by the vendor.
Priority Baseline · Gap Register
What Is Done, What Is Not — On the Record
The 360° product audit produced a P0–P3 register. This section tracks it in public, in the same spirit as the benchmark above: statuses are gate-verified where the work lives in this suite, and honestly marked EXTERNAL where they require a backend, a partner bank, or an auditor. This register is append-only and machine-readable; a release gate fails if the counts drift from the claims.
P0-1
Claims reconciliation
DONE
suite gate-enforced; data-room updates in founder queue
P0-2
Deterministic Calculation Engine
DONE
67 oracle tests · idempotence-gated
P0-3
Financial spreading MVP
EXTERNAL
build-vs-partner decision (Ocrolus/Inscribe)
P0-4
Policy Engine
DONE
44 tests · versioned pack · replay needs partner history
P0-5
Server-side evidence store + ledger
EXTERNAL
requires backend
P0-6
Eval harness + golden suite
SPLIT
harness + gating DONE (156 differential cases) · adjudicated corpus is the external labor item
P0-7
Critic v1
DONE
advisory-only · wired · gated
P0-8
Certified LOS sandbox
EXTERNAL
bank sandbox required
P0-9
Adverse-action engine
DONE
24 tests · 3 notice regimes · clocks
P0-10
SOC 2 + pen test + security posture
EXTERNAL
auditor engaged post-financing · browser-key disclaimer + tier decision DONE
P0 · 10 items — see 360° audit §34 for the full register
P1-12
Fair-lending DI methodology
PARTIAL
screens need backend · methodology doc queued
P1-13
MRM validation kit
PARTIAL
500-replay protocol published · kit documents queued
P1-16
Prompt-injection defense
DONE
v3.7: hardened preambles on all live-AI calls + schema rejection
P1-19
Two-clock latency honesty
DONE
published per scenario · SLOs need production
P1-21
Retention matrix
DONE
privacy draft + Impl. Guide §08
P1-23
Global cash flow incl. 1040/K-1
PARTIAL
v3.7: pgcf/pdscr/sbaGlobal formulas + phantom-income guard DONE · document ingestion external
P1-24
Honest full-cost ROI
DONE
total-cost headline · Gate 13
P1-11/14/15/17/18/20/22/25
Workflow · registry · KYC · doc-auth · IAM · reconciliation · SIEM · drill-down
EXTERNAL
backend / partner / production scope
P1 · 15 items — see 360° audit §34 for the full register
P2
Sensitivity strip
DONE
v3.7: [sens@1.0] decision distances, ledgered
P2
Examiner export pack
DONE
v3.7: one-click exam binder — artifacts + AAN + hash-chained ledger + policy pin
P2
Remaining 18
ROADMAP
policy simulation · appraisal review · borrowing-base automation · exam-binder-at-scale · VPC · et al.
P2 · 20 items — see 360° audit §34 for the full register
P3
All 20
ROADMAP
future differentiation — see 360° audit §34
P3 · 20 items — see 360° audit §34 for the full register
Release Log · Append-Only
2026-07-27
v3.1–v3.3
Wave 1–3: governance scenarios · PolicyEngine · eval harness · 92.7% registry
2026-07-27
v3.4
AAN engine + Critic v1 · explicit agent recommendations
2026-07-28
v3.4.x
Sign-off fixes · ROI deep-fix (Gate 13) · satellite alignment
2026-07-28
v3.5.x
Visual-system unification (Gate 14) · embedded-PHP watermark removal
2026-07-28
v3.6.x
Byte-identical universal footer · defensible ROI headline (total-cost basis)
2026-07-28
v3.7
pgcf/pdscr/sbaGlobal formulas · sensitivity strip · examiner binder · injection hardening (Gate 15)
2026-07-28
v3.7.1–5
Gap register published · consistency passes · idle-state polish · in-browser Binder Verifier + projector hardening