Skip to main content
CAIBots
CAIBots
Credit Underwriting AI · V4.0
Demo Environment · Synthetic Data
Registry-Computed Document
Coverage Benchmark · Progressive Build Plan

Underwriting Coverage,
Measured — Not Claimed

Three-dimension benchmark of the CAIBots Credit Underwriting platform — decision paths × loan products × regulatory events — mapped against the coverage a bank credit function is examined on. Every percentage on this page is computed at load time from the registry in this document's source, in the same discipline as the platform's Calculation Engine: numbers are outputs of declared inputs, never assertions.

Coverage · two numbers, one vocabulary
71.8%
Demonstrated
Decision paths, products and regulatory events exercised by scenarios S01–S20. Computed from the declared registry.
92.7%
Proven
156-case harness, 156 passed, 51 of 55 eligible items proven against a Python/Decimal differential oracle — an independent reimplementation of every formula and rule.
Why two numbers. Demonstrated is what runs in the demo today. Proven is what an evaluation harness has verified against an independently reimplemented oracle — a higher bar than a demo lane, because it tests the logic rather than the rendering. Across the CAIBots suite only Credit Underwriting currently publishes a Proven figure. The other three show an em-dash, which is the honest answer until a harness exists. Both numbers use the same vocabulary so the four estates can be compared without translating between them.

Where this denominator comes from

Coverage is only meaningful if the denominator is. This one is a split: one dimension is anchored to published regulation and two are not, and the difference is stated rather than blurred.

DimensionItemsAnchor
Regulatory events16 Externally anchored. Every item is a named CFR part or statute — Reg B 12 CFR 1002, HMDA 12 CFR 1003, CRA 12 CFR 25, FCRA 15 U.S.C. 1681, Flood 42 U.S.C. 4012a, Reg O 12 CFR 215, SCRA 50 U.S.C. 3937, FIRREA appraisal 12 CFR 34, model risk SR 11-7. Each row cites its own basis.
Decision paths21 Self-declared. No regulator publishes a canonical list of credit decision paths. These are the routes an origination system must be able to take — auto-approval through committee — enumerated by us.
Product types21 Self-declared. Any two institutions would draw this list differently. Ours spans consumer installment to syndicated Term B.

Selection criterion for the self-declared dimensions. An item is in the denominator if a credit decision system could reasonably be expected to handle it and its absence would be a gap a lender would notice. Items excluded by design are listed in the Out of Scope table below with a reason each — stating what is not counted is what makes the rest defensible.

What this is not. It is not an industry standard, and a different institution would produce a different denominator. The regulatory dimension can be checked against the CFR; the other two can only be checked against the criterion above. That is the honest position, and it is better said here than discovered by a reviewer.

Registry-Computed3 Dimensions4 Build WavesCurated vs. OCC Handbook / FFIEC Exam TaxonomyCompanion to Capability Matrix
01 · How to Read This Benchmark

Coverage Tiers — and What "World-Class" Means Here

World-class coverage is not scenario count. No credible buyer benchmarks an underwriting platform by counting demo scenarios; past ~18–20 exemplars a demo bloats without persuading anyone. World-class means four things: (1) a complete, examiner-shaped taxonomy of what "coverage" even is; (2) honest tiering of each cell against evidence; (3) computed numbers that update when the registry updates; and (4) coverage ultimately proven by an evaluation suite (Wave 3), not narrated by a demo. This document is that instrument.

Demo-Covered — interactive scenario exists (S01–S14) Wave 1 — path scenarios S15–S18 Wave 2 — guardrail gauntlet + product adds Eval Suite — golden-file cases, Wave 3 (audit P0-6) Roadmap — dated, not yet scheduled into a wave Out of Scope — excluded from denominators, rationale stated

Tier discipline mirrors the Capability Matrix's "Demo-Validated Design" standard from the claims-reconciliation pass: a Demo-Covered cell means the behavior is exercised in the interactive demo on synthetic data — production implementation is validated during the 90-day pilot. Eval-Covered (Wave 3) is the stronger claim: ≥3 golden-file cases per cell executed by the automated evaluation harness on every release. Out-of-scope items are excluded from every denominator so percentages cannot be inflated by scoping.

02 · Composite Scorecard

Coverage by Stage — Computed

Weighted composite across the three dimensions (weights shown below, declared in the registry). The progression is deliberately honest: the eval-stage target is below 100% because disparate-impact analytics remains a dated roadmap item (Phase 2 · Q4 2026) and two cells stay roadmap — a benchmark that reaches 100% before the product does is marketing, not measurement.

03 · Dimension A — Decision Paths

Decision-Path Coverage

The paths a credit decision can travel — the dimension buyers probe hardest, because it is where governance lives. The original fourteen scenarios all showed the system succeeding — a systemic gap. Wave 1 (shipped) closed it: S15–S18 show the system encountering trouble — abstaining, holding, escalating, and being overruled — which is what a CCO, an MRM officer, and an examiner actually ask to see.

04 · Dimension B — Loan Products

Product Coverage

Product coverage is claimed per-product, never as "every credit type you underwrite." Three products are declared out of scope with rationale and excluded from all denominators — stating a boundary is a strength in diligence; silence is the only wrong answer. Agricultural production credit shipped in Wave 2 (S20) — it was the highest-frequency community-bank objection.

05 · Dimension C — Regulatory Events

Regulatory-Event Coverage

The compliance moments a file must survive. Five of the missing events are deterministic policy checks, not scenarios — which is why Wave 2's centerpiece — now shipped — is a single Guardrail Gauntlet scenario (S19) that fires flood determination, legal lending limit, Reg O, HVCRE, and the SCRA screen in one run, wired to the deterministic Policy Guardrail Engine (P0-4, 9 citable rules, 44-test oracle suite) rather than scripted text.

06 · Progressive Build Plan

Four Waves to World-Class

Each wave has entry criteria, exit gates, and the audit P0 items it depends on. The sequencing rule: governance paths before products, products before proof, proof before production claims.

WAVE 1 ✓ SHIPPEDGovernance Path Scenarios — S15–S18

S15 Abstention / Degraded Mode (bureau feed fails mid-run → graceful degradation, agent abstains, partial packet routed to human with SLA) · S16 Verification Hold (4506-C transcript materially disagrees with CPA statements → hold + fraud-risk referral; showcases the Calculation Engine's double-count guard) · S17 Reg B Incompleteness → Withdrawal (the third Reg B path the demo currently skips) · S18 Officer Override with Telemetry (override + mandatory rationale, logged as MRM drift-monitoring data).

Entry Criteria
Calculation Engine v1.0 integrated (done — P0-2); claims reconciliation complete (done — P0-1); scenario evidence declared in CALC_EVIDENCE for S16.
Exit Gates
All four scenarios run full HITL choreography with hash-chained ledger entries; S15 demonstrates abstention (no fabricated values); S17 issues the 30-day incompleteness clock; S18 override telemetry appears in the audit export.
WAVE 2 ✓ SHIPPEDGuardrail Gauntlet + Product Adds — S19–S20

S19 Guardrail Gauntlet: one CRE construction credit that trips five deterministic regulatory checks in a single run — flood mandatory-purchase, legal lending limit, Reg O insider, HVCRE classification, SCRA/MLA — plus an SLA breach that actually fires and escalates. Built on the Policy Engine (audit P0-4), so each check is a versioned rule evaluated against Calculation Engine artifacts, not scripted text. S20 Agricultural Production Credit: FSA-guaranteed operating line with commodity-cycle stress — the most-named community-bank gap.

Entry Criteria
Policy Engine MVP (P0-4): versioned JSON rule set, deterministic evaluator, replay test against S01–S18 producing identical exception lists.
Exit Gates
All five S19 checks are rule-engine evaluations with rule IDs in the ledger; S20 stress uses engine formulas (price/yield haircuts); zero hand-typed ratios introduced.
WAVE 3 ✓ SHIPPEDEval-Proven Coverage — the Golden-File Suite

Coverage graduates from narrated to proven (audit P0-6): every non-roadmap registry cell gets ≥3 golden-file cases (input package → expected artifacts, exceptions, decision path, notices) executed by an automated harness on every release, with the pass rate published in this document. Demo scenarios remain the narrative layer — one instructive exemplar per archetype — while the eval suite carries the coverage claim.

Entry Criteria
Waves 1–2 shipped; Calculation Engine + Policy Engine artifacts stable; memo verifier (P0-7) available to score narrative outputs against artifacts.
Exit Gates
≥90% of eligible cells eval-covered at ≥3 cases; 100% pass rate as release gate; regression run required for any formula or rule version bump.
WAVE 4Production-Verified — Pilot Evidence
unlocks a tier this document cannot grant itself

The tier above Eval-Covered exists only after a 90-day institutional pilot: shadow-mode concordance on the partner's own book (≥90% decision alignment, ≥98% spread accuracy — the reconciled pilot gates), MRM validation sign-off on the 500-application replay, and fair-lending review. This document deliberately has no way to mark a cell "Production-Verified" — that column is added by pilot results, not by editing the registry.

Entry Criteria
Wave 3 exit gates; one certified LOS sandbox integration (P0-8); server-side evidence store & ledger (P0-5); security baseline (P0-10).
Exit Gates
Pilot exit criteria per the Pilot Structure document — the gates this suite's own documents already publish.
07 · Honest Limits

What This Benchmark Is Not

This taxonomy was curated by CAIBots from the structure of the OCC Comptroller's Handbook credit-risk booklets and FFIEC examination topics; it is a vendor-authored benchmark, not a certification, and no regulator has endorsed it. During a pilot, the partner institution should map this registry onto its own product set and exam findings and re-weight accordingly — the registry is in this document's source precisely so that is a five-minute edit, after which every number on this page recomputes. Percentages measure breadth of coverage, not depth or accuracy of any cell; depth is Wave 3's job. And a benchmark authored by the builder shares the limitation this platform's own SR 11-7 language now states about itself: independent validation is performed by the institution, not claimed by the vendor.

Priority Baseline · Gap Register

What Is Done, What Is Not — On the Record

The 360° product audit produced a P0–P3 register. This section tracks it in public, in the same spirit as the benchmark above: statuses are gate-verified where the work lives in this suite, and honestly marked EXTERNAL where they require a backend, a partner bank, or an auditor. This register is append-only and machine-readable; a release gate fails if the counts drift from the claims.

P0-1
Claims reconciliation
DONE
suite gate-enforced; data-room updates in founder queue
P0-2
Deterministic Calculation Engine
DONE
67 oracle tests · idempotence-gated
P0-3
Financial spreading MVP
EXTERNAL
build-vs-partner decision (Ocrolus/Inscribe)
P0-4
Policy Engine
DONE
44 tests · versioned pack · replay needs partner history
P0-5
Server-side evidence store + ledger
EXTERNAL
requires backend
P0-6
Eval harness + golden suite
SPLIT
harness + gating DONE (156 differential cases) · adjudicated corpus is the external labor item
P0-7
Critic v1
DONE
advisory-only · wired · gated
P0-8
Certified LOS sandbox
EXTERNAL
bank sandbox required
P0-9
Adverse-action engine
DONE
24 tests · 3 notice regimes · clocks
P0-10
SOC 2 + pen test + security posture
EXTERNAL
auditor engaged post-financing · browser-key disclaimer + tier decision DONE
P0 · 10 items — see 360° audit §34 for the full register
P1-12
Fair-lending DI methodology
PARTIAL
screens need backend · methodology doc queued
P1-13
MRM validation kit
PARTIAL
500-replay protocol published · kit documents queued
P1-16
Prompt-injection defense
DONE
v3.7: hardened preambles on all live-AI calls + schema rejection
P1-19
Two-clock latency honesty
DONE
published per scenario · SLOs need production
P1-21
Retention matrix
DONE
privacy draft + Impl. Guide §08
P1-23
Global cash flow incl. 1040/K-1
PARTIAL
v3.7: pgcf/pdscr/sbaGlobal formulas + phantom-income guard DONE · document ingestion external
P1-24
Honest full-cost ROI
DONE
total-cost headline · Gate 13
P1-11/14/15/17/18/20/22/25
Workflow · registry · KYC · doc-auth · IAM · reconciliation · SIEM · drill-down
EXTERNAL
backend / partner / production scope
P1 · 15 items — see 360° audit §34 for the full register
P2
Sensitivity strip
DONE
v3.7: [sens@1.0] decision distances, ledgered
P2
Examiner export pack
DONE
v3.7: one-click exam binder — artifacts + AAN + hash-chained ledger + policy pin
P2
Remaining 18
ROADMAP
policy simulation · appraisal review · borrowing-base automation · exam-binder-at-scale · VPC · et al.
P2 · 20 items — see 360° audit §34 for the full register
P3
All 20
ROADMAP
future differentiation — see 360° audit §34
P3 · 20 items — see 360° audit §34 for the full register
Release Log · Append-Only
2026-07-27
v3.1–v3.3
Wave 1–3: governance scenarios · PolicyEngine · eval harness · 92.7% registry
2026-07-27
v3.4
AAN engine + Critic v1 · explicit agent recommendations
2026-07-28
v3.4.x
Sign-off fixes · ROI deep-fix (Gate 13) · satellite alignment
2026-07-28
v3.5.x
Visual-system unification (Gate 14) · embedded-PHP watermark removal
2026-07-28
v3.6.x
Byte-identical universal footer · defensible ROI headline (total-cost basis)
2026-07-28
v3.7
pgcf/pdscr/sbaGlobal formulas · sensitivity strip · examiner binder · injection hardening (Gate 15)
2026-07-28
v3.7.1–5
Gap register published · consistency passes · idle-state polish · in-browser Binder Verifier + projector hardening