Skip to main content
SYNTHETIC MARKETS — every issuer, figure, and event on this page is fictional · This page documents what the demo suite PROVES and what it does NOT · Not research, not investment advice
Benchmark & Coverage · Suite V4.0

What Is Measured.
What Is Not Done.

Every claim the Investment Research demo makes is either mechanically gated, derived from declared inputs, or listed below as an open gap. This page is the benchmark of record for the suite — including the part most product pages omit: the register of what has not been built or proven.

Coverage · two numbers, one vocabulary
69.0%
Demonstrated
13 research areas in full, 3 in part, of 21. Denominator declared and enumerated on this page.
28/28
Proven
Every case agrees. IRVAL conviction composites, tiers and valuation bases were independently reimplemented in Python with Decimal arithmetic — from the published specification, not transcribed from the JavaScript — and every case executes twice. 28 of 28 agree across both implementations. Mutation-tested: injected faults were each caught, including a tampered policy threshold that consistency-checking alone would have missed.
Why two numbers. Demonstrated is what runs in the demo today. Proven is what an evaluation harness has verified against an independently reimplemented oracle — a higher bar than a demo lane, because it tests the logic rather than the rendering. Across the CAIBots suite only Credit Underwriting currently publishes a Proven figure. The other three show an em-dash, which is the honest answer until a harness exists. Both numbers use the same vocabulary so the four estates can be compared without translating between them.

Where this denominator comes from — and why there is no external one

No published standard enumerates investment research capability. That is worth saying plainly, because the other estates in this suite can point at one and this estate cannot.

StandardWhat it actually governs
FINRA 2241Analyst conflicts of interest, disclosure and firewalls. Conduct, not capability — it says nothing about what analysis should be performed.
Reg ACAnalyst certification that views are their own honest opinion. A certification requirement, not a scope of work.
MiFID IIResearch unbundling and inducements. Commercial arrangement, not analytical coverage.
CFA ROSResearch objectivity standards — independence and integrity of process. Again conduct.

Every one of these is a conduct regime. None enumerates the analytical work a research function should be able to do, and buy-side and sell-side would draw that list differently from each other in any case. Inventing a citation here would be worse than admitting there is none.

So the denominator is self-declared, with a published criterion. An item is in the 21 if it is analytical work a research desk is routinely asked for and whose absence a portfolio manager or IC would notice — earnings response, initiation, rating change, credit recovery, merger arbitrage, macro sweep, activist situation, and so on. Each is enumerated on this page with the lane that exercises it.

Excluded by design. Trade execution and order routing (an OMS/EMS function, not research). Portfolio construction and optimisation (a PM function). Client suitability and advice (a wealth function under a different regime). Proprietary alpha models (institution-specific by definition). Each is excluded because it belongs to a different system, not because it is hard.

What 69.0% therefore means. Thirteen areas in full and three in part of twenty-one that we enumerated and defended above. It is checkable against the register printed on this page and against nothing else — which is a weaker claim than KYC's FFIEC citation, and is stated as such rather than dressed up.

Quantified Coverage

How much of the territory this demo exercises

69%
Demonstration Coverage
Denominator stated, not implied: 21 research-and-compliance capability areas relevant to a research deployment. 13 exercised in full · 3 in part (half-weight) · 5 not exercised — itemised below, expanded in “What Is NOT Done”. This measures the DEMONSTRATION against that denominator, not any production research function, and it is intentionally not 100%.
EXERCISED IN FULL · 13
✓ Earnings-event research (flash → note) — flash lanes
✓ Coverage initiation (full package) — initiation lanes
✓ Valuation cross-check (DCF vs comps) — Valuation Agent
✓ Fundamental analysis from filings — Fundamental Agent
✓ Sentiment and positioning read — Sentiment Agent
✓ Macro overlay / FOMC events — Macro Agent
✓ Risk and scenario bounding — Risk Agent
✓ Credit / IC committee workflow — committee lane
✓ Reg AC attestation chain — cert → sign → distribute
✓ Supervisory Analyst release gate — role signatures
✓ MNPI / restricted-list hard block — statutory-class, sandbox-locked
✓ Numeric grounding before distribution — grounding check
✓ Hash-chained audit + independent derivations — 6 gates · 10 derivations
EXERCISED IN PART · 3
◐ Model-risk governance (SR 11-7 class) — sandbox locks + versioned pack shown; validation out of demo scope
◐ Recordkeeping (17a-4 class) — session chain + export; production WORM store is pilot scope
◐ Best-execution / trade linkage — portfolio lane references it; execution systems not exercised
NOT EXERCISED · 5
○ Live-AI output quality (no evaluation corpus yet)
○ Licensed market-data integrations (Bloomberg/FactSet/EDGAR simulated)
○ Production distribution plumbing (entitlements, CRM, mailer)
○ Non-equity asset classes at depth (rates, FX, structured)
○ Multi-language / non-US regulatory regimes
The Release Battery

Six Gates, or No Release

No file ships unless the full battery prints ALL 6 IR GATES PASS — packaging is literally grep-gated on that string. Every founder-caught field defect became a permanent assertion.

GATEWHAT IT PROVES
GATE 1 · CLAIMSBans stale counts, era text, absolutist compliance language, weak-hash relics — including en-dash variants that once evaded ASCII bans.
GATE 2 · STRUCTURETag balance and full script parse on every page.
GATE 3 · VISUAL CANONOne nav, one byte-identical footer, banner canon, version stamps, feature presence — 60+ assertions.
GATE 4 · HEADLESS RUNTIMEPages must RUN clean, not merely parse — null-for-absent-ids shim born from a field crash.
GATE 5 · REAL DOMjsdom executes every page: zero errors, computed-visible nav, BOTH auth paths of the demo must yield all scenario pills; pure-SHA parity vs node crypto.
GATE 6 · VALUATION ORACLEIndependent reimplementation of every engine formula, asserted against the page's own declared inputs — ten derivations, six composites, drift-bans.
The Valuation Oracle

Ten Derivations, Independently Reimplemented

Displayed figures are computed at runtime by IRVAL v1.0 from declared inputs — and Gate 6 reimplements every formula independently and asserts the outputs against the page's own data. Conviction tiers are composites of five declared dimensions (threshold: ≥75 HIGH · ≥50 MEDIUM). You can reproduce all of it yourself: export the ⬖ Research Binder and drop it on ⛊ Verify — your browser recomputes every SHA-256 and every derivation with no server involved.

SCENARIOTICKERBASISDECLARED INPUTSFORMULADERIVED
SC01VYRApricePT $215 · price $183(PT − price) / price+17.5%
SC02SFGpricePT $240 · price $213.70(PT − price) / price+12.3%
SC03PHCrecoveryentry 100 · recovery 55(recovery − entry) / entry−45.0%
SC04VLTNpricePT $300 · price $395.30(PT − price) / price−24.1%
SC05CLSPpricePT $465 · price $410(PT − price) / price+13.4%
SC06Nimbusarbtarget $36.25 · offer $38.50(offer − target) / target+6.2% arb
SC08pricePT $540 · price $478.70(PT − price) / price+12.8%
SC09pricePT $420 · price $480(PT − price) / price−12.5%
SC10pricePT $38 · price $72(PT − price) / price−47.2%
SC12HLXpricePT $96 (candidate) · price $82.50(PT − price) / price+16.4%
Three scenarios are honestly non-numeric by design: SC07 (portfolio-level reweight), SC11 (interdicted — no research produced), SC13 (abstained — dual basis preserved, neither published).
Governance Coverage

Thirteen Scenarios · Policy-Laned · Verifiable

Routing lanes come from the fingerprinted policy pack (IR-2026.07) — the same pack whose hash stamps every ledger row. Change one rule in the ⚖ Sandbox and watch the fingerprint move; the restricted-list hard block is displayed locked: statutory-class controls are not sandboxable.

SCSCENARIOLANEHUMAN AUTHORITYENGINE-DERIVED
01Earnings FlashPRINCIPAL_FASTAuto-clear · rated note certified + principal-releasedYES
02Coverage InitiationPRINCIPALAnalyst sign-offYES
03Credit · IC CommitteeCOMMITTEEIC signaturesYES (recovery basis)
04PM Ad Hoc DowngradePRINCIPALAnalyst sign-offYES
05Perpetual MonitorPRINCIPALAnalyst sign-offYES
06M&A Event-DrivenCOMMITTEEIC signaturesYES (arb basis)
07FOMC Macro SweepPRINCIPALAnalyst sign-offN/A — portfolio-level reweight, no single PT by design
08Activist / 13DPRINCIPALAnalyst sign-offYES
09Earnings Miss DowngradePRINCIPALAnalyst sign-offYES
10Short Thesis InitiationCOMMITTEEIC signaturesYES
11Restricted / MNPIBLOCKInterdicted — no research producedN/A — blocked upstream by design
12Analyst OverridePRINCIPALCandidate sealed · override + rationale ledgeredYES (candidate derived)
13Evidence ConflictCOMMITTEEAbstained · routed with conflict quantifiedN/A — abstained by design (dual basis preserved)
Scored behaviors, not error states: abstention correctness (SC13) and override telemetry (SC12 — candidate sealed beside verdict and rationale) are benchmark dimensions of a governable research system.
The Gap Register — Public By Design

What Is NOT Done

A benchmark that only lists what works is marketing. These are the open gaps, stated plainly; each one is pilot or production scope, and several are hard gates.

01
Live-AI output quality
Unproven. Live Custom calls Claude with your key; research quality of live generations has no evaluation corpus yet.
02
Market-data integrations
Bloomberg / FactSet / EDGAR feeds are simulated. Licensed integrations are pilot scope, not demo reality.
03
Production runtime & SLA
All runtimes are demo-compressed orchestration; projected production figures are design targets (two-clock model), not measurements.
04
Institutional MRM validation
Everything here is engineering validation. It is NOT a client institution's model-risk validation and is never represented as such.
05
Statutory wording
Reg AC / Reg FD / MiFID II descriptions are PENDING counsel validation — a hard gate before any external compliance claim.
06
SOC 2 Type II
A funded post-raise milestone. 'SOC 2-Aligned' describes control design intent, not an audit outcome.
07
Valuation depth
IRVAL v1.0 derives upside, recovery, arb spread, and composites. Full DCF build-ups, SOTP, and comps normalization are not yet mechanized.
08
Access control
Page credentials are client-side courtesy gates, source-visible. Production requires SSO / RBAC / tenant isolation.
09
Compliance controls
Restricted-list / MNPI / Reg FD behavior is policy-pack-driven within the demo. Production-grade surveillance and wall enforcement are integration scope.
10
Workflow integrations
No OMS, research-archive, CRM, or distribution integrations exist today.
11
Evaluation corpus
The oracle gates prove arithmetic. A research-quality benchmark (factual accuracy, citation fidelity, analyst agreement) is pilot scope.
12
Pricing & pilot terms
Indicative until a scoped design-partner agreement exists.
Run the Demo — Then Verify It ⛊ Book a Call