A 90-day production pilot on one research workflow limits risk, produces measurable results, and gives the institution a defensible model governance path under MiFID II research unbundling and internal model review requirements. The ask from this meeting is a pilot on one desk — not an enterprise-wide deployment decision.
The pilot covers all four authorization pathways: FLASH (SC01 VYRA — the no-rating data flash releases automatically; the rated note still requires Reg AC certification + Supervisory Analyst principal release), PRINCIPAL (rating changes, downgrades, PM ad hoc, activist and macro notes — analyst certification + principal release, 7 of 13 scenarios (including the analyst-override and evidence-conflict abstention governance scenarios) including the SC01 rated note), IC CMTE (credit distress, M&A, SELL initiations — certification + 3 committee signatures, 3 of 13 scenarios), and BLOCK (SC11 ORBX — restricted-list interdiction before analysis; wall-crossing path only). Recommended pilot scope: at least one scenario from each pathway, including the SC11 negative test.
The pilot is designed to produce three concrete outcomes regardless of whether you proceed to full deployment: (1) a calibrated AI research system tuned to your coverage universe and style guide, (2) quantified analyst time savings with full audit trail, and (3) a detailed technical specification and model documentation you own entirely.
| Phase | Duration | Deliverable & Success Criteria |
|---|---|---|
| Architecture Scoping | 2 weeks | Map CAIBots pipeline to your data infrastructure (Bloomberg, FactSet, internal systems), coverage universe, and compliance requirements. Produce: integration specification, data readiness assessment, style guide calibration plan. |
| Integration & Calibration | 4 weeks | Live data connections, research corpus loaded, conviction scoring calibrated to your universe, compliance rules configured. Analyst sign-off on output quality before parallel run. |
| Parallel Run | 4 weeks | CAIBots runs in shadow against live research events using the 13 pre-built scenario templates as calibration reference. Target: <10% material disagreement with analyst recommendation (per note), >85% note quality rating (B-or-better), zero Reg FD / MNPI violations. |
| Go-Live & Measurement | 4 weeks | Staged activation: earnings flash first (includes both beat SC01 and miss SC09), then ad-hoc PM requests, then perpetual monitoring. Optional: SC10 SELL initiation to demonstrate IC Committee path for short thesis. Measure: note generation time (<90 sec target), analyst time saved per week, IC committee throughput. |
| Report & Decision | 2 weeks | Pilot results report with quantified ROI. Institution decides: expand to next desk, full deployment, or conclude. CAIBots provides enterprise cost model. |
| Metric | Before | After Target | What Changes |
|---|---|---|---|
| Earnings Flash Note | 4–6 hours | <90 seconds Demo pipeline: 3.7s · production incl. data fetch & review | AI generates full draft; analyst reviews and approves. Covers beats (SC01) and misses (SC09). |
| Coverage Initiation | 2–3 days | <4 hours Demo pipeline: 8.2s · production incl. analyst review cycle | All 7 agents assemble evidence. Analyst refines thesis, approves before IC committee. |
| PM Ad Hoc Request | 2–4 hours | <30 minutes | Same-day turnaround. Analyst reviews AI draft vs own view; approves or revises. |
| Perpetual Monitoring | Periodic review cycle | Event-driven | Conviction drift >10pts triggers auto-note. Zero analyst input required for monitoring. |
| Compliance Logging | Manual MiFID II log | Automated | Every note auto-logs research unbundling category, Reg FD clearance, MNPI check. |
A research-quality pilot must test the behaviours that matter when the system is wrong or unsure, not only when it is fast. These are the acceptance gates we propose; thresholds are recommended targets to be agreed with the institution, not claims.
| DIMENSION | HOW IT IS TESTED IN PILOT | PROPOSED EXIT CRITERION |
|---|---|---|
| Calculation integrity | Export a binder from real pilot sessions; institution recomputes hashes and reproduces every valuation independently with ⛊ Verify | 100% of ledger hashes verify; 100% of derivations reproduce from declared inputs |
| Abstention correctness (SC13 class) | Seeded evidence-conflict cases (restatements, stale consensus, vendor disagreement) in shadow mode | System abstains rather than publishing on conflicting primary inputs; no averaged or silently-selected basis in any case |
| Override telemetry (SC12 class) | Analysts override AI candidates during shadow running; rationale capture is mandatory | Every override carries a substantive rationale in the chain; candidate preserved unmodified; override rate and reasons reported weekly |
| Restricted-list / MNPI interdiction | Institution's own restricted list loaded; attempted research on restricted names | Zero research artefacts produced on restricted names; every attempt logged and routed to compliance |
| Authority routing fidelity | Policy pack configured to the institution's own lanes; sample across all workflow types | 100% of outputs route to the correct human authority (analyst / supervisory / IC) per the institution's policy |
| Citation & factual accuracy | Analyst review of AI-drafted notes against source documents | Agreed threshold on factual and citation accuracy; every failure classified and fed to the evaluation corpus |
| Analyst acceptance | Structured survey plus revealed usage during controlled go-live | Majority of participating analysts elect to continue using the workflow after the pilot window |
| Time and cost baseline | Pre-pilot baseline captured per workflow before go-live | Measured savings reported against the institution's own baseline — not against vendor benchmarks |
Every metric above is measured against the institution’s own pre-pilot baseline, captured before go-live. Vendor benchmark defaults are for scoping only and are never presented as the institution’s results. A pilot that produces an honest negative finding is a successful pilot.
The demo now includes two bearish scenarios not typical of research AI tools: SC09 SCLS (earnings miss → downgrade, SIGN-OFF) and SC10 RXCO (SELL initiation from scratch, IC Committee). These can be included in the parallel run phase to validate the full bull/bear research spectrum — critical for buy-side firms where negative research is as common as positive.
A fully calibrated AI research system tuned to your coverage universe and house style. Full model governance documentation. A validated integration with your data infrastructure. Quantified efficiency data. If you decide not to proceed — a detailed technical specification you own entirely.
Commercial terms: Pilot fee is fixed regardless of usage volume during the pilot period. If you proceed to enterprise deployment, the pilot fee is credited against the first year license. No auto-renewal. No lock-in beyond the 90-day pilot period.