Run AI Pilots - From Idea to Evidence to Scale (Federato, Five Sigma, Akur8 Deploy-to-Filing)
Running an insurance AI pilot at the executive scale is not the same as testing a vendor. A pilot at L5 is a sequenced evidence-generation program with named gates at 90 days, 180 days, and 12 months, an explicit thesis the chief actuary signed before the pilot started, an instrumented champion-control parallel run that produces ASOP-41-compliant attribution, a kill protocol the AI committee will actually execute, and a path-to-scale that the executive committee authorizes in advance so the Horizon-2 production rollout doesn't relitigate the pilot's gate-pass decision. The three canonical 2026 pilot patterns covered in this lesson - Federato RiskOps agentic underwriting on professional liability submissions, Five Sigma agentic claims on auto first-party-property, and Akur8 deploy-to-filing on small-commercial property pricing - are the ones the rated-carrier population is actually piloting at scale right now, and they map directly to the three named Horizon-1 pilots in the playbook from Lesson 1. This lesson is the pilot operating system: the 90/180/12 gate logic, the kill protocol that prevents zombie pilots, the three canonical pilot designs in detail, and the path-to-scale handoff that converts pilot evidence into Horizon-2 production capability.
The Pilot Thesis the Chief Actuary Signs Before Day One
Before a pilot starts, the chief actuary signs a pilot thesis document that states the named in-scope book, the expected combined-ratio impact range (base case and downside case), the attribution methodology that will separate AI lift from market conditions and rate action, the champion-control parallel-run protocol, the success criteria at 90/180/12 gates, and the kill criteria at each gate. The thesis document is signed by the chief actuary, chief underwriter (for UW pilots) or chief claims officer (for claims pilots), CRO, and AI committee chair. The signed thesis is the document the executive committee references when the gate decision arrives - without it, gate decisions devolve into post-hoc storytelling.
Sample thesis for Federato RiskOps on professional lines: in-scope book is professional liability submissions in the carrier's primary appointment territories, expected impact 0.8-1.1 combined-ratio points sustained at 12 months attributable to AI under documented methodology, base case 0.95 combined-ratio improvement, downside case 0.3 with a kill trigger at month 9 if measured impact is below 0.4 points on shadow-mode parallel run, success criteria at 90 days includes shadow-mode parallel run operational and submission throughput within plan, 180-day criteria includes measurable appetite-discipline impact on the controlled cohort, 12-month criteria includes 0.6+ AI-attributable combined-ratio points and complaint-ratio holding against control.
The 90-Day Gate - The Shadow-Mode Evidence Bar
The 90-day gate is the operational-readiness check, not the impact check. The bar: shadow-mode parallel run operational (the AI is making recommendations on real submissions or claims that the human underwriter or adjuster sees alongside their own decision, with both decisions logged for later comparison), data fabric flowing cleanly (the AI has access to the data it needs without integration breaks), governance posture established (the use case in the algorithm inventory, the model card current, the bias-testing pipeline operational on the controlled cohort, the incident-response runbook tested with one tabletop exercise), and team adoption documented (the underwriters or adjusters in scope have completed training and are engaging with the AI's recommendations rather than ignoring them).
The 90-day gate decision is binary: shadow-mode operational and clean - proceed to instrumented measurement period. Shadow-mode not operational or clean - extend by 30-60 days with a remediation plan, or kill the pilot. Carriers who pass the 90-day gate on a low bar pollute their later attribution measurement; the gate has to be rigorous about operational cleanliness before the impact measurement begins.
The 180-Day Gate - The Direction-and-Trend Check
The 180-day gate is the directional-evidence check. The bar: measurable lift visible on the controlled cohort against the matched control, lift trending in the expected direction (the curve is moving, even if it has not reached final magnitude), no material complaint-ratio degradation, no regulatory red flags surfaced in the use case's documentation review, vendor performance within SLA, no material adverse customer-experience signal.
The 180-day gate decision has three branches: positive directional evidence with no red flags - proceed to 12-month measurement at full scale on the controlled cohort. Mixed directional evidence (some metrics moving, others flat) - narrow scope and continue, with explicit kill trigger at month 9 if narrowed scope doesn't produce evidence by then. Negative directional evidence or material red flag - kill the pilot, document the lessons, reallocate the capital. The 180-day kill is the most important kill in the pilot operating system because it prevents the zombie-pilot failure mode where a pilot runs to 12 months and produces inconclusive evidence that nobody can act on.
The 12-Month Gate - The Attribution and Path-to-Scale Decision
The 12-month gate is the attribution-and-scale decision. The bar: AI-attributable combined-ratio impact measurable under the chief actuary's signed methodology against the controlled cohort, sustainability assessment from chief actuary and CRO (does the lift hold under the carrier's portfolio mix and the next twelve months of market conditions?), governance posture clean (algorithm inventory current, model cards refreshed, bias testing documented, incident-response tested twice), vendor concentration acceptable (the pilot's vendor isn't pushing the carrier past the 28% concentration cap on critical decision flow), and path-to-scale operationally feasible (the carrier has the talent and infrastructure to extend the pilot to the full in-scope book in Horizon 2).
The 12-month gate decision is the pilot-to-production transition the playbook anticipates. Success at the gate means the pilot transitions to Horizon-2 production rollout with named milestone gates for full-scope extension; partial success means the pilot transitions with narrower scope; failure means the pilot retires with documented lessons and the freed capital reallocates per the executive committee's review.
The Kill Protocol That Prevents Zombie Pilots
The zombie-pilot failure mode kills more transformation programs than any other: a pilot runs past its evidence horizon, produces inconclusive data, gets extended for "one more quarter" repeatedly, and consumes capital and attention while never reaching a decision. The kill protocol prevents this with explicit triggers and a documented decision cadence.
Kill triggers at each gate: 90-day - shadow-mode not operational, data fabric breaks, team not engaging, or governance posture not established. 180-day - no measurable directional evidence on the controlled cohort, material complaint-ratio increase versus control, regulatory red flag surfaced, vendor SLA breach. 12-month - AI-attributable impact below the downside case in the signed thesis, sustainability concerns from chief actuary or CRO, vendor concentration pushing past acceptable, or path-to-scale not operationally feasible.
The kill decision cadence: AI committee reviews pilot status quarterly with the chief actuary, chief underwriter or chief claims officer, CRO, and chief AI officer. Any of those executives can flag a kill consideration; the AI committee chair convenes a focused review within 14 days; the decision is documented with appeal path. Honest pilot retirement is a credibility action; running a doomed pilot to the next gate is a credibility loss.
Pilot Pattern One - Federato RiskOps Agentic Underwriting
The Federato RiskOps pilot is the canonical agentic-underwriting pilot at the rated-carrier scale in 2026. Scope: professional liability submissions in the carrier's primary appointment territories, typically 8,000-15,000 submissions over a 12-month pilot period at a $1.2B specialty commercial carrier. The Federato workbench ingests submissions, applies appetite-discipline logic, surfaces portfolio-aware recommendations, and supports agentic dispositioning where the carrier has authorized it.
Pilot design: champion-control parallel run on a matched cohort, with Federato-augmented underwriters working a portion of submissions in shadow mode for the first 90 days and in production for months 4-12; matched-cohort underwriters working the control submissions on the traditional workbench. The chief underwriter sets the appetite logic; the chief actuary signs the attribution methodology; the CRO monitors vendor concentration and incident-response coordination; the AI committee reviews quarterly.
Sample evidence at 12-month gate: AI-attributable 0.8 combined-ratio points on the controlled cohort with chief-actuary-signed attribution; submission cycle-time compression 35-45% on the controlled cohort; appetite-aligned binding ratio 87% on Federato cohort vs 79% on control; underwriter productivity 22% higher on the controlled cohort; complaint ratio holding against control with no material adverse signal. Path to Horizon 2: extend to professional lines across all appointment territories, with named milestone gates at three-month intervals.
Pilot Pattern Two - Five Sigma Agentic Claims
The Five Sigma pilot is the canonical agentic-claims pilot in 2026, with reference deployments at Starr (2025) plus Sutherland partnership and Covéa's parallel Shift Claims deployment serving as the comparable agentic-claims architectures. Scope: auto first-party-property claims (Tier 1 complexity) in selected states, typically 3,500-6,500 claims over a 12-month pilot period at a $1.2B carrier with material auto exposure.
Pilot design: tiered shadow-mode parallel run where Five Sigma handles a portion of Tier-1 claims in shadow mode with adjuster sign-off on every disposition for the first 90 days, then progresses through limited production (25% of Tier-1 with sign-off) at days 90-180, then full Tier-1 production at days 180-365. Champion-control parallel run on a matched cohort of Tier-1 claims handled traditionally. Chief claims officer's complaint-ratio monitoring discipline governs the tier expansion cadence.
Sample evidence at 12-month gate: 0.6 AI-attributable combined-ratio points on the controlled Tier-1 cohort (ALAE reduction 48%, cycle-time compression 62%, severity discipline 12%); complaint ratio at 0.98x control (no material degradation); reserve-development pattern clean on Tier-1; chief actuary attestation on sustainability. Path to Horizon 2: extend to homeowners and small-commercial property first-party claims through the model-promotion process from Ch2-3.
Pilot Pattern Three - Akur8 Deploy-to-Filing
The Akur8 deploy-to-filing pilot is the canonical pricing-platform pilot in 2026. Scope: small-commercial property or homeowners pricing in selected states, with Akur8's transparent GLM/GBM running the indication and Akur8's Rate Repo and Deploy modules supporting the filing pipeline. Typical pilot scope at $1.2B specialty carrier: one state-line combination (e.g., homeowners in Georgia) with full rate-indication, filing, and post-effective monitoring through Akur8.
Pilot design: the pricing platform runs in parallel with the carrier's existing pricing tool for the first 90 days (shadow mode on rate indications), then the chief actuary selects the Akur8 indication for the next state-line filing at day 90-180, then the filing goes effective and 6-12 months of post-effective monitoring closes the pilot. Akur8 Discover (with the Matrisk acquisition deepening filing intelligence) supports competitor-filing monitoring. The chief actuary signs every step; the algorithm inventory captures the rating logic; the chief compliance officer ensures Akur8's filing-explainability artifacts satisfy state DOI requirements (Colorado Reg 10-1-1, NAIC AI Systems Evaluation Tool documentation, state-specific bulletins).
Sample evidence at 12-month gate: rate-adequacy improvement on the in-scope state-line, measurable through actual-vs-expected loss-ratio on the post-effective book at 9-12 months; filing-cycle-time compression 45-60% versus the carrier's traditional filing process; explainability artifacts cleanly satisfy state DOI bulletin requirements with no follow-up RFAI; chief actuary attests Akur8 produces signature-ready rate-indication work product. Path to Horizon 2: extend Akur8 to additional state-line combinations on the same line, then to adjacent lines, per the carrier's filing roadmap.
The Path-to-Scale Handoff - From Pilot to Production
The pilot-to-production handoff is where most transformation programs lose momentum. The handoff has six elements: documented pilot evidence pack (including the 12-month gate decision memo with chief actuary attribution), production architecture readiness (MLOps platform supports production deployment with monitoring, observability, and rollback), governance artifacts current (algorithm inventory updated for production scope, model card refreshed for production version, bias testing rerun on production cohort, incident-response runbook updated), production scope definition (named lines, named territories, named milestone gates for full-scope extension in Horizon 2), production talent assigned (Algorithm Inventory Owner, AI Product Manager, MLOps Lead all attached to the production capability), and treaty-broker briefing on production scope so the next renewal can address the AI capability in cession language.
The handoff is documented in a production-transition memo signed by the AI committee chair, chief underwriter or chief claims officer (or chief actuary for pricing pilots), chief AI officer, and CRO. The memo authorizes the Horizon-2 production rollout with named milestone gates and resource allocation. Without the memo, the pilot's lessons get rediscovered in production and the path-to-scale resets to a new pilot with the same questions.
The Pilot Instrumentation Stack
The instrumentation that makes attribution honest covers six surfaces. (1) Champion-control parallel run with matched-cohort design so the chief actuary can compute AI-attributable lift. (2) Decision logging - every AI recommendation and every human decision captured with timestamp, input data, and reason chain. (3) Outcome tracking - every bound submission's loss development tracked to attribution close, every closed claim's actual final loss vs reserve estimate tracked, every rate-filed indication's post-effective loss-ratio tracked. (4) Customer-experience monitoring - complaint ratio against control, NPS or customer-survey signal, downstream cancellation and non-renewal patterns. (5) Vendor performance monitoring - SLA adherence, incident history, support responsiveness, model performance metrics from the vendor's MLOps platform. (6) Governance artifacts - algorithm inventory entry currency, model card refresh cadence, bias testing cadence, incident-response tabletop history.
Without all six surfaces, attribution is contestable and the 12-month gate decision is fragile. With all six, the AM Best analyst, the DOI examiner, the treaty broker, and the carrier's auditor have the documentation they need at their respective review moments.
Key Takeaways
- The pilot thesis the chief actuary signs before day one anchors the gate decisions. Named in-scope book, expected combined-ratio range (base and downside), attribution methodology, champion-control parallel-run protocol, success criteria at 90/180/12, kill criteria at each gate. Signed by chief actuary, chief UW or chief claims officer, CRO, AI committee chair.
- 90-day gate: operational-readiness check. Shadow-mode parallel run operational, data fabric clean, governance posture established (algorithm inventory entry, model card, bias testing, incident-response runbook tabletop), team adoption documented. Binary: proceed to measurement or extend with remediation.
- 180-day gate: direction-and-trend check. Measurable lift on controlled cohort, trending in expected direction, no complaint-ratio degradation, no regulatory red flag, vendor SLA. Three branches: positive direction → proceed to 12-month at scale; mixed → narrow and continue with month-9 kill trigger; negative → kill and reallocate. The 180-day kill is the most important kill - it prevents zombie pilots.
- 12-month gate: attribution-and-scale decision. AI-attributable combined-ratio impact under chief actuary methodology, sustainability assessment, governance posture clean, vendor concentration acceptable, path-to-scale operationally feasible. Pilot transitions to Horizon-2 production with named milestone gates.
- Federato RiskOps pilot pattern: professional lines submissions, 8,000-15,000 over 12 months at $1.2B specialty carrier. Champion-control matched cohort, 0.8 AI-attributable combined-ratio points typical at 12 months, 35-45% cycle-time compression, 87% appetite-aligned binding vs 79% control.
- Five Sigma agentic claims pilot pattern: auto first-party-property Tier 1, 3,500-6,500 claims over 12 months. Tiered shadow-mode → limited production → full Tier-1 production. 0.6 points AI-attributable (ALAE -48%, cycle-time -62%, severity discipline 12%) with complaint ratio at 0.98x control.
- Akur8 deploy-to-filing pilot pattern: one state-line combination, shadow indications → live filing → post-effective monitoring. Rate-adequacy improvement measurable at 9-12 months; filing-cycle-time compression 45-60%; explainability artifacts satisfy state DOI bulletin requirements with no follow-up RFAI.
- The path-to-scale handoff has six elements: documented evidence pack, production architecture readiness, governance artifacts current, production scope definition, production talent assigned, treaty-broker briefing. Production-transition memo signed by AI committee chair, chief UW/claims/actuary, chief AI officer, CRO authorizes the Horizon-2 rollout.
Skill.re