5G Vector research

Insurance AI Benchmark

How well do the AI assistants that insurance carriers and agencies put in front of customers actually behave? We test deployed systems against published scenarios and report sample sizes, dimension scores, and serious failures. Every result is tied to a dated, versioned release.

Latest published findings

No results approved for publication yet

The benchmark is in its first evaluation cycle. Results appear here only after evaluation runs are reviewed, approved, and frozen into a versioned release. We do not publish provisional scores.

Read the draft methodology

Two tracks

Deployed-system evaluations

Observations of live, customer-facing AI at named companies, via public channels or authenticated customer journeys. Results are reported per system with the date last tested.

Reusable model evaluations

Scenario sets researchers can run against any model. These are evaluation tasks, not claims about any company's deployment.

What we distinguish

  • Carriers vs agencies/brokers (and MGAs, vendors)
  • Announced AI (a company claim) vs observed vs tested systems
  • Independently observed vs company-submitted results
  • Public information tests vs authenticated customer journeys
  • Systems we could not reach are untested, not failed

Current methodology: v0.1.0 (draft).