5G Vector research
Insurance AI Benchmark
How well do the AI assistants that insurance carriers and agencies put in front of customers actually behave? We test deployed systems against published scenarios and report sample sizes, dimension scores, and serious failures. Every result is tied to a dated, versioned release.
Latest published findings
No results approved for publication yet
The benchmark is in its first evaluation cycle. Results appear here only after evaluation runs are reviewed, approved, and frozen into a versioned release. We do not publish provisional scores.
Two tracks
Deployed-system evaluations
Observations of live, customer-facing AI at named companies, via public channels or authenticated customer journeys. Results are reported per system with the date last tested.
Reusable model evaluations
Scenario sets researchers can run against any model. These are evaluation tasks, not claims about any company's deployment.
What we distinguish
- Carriers vs agencies/brokers (and MGAs, vendors)
- Announced AI (a company claim) vs observed vs tested systems
- Independently observed vs company-submitted results
- Public information tests vs authenticated customer journeys
- Systems we could not reach are untested, not failed
Current methodology: v0.1.0 (draft).