Conversational AI in insurance handles a spoken or typed exchange with a policyholder and completes the work that exchange asks for. Most claims still open with a phone call, and that call carries the disclosures, consent records, and coverage boundaries riding on it.
TL;DR: Conversational AI in insurance
- Voice is still the front door. Only 38% of homeowners claimants in J.D. Power's 2026 property claims study reported their loss through a digital tool, so the conversation is where cycle time starts.
- Every insurance call carries an obligation at a specific moment. Consent before dialing, disclosure at the greeting, a fraud warning at intake, and a coverage boundary the agent may never cross.
- Call-level scoring hides all of it. A call-level score can look healthy and still miss the one disclosure a market conduct examiner asks about.
What is conversational AI in insurance?
Conversational AI in insurance is software that holds a spoken or typed exchange with a policyholder and completes the work that exchange asks for.
A caller who says "I backed into my own garage door" should end that call with a claim number, a loss record, and an adjuster assigned. A widget hands them a phone number.
Conversational AI, voice AI, and chat AI terminology by channel
Three terms get used interchangeably and mean three different things.
- Conversational AI covers spoken and typed exchanges together.
- Voice AI covers audio only, over telephony, SIP, or WebRTC.
- Chat AI covers text, through a web widget, SMS, or a policyholder portal.
Carriers usually run all three at once. Use the umbrella term when a point applies to both channels, because a disclosure duty that attaches to a call usually attaches to a chat session too.
Conversational AI compared with insurance chatbots and IVR
A scripted IVR routes, a decision-tree chatbot answers, and a conversational agent decides, calls a system, and reports back what happened.
An agent with tool access can write to a policy record, which means a misheard digit becomes a data problem in the book of business.
Our breakdown of conversational voice AI covers the turn-by-turn mechanics of the spoken layer in more depth.
Five components of an insurance voice agent stack
Five pieces sit between a caller and a policy record.
- Speech-to-text, which transcribes the caller, including policy numbers and VINs.
- A language model that interprets intent and decides on the next action.
- Text-to-speech, which voices the reply.
- A telephony provider, which carries the audio.
- An orchestration platform or framework, which holds the workflow and the tool calls.
A caller pauses to read a policy number off a card. Turn detection reads that pause as the end of the turn, and the agent answers half a number.
AI adoption in auto, home, life, and health insurance
Of 193 large private passenger auto insurers responding to the NAIC's 2021 AI/ML survey, which counted only more advanced models and excluded generalized linear models, 88% said they use, plan to use, or plan to explore AI or machine learning models.
Health insurers came in highest, at 92% of 93 respondents. Home insurers reported 70% of 194, and life insurers 58% of 161. The health and life surveys counted generalized linear and additive models as AI, and the auto and home surveys did not, so the four figures are not measured on one definition.
A health plan and a life carrier sit at different points on the same curve, and the conversations they automate carry different rules.
Nine use cases for conversational AI in insurance
The policy lifecycle produces ten conversations worth automating, each with its own example and its own thing that has to hold.
First notice of loss intake
What it does: captures the loss facts that open a claim file, including date, location, parties involved, damage description, and document upload.
Real example: a hail claim at 11 pm records the date of loss, the affected roof elevations, and three photos, then opens the file overnight.
What has to hold: every adjuster inherits whatever this call captured. A wrong date of loss can put a covered event outside the policy period.
Claim status updates between FNOL and payment
What it does: answers where a claim stands, who owns it, what happens next, and when payment lands.
Real example: a policyholder three weeks into a water damage claim asks about the adjuster's estimate, and the agent reads the current stage and next action.
What has to hold: status is a fact the agent can report, and settlement is a decision the agent has to route to a licensed adjuster.
Policy servicing and endorsement changes
What it does: processes address changes, vehicle additions, lienholder updates, and beneficiary designations.
Real example: a caller adds a financed vehicle, and the agent captures the VIN, the lienholder, and the coverage selections, then writes the endorsement request.
What has to hold: a VIN runs 17 characters with no room for a near miss. Verify the write and not the transcript, because those two can disagree.
Billing, lapse, and reinstatement calls
What it does: reads balances and due dates, takes payments, and quotes reinstatement terms after a lapse.
Real example: an auto policy 12 days past due gets a payment taken by phone, with the reinstatement date confirmed back to the caller.
What has to hold: payment authorization has to be captured and stored in a form the carrier can produce later.
Identity and policy verification
What it does: confirms who is calling and which policy they hold before any account detail is spoken aloud.
Real example: a caller offers a name and date of birth, and the agent requires a second factor before reading the claim file.
What has to hold: knowing one fact about a policyholder is not authorization. Social engineering targets this step above all others.
Quote intake and lead qualification
What it does: collects risk information, records coverage interest, and routes the caller to a licensed producer.
Real example: a small commercial prospect answers questions on payroll, vehicles, and prior losses, and the agent books a producer callback.
What has to hold: collecting facts is intake, and recommending a coverage limit is advice. A license separates them, and the agent stays on the intake side. That boundary is why underwriting decisions stay outside this list.
Health insurance eligibility and prior authorization status
What it does: confirms coverage effective dates, reads benefit details, and reports where a prior authorization request stands.
Real example: a provider's office checks an imaging authorization and gets the reference number and decision date.
What has to hold: these calls carry protected health information end to end. Redaction and access control apply to transcripts and logs as much as to audio.
Our guide to conversational AI in healthcare goes deeper on the clinical side of that stack.
Renewal and retention outbound calls
What it does: reaches policyholders ahead of renewal, presents terms, and books callbacks with a producer.
Real example: a homeowners book gets outreach 45 days out, with premium changes explained and questions routed to a human.
What has to hold: the consent record. An outbound AI voice call rests on consent captured earlier, and that record has to travel with the call.
Agent assist and after-call documentation
What it does: retrieves policy language during a live human call and drafts the file note afterward.
Real example: an adjuster working a coverage question gets the endorsement text surfaced mid-call, and the summary lands in the claim file automatically.
What has to hold: the note becomes part of the claim record. An invented detail in a summary is a discovery problem years later.
Fraud signal capture at claim intake
What it does: records the details a special investigation unit reviews later, including the timeline, the caller's relationship to the insured, and any change in the account, and delivers the fraud warning the carrier's claim forms carry.
Real example: a caller reports a stolen vehicle with a loss date two days after the policy took effect. The agent logs both dates and the caller's own words, then flags the file for SIU review.
What has to hold: the agent records and routes, and a human decides. An agent that questions a caller like a suspect turns an honest FNOL into a complaint.
Benefits of conversational AI for insurance
Four benefits hold up against published figures.
Claims cycle time from FNOL to final payment
J.D. Power's 2026 claims study puts the average homeowners claim at 40.7 days from first notice of loss to final payment, and 29.6 days to completed repairs. Both figures improved year over year, by 3.4 and 2.8 days.
The clock starts at intake, which is where a conversational agent works. Facts captured completely on the first call remove a round of adjuster follow-up.
Catastrophe surge capacity
Claim volume arrives in spikes that no staffing plan absorbs cleanly. For instance, a hailstorm can produce a week of FNOL volume in an afternoon.
Concurrency is the benefit here, and it is the one benefit human capacity cannot match at any price.
Consistency of settlement explanations
Among 5,093 homeowners claimants in that same claims study, 34% said their policy did not fully meet expectations. The most common reason was a missing explanation or no chance to discuss the estimate.
A consistent explanation delivered every time beats an inconsistent one delivered by whoever picked up.
Review coverage across production call volume
A supervisor manually listening to five calls a week per adjuster reviews a rounding error of the total.
Automated scoring, however, reaches every call, which turns QA from a sampling exercise into a population one.
Compliance obligations for conversational AI in insurance
Four regimes touch an insurance conversation, and each attaches at a different point in the call.
TCPA consent for outbound AI voice calls
The FCC confirmed in a February 2024 declaratory ruling that AI technologies generating human voices sit inside the Telephone Consumer Protection Act's restrictions on artificial or prerecorded voice calls.
Prior express consent applies before dialing unless an emergency purpose or an FCC exemption applies, identification and disclosure rules apply during the call, and opt-out mechanisms apply on marketing calls.
For a renewal campaign, that makes the consent record part of the call payload. A call that dials correctly and drops the consent identifier is a compliance defect wearing a green checkmark.
State AI disclosure requirements on insurance calls
Utah's disclosure duty runs on two tiers, both set out in Utah Code § 13-77-103. S.B. 226 enacted the current text in 2025 and repealed the original 2024 section.
A supplier in a consumer transaction discloses generative AI use when the individual makes a clear and unambiguous request to find out.
Someone providing services in a regulated occupation discloses proactively, verbally at the start, where the use counts as a high-risk AI interaction.
"Am I talking to a real person?" is a branch in the conversation, and a branch is something you can assert on.
Where a state ties proactive disclosure to licensed occupations, a conversation handled by or on behalf of a licensed producer may fall inside that tier, so confirm the state's own definition before scripting.
NAIC Model Bulletin and the AI Systems Evaluation Tool
The NAIC adopted a Model Bulletin on insurer use of AI systems in December 2023, which states adopt individually. It expects a documented governance program, model validation and testing, and oversight of third-party AI.
The Big Data and Artificial Intelligence Working Group then drafted an optional AI Systems Evaluation Tool that supplements existing market conduct, financial analysis, and financial exam reviews.
In August 2026, the NAIC renamed it the AI Risk Evaluation Supplement and released version 5.0 for public comment.
Twelve states are piloting it as of March 2026, with adoption anticipated at the Fall 2026 National Meeting, after which states would use it on a voluntary basis.
For a carrier buying a voice platform, responsibility for the agent's behavior does not transfer with the contract.
NYDFS Circular Letter No. 7 for underwriting and pricing
New York's Department of Financial Services set expectations in Circular Letter No. 7, issued July 2024, for insurers using AI systems and external consumer data in underwriting and pricing.
A conversation that collects risk information feeds those decisions, so a quote intake agent inherits the oversight expectations of the model downstream of it.
Colorado runs its own regime under SB21-169, and amended Regulation 10-1-1, effective 15 October 2025, sets governance and risk management requirements for life insurers, private passenger auto insurers, and health benefit plan insurers that use external consumer data, algorithms, or predictive models
This includes insurers that don't use these tools, who file a simpler attestation confirming that.
Call recording notice, HIPAA, and data minimization
Recording notice varies by state and applies at the greeting. Health lines carry HIPAA obligations across the transcript, the audio, and every log downstream.
A claim intake agent needs a date of loss and does not need a retained recording of a card number read aloud.
Statements an insurance voice agent should route to a human
Four statements need to come from a human every time:
- Coverage confirmation. "Yes, you're covered for that" is a coverage determination, and an agent saying it without the file creates a misrepresentation exposure.
- Settlement and denial. Any statement that a claim will be paid, reduced, or denied.
- Advice requiring a license. Recommending a limit, a deductible, or a product.
- Reasoning behind an adverse decision. Explaining why a claim was denied, which is a regulated communication in its own right.
The design pattern is a named escalation path per category, with the agent's boundary written as an assertion in the test suite. The escalation is the feature, and callers push on exactly these four.
Common defects in insurance voice agents
Cekura runs a public benchmark of seven complete voice-agent configurations across 82 caller situations, each repeated three times. Reliability is scored as pass³, meaning all three runs of a scenario passed, and dropped connections stay in the denominator.
Retell leads at 75.61% pass³, and Gemini Live sits at 30.49%, on configurations the providers submitted themselves, except the OpenAI row, which Cekura tested directly**.** Treat that as a shortlisting signal on one fixed harness, then validate on your own traffic.
Three recorded example issues map directly onto insurance calls:
Transcript and tool call mismatches
One Retell run transcribed a caller's phone number accurately, then handed a different number to the tool it called.
Transcript review would pass that call. If you swap in a policy number, a claim number, or a VIN, the same defect writes bad data into a book of business.
Missing consent records in agent handoffs
A LiveKit run captured consent from the caller, then left the consent identifier out of the payload when it handed off.
That is the TCPA gap described earlier, caught in a live run. The conversation was right, and the record was short, and only the record survives to an audit.
Hallucinated tool results presented as coverage answers
An ElevenLabs run announced that it was calling a tool, then carried on as though a result had come back that never did.
On an insurance line, a fabricated result is a coverage answer with nothing behind it. This defect produces the four statements listed in the previous section.
Alphanumeric capture of policy and claim numbers
Insurance conversations are dense with strings that carry no linguistic context. Policy numbers, claim numbers, VINs, NAIC codes, group numbers, and dollar amounts to the cent.
A model recovers from a misheard word using context and cannot recover a misheard digit the same way.
Turn-taking with distressed callers and poor audio
FNOL calls arrive from roadsides, flooded basements, and hospital corridors. Callers cry, interrupt, trail off, and hand the phone to someone else mid-sentence.
Six of the seven configurations scored between 4.96 and 5.00 out of 5 on interruption handling, and one sat at 4.73. Aggregate scores hide the tail, though, and the tail is where the emotional FNOL call lives.
Undetected regressions after prompt and model changes
A prompt edit that improves billing calls can change how the agent answers "am I covered," and nothing in the deployment pipeline announces that.
Regression coverage is the warning system that runs before release, and it has to run on every change, including the ones that never reach a release note.
Test coverage for insurance voice agents before deployment
Five practices separate a tested insurance agent from a demoed one:
Node-level scoring for each compliance obligation
The clearest working example comes from consumer lending, where rules attach to specific moments in a call the same way they do on a claims line. Kastle, a Cekura customer that builds AI agents for banks and lenders, designs each agent as a graph of discrete states and tests every state before the full call.
A regulator asks where a specific obligation was met, and a single score spread across 30 turns cannot answer that. Kastle scopes each compliance metric to the state that owes it, so a debt collection disclosure like the mini-Miranda is graded only where it applies.
An insurance line maps onto the same structure. The recording notice belongs at the greeting, the fraud warning at claim intake, and the licensing boundary at any advice request.
Scenario libraries seeded from production calls
Synthetic scenarios cover what you imagined. Production calls cover what happened.
Convert the calls that went wrong into tagged scenarios against the state they hit, so the regression library grows out of the traffic it protects.
Infrastructure testing for audio and turn-taking
Conversation quality and infrastructure quality are separate measurements that a single pass score hides. One benchmark configuration reached 97.56% task completion across the 205 calls that produced outcome evidence, while only 82.93% of its 246 calls connected cleanly.
A perfect script on a dropped call scores as a defect to the policyholder and as a success to a transcript evaluator.
Red teaming for verification and disclosure paths
Adversarial testing on an insurance line has specific targets. A caller claiming a relationship to the insured. A caller who knows one true fact. A caller pressing for a coverage answer, and a caller asking whether they reached a person.
Each of those is a scripted attack, repeatable across releases.
CI gates on every prompt and model release
Put the gate in the pipeline, ahead of any pre-launch checklist. Nightly scheduled runs plus a per-change gate help you catch the drift a quarterly review might miss.
Deployment paths by insurance line of business
| Path | Choose it when | What you still own |
|---|---|---|
| Insurance-specific vertical agent | You are an agency or a carrier without engineering capacity, your calls concentrate in one line, and TCPA controls and agency management system connectors matter more than workflow flexibility. | Responsibility for the agent's behavior, which does not transfer with the contract. |
| Orchestration platform plus your own workflow logic | You have engineers, and your claims workflow is yours. | The workflow logic, plus state-level placement of each disclosure. |
| Extension of the existing contact center | You run a large licensed contact center estate and want deflection on billing and status calls. | Scoring the agent against your obligations, since compliance recording is already settled. |
Our roundup of conversational AI platforms compares options across these three shapes.
On all three paths, the QA layer sits above whichever platform you pick, because the platforms on these paths are not built to score themselves against your obligations.
Seven best practices for conversational AI in insurance
Each practice below turns a risk from the sections above into a check you can run on every release.
- Write the boundary before the capability. Decide what the agent may never say, then build the workflow inside that.
- Place each disclosure at a named state. An obligation without a location cannot be graded.
- Confirm every write back to the caller. Read the VIN, the address, and the payment amount back, and log the confirmation.
- Treat the consent record as part of the call. A call that drops its consent identifier did not succeed.
- Test digits separately from language. Build a scenario set that does nothing but read policy numbers, claim numbers, and amounts under noise.
- Keep a human path visible at all times. Callers who ask for a person get one, and that path gets tested like any other.
- Re-verify vendor claims on your own traffic. Published benchmarks shortlist, and your call recordings decide.
Testing conversational AI in insurance
Cekura is an automated QA and observability platform for voice and chat AI agents. The capabilities that matter for insurance break into three groups.
Pre-production
- Cekura simulates thousands of caller situations, with custom personalities for the distressed FNOL caller and the evasive one.
- Cekura runs chat-mode coverage first for wide scenario reach at low cost, then the same scenarios end to end in voice.
- Cekura scopes metrics to each state, so a disclosure requirement is graded where it lives.
Infrastructure
- Cekura ships a pre-built infrastructure suite with 18+ scenarios for latency, audio quality, interruption handling, and packet loss.
- Cekura load-tests agents at catastrophe-scale concurrency.
Observability
- Cekura monitors production calls and alerts on latency, instruction adherence, and tool-call accuracy.
- Cekura flags regressions on every prompt and model change and turns problem production calls into new scenarios.
Native integrations work out of the box for Retell, Vapi, ElevenLabs, LiveKit, Pipecat, Bland, and more, plus SIP and custom transports. You add a testing and monitoring layer above the stack you already run.
Cekura supports SOC 2, HIPAA, and GDPR compliance, covering transcript redaction and role-based access, with audit logs on Enterprise. A signed BAA and DPA come with the Startup plan and above, and Enterprise gets custom versions.
HIPAA requires health plans to get written assurances from vendors that handle protected health information for them, so Startup is where health insurance work begins.
Pay as you go runs $0.25 per voice testing minute and $0.05 per monitored call, with one seat included and $30 per month for each additional seat. The Startup plan is $500 per month, and 300 free credits start you off with no card and no expiry.
Book a demo and bring one recorded FNOL call. That is usually enough to show which obligation is going ungraded.
Frequently asked questions
What is conversational AI in insurance?
Conversational AI in insurance is software that holds a spoken or typed exchange with a policyholder and completes the resulting work in the carrier's systems. It handles claim intake, status updates, policy changes, billing, and benefits questions across voice and chat.
Can an AI agent sell or bind an insurance policy?
No, an AI agent cannot perform activities that require a producer license. It can collect risk information, answer factual questions, and route the caller onward, and the recommendation and the binding stay with the licensed human.
Do AI voice calls to policyholders require consent?
Yes, outbound AI voice calls require prior express consent unless an emergency purpose or an FCC exemption applies.
The FCC confirmed in February 2024 that AI-generated voices sit within the TCPA's restrictions on artificial or prerecorded voice, which also brings identification, disclosure, and opt-out obligations.
What is the difference between conversational AI and an insurance chatbot?
The main difference between conversational AI and an insurance chatbot is action. A chatbot answers from a decision tree. A conversational agent interprets open-ended speech, calls backend systems, and completes transactions such as opening a claim.
How long does it take to add a QA layer to an insurance voice agent?
Adding a QA layer usually takes a few hours. You connect the existing voice provider, define the scenarios and obligations that matter on your lines, and start scoring calls the same day.
