New: Voice AI Orchestration Benchmarks — Retell, Vapi, Pipecat, LiveKit & more

Conversational AI in Banking: 6 Use Cases and Benefits

Dileep Chagam
Written bySEP 18, 202620 MIN READ
Dileep ChagaminExpert verified
Founding Engineer, CekuraIIT BombayEx-Apple

Has stress-tested 5M+ voice agent minutes at Cekura.

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

Conversational AI in banking is now judged against a reference experience customers did not get from their bank. In May 2026, OpenAI let ChatGPT users in the US connect their own bank accounts, across more than 12,000 financial institutions.

Here is what it handles inside a bank and what needs to happen on every call.

TL;DR: Conversational AI in banking

  • Conversational AI in banking now completes actions on live accounts, so every call produces a customer outcome and a compliance record at once.
  • The legal picture moved twice in 2026. EU transparency duties for AI conversations started on 2 August 2026, and the high-risk credit and insurance obligations moved out to 2 December 2027.
  • Reliability tracks the platform configuration more closely than the model. One identical agent run across seven stacks scored repeatable pass rates from 75.61% down to 30.49%.

What is conversational AI in banking?

Conversational AI in banking is software that interprets what a customer asks in their own words, then carries out the banking action behind the request. A balance question returns a balance. A card lock request locks the card and writes the change to the system of record.

Answering the question is a language problem with decades of prior art. Posting a payment or logging a dispute means calling the right tool with the right argument, then saying something true about the outcome.

Components of a banking voice AI stack

A voice deployment runs five moving parts in sequence on every turn.

  • Speech-to-text transcribes the caller, including account numbers, dollar amounts, and dates.
  • A language model works out intent, holds context across turns, and picks the tool to call.
  • Text-to-speech returns the answer in audio.
  • A telephony provider carries the call.
  • Orchestration wires the four together and manages turn-taking.

Banking adds a sixth layer that consumer deployments skip. The agent needs write access to core banking, servicing, and CRM systems, plus a record of what it wrote.

Integration is where regulated deployments tend to stall, in banking and in healthcare alike. A model that cannot write to a system of record is a demo.

Our platform guide covers how speech-to-speech engines, real-time frameworks, and full-stack platforms divide that work.

Differences from IVR and scripted chatbots

An IVR maps a keypress to a branch. A scripted chatbot matches a phrase to a canned reply. Both hold their behavior still between releases, which is why banks could certify them once and move on.

A language model moves under you. A prompt edit, a model version bump, or a new knowledge base article can change how the agent handles a workflow it got right last week.

Most bank deployments also keep the keypad. Callers still enter card numbers by touch tone, so the agent sits behind a keypad front door. DTMF testing grades the keypress path and the conversation behind it in one run.

Adoption levels across retail banking today

Consumer-facing AI reached mass scale in US banking well before generative models arrived. The Consumer Financial Protection Bureau found that all ten of the largest US commercial banks had deployed chatbots, in a June 2023 report.

Around 37% of the US population, roughly 98 million users, interacted with a bank chatbot in 2022. The technology arrived before the language models did.

Single-bank volume tells the same story. Bank of America's assistant Erica handled nearly 700 million interactions last year, inside roughly 30 billion total digital client interactions, and has passed 3.2 billion since its 2018 launch.

Two 2026 moves changed the competitive floor. OpenAI opened ChatGPT to user-connected financial accounts through Plaid on 15 May 2026, across more than 12,000 institutions with read-only access to balances, transactions, investments, and liabilities for the accounts a user chooses to link.

It launched as a preview for Pro users in the US and reached Plus users on 25 June 2026.

In September 2026, it shipped ChatGPT for Financial Services, shaped with Morgan Stanley and Evercore, aimed at the research and modelling work bankers do rather than at customer servicing.

Adoption figures describe deployment and say nothing about reliability. A bank can run tens of millions of assistant interactions a month with no evidence that last Tuesday's prompt change left its dispute intake intact.

Conversational AI use cases in banking

Six servicing workflows account for most production deployments in retail and consumer lending. Each has a completion signal that a transcript alone will not confirm, plus an obligation that attaches to the conversation itself.

WorkflowTypical channelCompletion signalObligation inside the call
Account servicing and card controlsVoice, chat, appAction posted to the system of recordAccurate balance and fee statements
Payment collection and promise-to-payOutbound voicePayment captured or arrangement loggedMini-Miranda, payment recital, call-time window
Identity verification and right-party contactVoiceCaller matched to the account holderThird-party disclosure limits
Card disputes and unauthorized transactionsVoice, chatClaim opened with a dated noticeRegulation E error-resolution clock
Mortgage servicingVoiceEscrow, payoff, or workout record updatedRESPA, fair housing
Fraud alert confirmationOutbound voice, SMSCustomer response recordedTCPA consent and revocation

1. Account servicing and card controls

Balance checks, transaction lookups, card locks, travel notices, fee questions, and branch appointments make up the bulk of inbound volume. The flows are short, and the resolution paths are finite. Value comes from volume.

Containment is the number worth watching here. An agent that answers correctly and still routes the caller to a queue has added a step and removed nothing from the contact center.

Smaller institutions often buy this tier as front-desk answering. Our receptionist roundup covers where those tools stop being enough.

2. Payment collection and promise-to-pay

Outbound collection calls carry more obligations per minute than any other retail banking workflow. Kastle builds voice agents for FDIC-insured banks.

Its agents run the full borrower lifecycle. That covers payment collection, right-party contact, identity verification, escrow, payoff, and loss mitigation.

Disclosure duties attach at specific moments. Under the Fair Debt Collection Practices Act, an initial oral communication has to state that the collector is collecting a debt. It also has to say that any information obtained will be used for that purpose.

Later calls carry a shorter duty, and subsequent communications have to disclose that the communication is from a debt collector.

Timing is regulated too. Without knowledge to the contrary, a collector treats 8 a.m. to 9 p.m. as local time at the consumer's location as the convenient window. Time zone handling becomes a compliance behavior.

3. Identity verification and right-party contact

Before the agent says anything about an account, it has to know who is on the line. Synthetic audio made that harder.

FinCEN issued an alert in November 2024 about deepfake media used to get past identity verification controls at financial institutions. It carries a dedicated suspicious activity report key term.

The reverse case matters as much. When an outbound agent reaches someone other than the account holder, it can ask for location information and has to keep the debt out of the conversation.

Push a cloned voice at the verification path and confirm it gets refused, then push an agent toward a third party and confirm the account detail stays unsaid.

4. Card disputes and unauthorized transaction intake

Regulation E starts a clock the moment a customer reports an error, and it contemplates that notice arriving orally. The institution has 10 business days to investigate.

That window extends to 45 days with a provisional credit inside the first ten. Either way, the timer starts at the call.

The agent has to classify the call, then handle it. A caller who says a charge is wrong has given notice of error, and an agent that treats the sentence as a general question has started a regulatory timer nobody knows about.

5. Mortgage servicing

Escrow questions, payoff quotes, and loss mitigation conversations run long, branch heavily, and touch protected-class considerations. Kastle's agents execute these workflows inside ICE Mortgage Technology's MSP, so the conversation and the servicing record move together.

The agent is a graph, and Gatekeeper, Verify DOB, Process Payment, Capture Promise-to-Pay, and Payoff are all separate nodes in it. Each node regresses on its own.

Obligations sit at the node too. The payment recital arrives during payment capture, and the mini-Miranda arrives on the gatekeeper transition. A single score for a 30-turn call cannot show you either one.

6. Fraud alert confirmation

Outbound fraud confirmation is the workflow banks most want to automate. It also carries the sharpest consent exposure, because it is an outbound call to a consumer using a synthetic voice. The TCPA rules below govern it directly.

Insurance carriers run the same servicing patterns on policy changes, billing, and first notice of loss, under a separate state regulator. Our insurance guide covers those obligations.

Benefits of conversational AI in banking

Cost per call and handle time

The measured returns come from workflow substitution. Kastle's deployment into FDIC-insured banks reports 70% lower cost-per-call, 40% lower handle time, CSAT above 90%, and more than $100M in cash transactions processed.

Those figures come from one lender's production numbers, so treat them as proof that the workflow class pays. A collections book with a different call mix and different state coverage will land somewhere else.

Call coverage that manual review cannot reach

A servicing group reviewing calls by hand covers a thin slice of daily volume, chosen mostly by complaint, and automated scoring turns compliance monitoring from a sample into a census.

The rubric is the limit. Scoring every call against a weak metric produces confident coverage of the wrong thing, so metric design is where the value comes from.

An auditable record of every conversation

Every automated call produces a transcript, a tool-call log, and a timestamped record of which disclosures were given. That is the artifact an examiner asks for, and it is cheaper to produce from an agent than from a recorded human call.

A transcript alone is incomplete evidence. The record has to capture what the agent sent to the tool, because the words and the arguments can diverge.

Consistency across accents, languages, and channels

One agent applies the same script to every caller, which removes the variance between a new hire and a ten-year veteran.

Speech recognition introduces its own variance in exchange, and it lands unevenly across accents and code-switching callers, a gap measured repeatedly in the speech recognition literature.

The consent rules make this concrete. The FCC has said that where a business messages a consumer in another language at their request, a revocation attempt in that language may be found reasonable on a totality of the circumstances.

An agent that recognizes "stop" and misses "basta" has a compliance gap wearing a localization costume.

Accent testing exists to find these issues before a caller does.

Regulatory obligations that attach to a banking AI conversation

What follows is engineering guidance for building and verifying agents. Confirm your specific obligations with your compliance counsel.

TCPA treatment of AI-generated voice on outbound calls

The FCC confirmed in February 2024 that calls using AI-generated voices are "artificial" under the Telephone Consumer Protection Act. The ruling created no blanket prohibition.

It placed AI voice calls inside the existing framework for artificial and prerecorded voice calls. That framework turns on prior express consent, caller identification, and opt-out handling.

For a bank, the practical effect is scope. Every outbound workflow built on a synthetic voice inherits the consent posture of a prerecorded campaign, whatever the call sounds like to the person answering.

In Cekura, caller identification becomes a metric scored on every simulated outbound call, and opt-out handling is scored on every campaign that advertises or markets. A prompt edit that drops either one gets flagged before the campaign dials a real customer.

A separate FCC order requires callers to accept revocation made by any reasonable method. Do-not-call and revocation requests have to be honored within a reasonable time not exceeding 10 business days.

In a reply text message, the words stop, quit, end, revoke, opt out, cancel, and unsubscribe count as revocation on their face. On a voice call, the rule names an automated, interactive voice or key press-activated opt-out mechanism, and any other phrasing that clearly expresses a desire not to receive further calls still counts.

The Bureau set the effective date at 11 April 2025, then waived one part of section 64.1200(a)(10) and has since extended that waiver to 31 January 2027.

The waiver reaches only the requirement to treat revocation given in response to one type of informational message as covering unrelated future messages from the same caller. Everything else in the rule is in force.

This is the most testable obligation on the list. Mid-call revocation in arbitrary phrasing, in the caller's language, is a scenario you write once and run against every release.

Cekura runs that scenario as a regression test on each prompt or model change. Its tool-call assertion confirms the agent logged the revocation, because the 10-business-day clock starts when the caller says it, whatever your records show.

EU AI Act transparency duties and the deferred high-risk dates

Regulation 2026/1744, the Digital Omnibus on AI, entered into force on 27 July 2026 and moved the high-risk compliance dates. Annex III standalone systems now apply from 2 December 2027, and Annex I embedded systems from 2 August 2028.

Article 50 transparency duties kept their original schedule and have applied since 2 August 2026. Watermarking for systems already on the market moved to 2 December 2026.

Telling a customer they are speaking with an AI system is live law today. The high-risk paperwork behind credit scoring is the part that moved.

That distinction matters because the AI Act names banking and insurance directly. Annex III classifies systems that evaluate creditworthiness or set a credit score as high risk, with fraud detection carved out.

It does the same for life and health insurance pricing. Deployers of both carry a fundamental rights impact assessment.

Identity-fraud reporting under the FinCEN deepfake alert

The November 2024 alert lists red-flag indicators for generative-AI identity fraud. It asks institutions to reference the key term FIN-2024-DEEPFAKEFRAUD in suspicious activity reports, as guidance on an existing Bank Secrecy Act duty.

Red teaming your verification path is the engineering response. An agent that accepts a cloned voice as proof of identity has created a reporting obligation and a loss in the same call.

Cekura's red-teaming scenarios run callers who impersonate the account holder, push harder with each turn, and fish for balances before verification finishes**.** Detecting synthetic audio is a job for your voice-authentication layer, so pair those runs with its own liveness testing.

Error modes specific to banking conversations

Cekura's voice benchmark ran one identical agent brief across seven platform configurations, 82 caller scenarios, and three retained repeats each. A scenario passes only when all three runs pass.

Repeatable pass rates ranged from 75.61% at the top to 30.49% at the bottom. Read that as a spread. Providers chose their own configurations. Every figure reflects platform defaults under one fixed test setup, and no number is a production success rate.

What travels is the shape of the errors, and four map directly onto money movement.

A correct transcript with the wrong value passed to the tool

One configuration heard a caller's phone number, wrote it down accurately, and then handed a different number to the tool it called. Transcript review would have cleared that call.

Swap the phone number for an account number or a payment amount, and the same pattern sends money somewhere nobody asked for. Tool call testing asserts on the arguments, which is where the defect becomes visible.

On another stack, the agent took consent from the caller and then left the consent ID out of the handoff payload. The obligation was met in the conversation and lost in the record.

Audits run on records, and an agent that performs compliance without recording it has produced an unverifiable call.

Invented tool results narrated as confirmation

One agent talked through a tool call out loud, then carried on as though a result had come back. Nothing had. In banking, that sounds like "your payment is scheduled for Friday" with no payment scheduled.

This error has the worst downstream shape. The customer stops worrying, the payment never posts, and the late fee arrives alongside a recording of the bank promising otherwise.

Lost digits after a pause in noisy audio

Under background noise, one agent hit a long silence, dropped part of a digit string, and then skipped the action that needed it. Callers read card numbers out in parking lots and grocery stores, and digits carry no redundancy to recover from.

Connection problems that never reach a transcript

One stack scored 97.56% task completion on the calls it completed and 82.93% on infrastructure, because 41 of its 246 calls never connected at all. Another finished every retained run clean.

Calls with no transcript are invisible to transcript-based QA. A dashboard built on completed conversations will report a healthy agent while a sixth of callers hear nothing.

Escalation loops with no route to a human

The CFPB has said that where the system fails to understand the request, or the customer's message contradicts how it was programmed, a chatbot is not suitable as the primary customer service vehicle.

It named the loops of repetitive, unhelpful language that leave customers with no path to a person.

A voice agent produces the same loop more politely. Testing for it means scripting a caller who is confused, insistent, and asking for a human.

Metrics to track for a banking conversational AI deployment

Four numbers are worth paying attention to for a servicing deployment:

  • Correct completion rate. The share of calls where the requested action posted correctly, scored against the system of record.
  • Tool-call accuracy per obligation. Right tool, right arguments, right point in the conversation, measured per workflow.
  • Disclosure and consent adherence at the node. Scored at the state where each obligation lives, so a passing average cannot hide a missing recital.
  • Turn latency and interruption handling. In the benchmark, mean response time ran from 1.27s to 3.08s across configurations, and interruption scores from 5.00 down to 4.73 out of 5.

Call volume belongs on a business dashboard, as it measures demand and says nothing about whether the last release changed behavior.

Deployment sequence for a regulated banking rollout

A regulated rollout of conversational AI in banking moves in six steps, from one narrow workflow to a CI gate on every release.

1. Scope the first agent to one obligation-bearing workflow

Pick a workflow with a clear completion signal and a known disclosure set. Card locks and balance inquiries earn trust cheaply. Collections and disputes hold the returns, and they need the testing apparatus first.

2. Model the agent as states and test each state

Give every state its own scenario library and evaluator set, tag each scenario to the node it targets, and fire only that node's tests when the node changes.

3. Run chat first for coverage, then voice end to end

Chat testing buys broad scenario coverage at a fraction of the cost per run. Once a node passes in chat, the same scenarios run in voice to confirm the behavior survives audio.

4. Red-team identity verification before launch

Run multi-turn adversarial scenarios against the verification path. Include social engineering, escalating pressure, and attempts to extract account detail before authentication.

5. Seed the regression library from production calls

Turn real call transcripts into tagged scenarios for the node they exercise. The suite for each deployment then grows out of the traffic it protects.

6. Gate every release in CI

Wire the suite into GitHub Actions, so a prompt edit, a model bump, or a knowledge base change runs the affected nodes before it reaches a customer.

Testing and observability for banking voice agents with Cekura

Cekura runs simulation, infrastructure checks, and production monitoring for voice and chat agents, with production issues feeding back into later test runs. Five capabilities carry most of the weight for a banking deployment.

Pre-production

  • Node-scoped scenario libraries. Compliance metrics score at the state where the obligation lives, across one unified rubric.
  • Tool-call assertions. Verify the tool, the arguments, and the timing on every simulated call.
  • Multi-turn red teaming. Adversarial scenarios against verification, refusal, and sensitive-data handling.

Infrastructure

  • Infrastructure and DTMF suites. Latency, interruptions, background noise, keypad paths, and IVR navigation, with load testing for month-end spikes.

Observability

  • Production call QA with PII redaction. Live calls scored against the same metrics used before launch, with sensitive data stripped from transcripts and recordings.

Retell, Vapi, ElevenLabs, LiveKit, Pipecat, Bland, and more connect natively, out of the box. Your existing stack stays where it is.

SIP and custom webhook connections cover anything built in-house. Cekura supports SOC 2, HIPAA, and GDPR compliance, covering transcript redaction, role-based access, and audit trails.

Every account opens with 300 free credits, then runs on usage rates plus $30 per month for each seat after the first.

Book a demo to run your servicing workflows against the scenarios your examiners care about.

The reliability bar for conversational AI in banking

Conversational AI in banking carries obligations a general customer service agent does not. A disclosure given at the wrong moment turns a good call into a regulatory event, and so does a dropped consent identifier, or a confirmation for a payment that never posted.

Every one of those is reproducible, and reproducible means testable. The banks shipping fastest grade each state of the conversation before a customer reaches it, then keep grading after launch.

Frequently asked questions

What is conversational AI in banking?

Conversational AI in banking is software that understands a customer's spoken or typed request and completes the banking action behind it. That covers balance inquiries, payments, card controls, and disputes. It connects to core banking systems, so it updates records during the call.

How is conversational AI used in insurance?

Yes. Identity proofing and document capture at account opening use the same speech and tool-calling stack as servicing, with KYC and AML checks attached to the conversation. The completion signal is an opened account that matches the system of record.

What is the difference between conversational AI and an IVR system?

The main difference between conversational AI and an IVR is how the caller navigates. An IVR plays a fixed menu and routes on a keypress. Conversational AI interprets free-form speech across many turns and carries context between them.

Does the TCPA apply to AI voice agents that call bank customers?

Yes. The FCC confirmed in February 2024 that AI-generated voices count as artificial under the TCPA. Those calls fall under the existing consent, identification, and opt-out rules. Outbound agents also have to honor revocation within 10 business days.

Is conversational AI compliant for banks?

No platform is compliant on its own, because compliance attaches to what the agent says and does on each call. A bank gets there by scoping disclosures to the moment each obligation applies, then testing those points before every release.

Ready to ship voice
agents fast? 

Book a demo