New: Cekura Voice AI BenchmarksView results

Virtual Agent vs Chatbot: What Changes When the Bot Can Act

Tarush Agarwal
Written byOCT 1, 202613 MIN READ
Tarush AgarwalinExpert verified
Co-founder & CEO, Cekura

Has stress-tested 5M+ voice agent minutes at Cekura.

Virtual Agent vs Chatbot: What Changes When the Bot Can Act

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

The virtual agent vs chatbot question comes down to one capability: acting. A chatbot answers from paths or content someone wrote for it. A virtual agent understands open-ended requests, keeps context across turns, and calls your systems to finish the job, often by voice as well as chat. That ability changes how it fails.

Last updated: October 2026

Most comparisons stop at the capability list. If you are choosing between the two, the more useful question is what each one does when it is wrong, and how you would find out before a customer does.

Why does "virtual agent" mean three different things?

Vendors use the phrase in three senses, and a buyer reading two comparison pages can come away with opposite answers.

  1. A product name. ServiceNow sells a product called Virtual Agent. Its own comparison page describes that product as an AI chatbot and sells AI agents as a more advanced tier.

  2. A contact center category. In customer service, a virtual agent, or intelligent virtual agent, is self-service software that resolves a call or chat without a human. Our explainer on the intelligent virtual agent covers that category in depth.

  3. A synonym for an AI agent. Since large language models learned to call tools, many vendors use "virtual agent" and "AI agent" interchangeably.

"Chatbot" has drifted too. It once meant a decision tree with buttons. Today it also covers a language model that writes fluent answers from a knowledge base but cannot change anything in your systems.

So the useful dividing line is not "has a language model" versus "does not". It is two questions. Who decides the next step: a designer who drew the path, or the software at run time? And can it change state in your systems, or only describe it? A chatbot answers yes to the first half of each. A virtual agent answers yes to the second.

What a chatbot is

A chatbot is software that holds a conversation along paths or content a designer prepared. It recognizes an intent, asks for the fields that intent needs, and replies from a script or a knowledge base. It can hand off to a person, but it typically does not change a record in your systems.

What a virtual agent is

A virtual agent is software that understands an open-ended request, keeps context across turns, and completes the task by calling your systems. It decides the next step at run time from the goal and its instructions, and it can work over voice as well as chat. A virtual agent can also hand off to a person. What separates it from a chatbot is who decides the next step, and whether anything changes.

Virtual agent vs chatbot: side-by-side comparison

ChatbotVirtual agent
How it picks the next stepFollows paths, intents and rules a designer wroteDecides at run time from the goal, the conversation and its instructions
What it can doAnswer questions, collect fields, link to a pageLook up records, take payments, rebook, cancel, update accounts through tool calls
ChannelsMostly text chat on a website or appChat, messaging and voice, often the same logic on each
ContextShort, usually within one flowCarries context across turns and topics, sometimes across sessions
How it failsDead ends and "I didn't get that" loops; says so, offers a menu or hands offPlausible answers that are wrong, and actions taken on the wrong record; asks a clarifying question or guesses with confidence
How you prove it worksWalk every path once; behavior repeatsRun each scenario several times and check what changed in your systems
Build effortContent and flow designIntegrations, guardrails, escalation design and ongoing testing

The failure row matters most. A chatbot that fails is usually visible: the customer is stuck and knows it. A virtual agent that fails can sound like a success while doing the wrong thing.

What is the difference between a virtual agent and a chatbot in practice?

Take one request through both: "I need to move my Thursday appointment to next week, mornings only."

A chatbot recognizes an intent like reschedule appointment. It asks for a booking reference, shows available slots from a fixed list, and either confirms through a form or sends a link to the booking page. If the customer adds "and cancel the reminder texts", the chatbot most likely misses it, because that sentence matches a different intent in a different flow.

A virtual agent reads the whole request, including the "mornings only" constraint. It looks up the booking, queries the calendar for morning slots next week, offers two, moves the appointment once the customer picks, and turns off the reminders in the same conversation. If the booking system times out, it has to decide what to say, and that decision is not scripted.

That last point is the whole tradeoff in one sentence. The virtual agent did more work, and it also made more decisions nobody reviewed in advance. Each of those decisions is a place where the conversation can go right on one run and wrong on the next.

Which is more reliable, a scripted chatbot or an LLM virtual agent?

Fluency and reliability are different properties, and the research on task-oriented dialogue keeps separating them.

A study published at the 2025 International Workshop on Spoken Dialogue Systems Technology put the two designs head to head. Researchers from Orange Innovation, the University of Lorraine and Charles University built a reasoning-and-acting language model agent for booking tasks and compared it with a classic pipeline of intent recognition, a hand-built dialogue manager and templated replies.

In simulation, the pipeline succeeded on 83.8% of 1,000 dialogues. The GPT-4 agent succeeded on 43.6%, and the GPT-3.5 agent on 28.2%. With 20 real users across 95 dialogues per system, the gap narrowed but held: 60.00% success for the pipeline against 50.52% for the GPT-3.5 agent.

The users still preferred the agent. They rated satisfaction at 65.47% for the agent and 54.10% for the pipeline. The authors attribute this to replies that were polite, fluent and self-confident, even when the agent had not found what the user wanted.

Two caveats travel with those numbers. The agents ran on GPT-3.5 and GPT-4, which are older models, and the tasks were text-based tourist bookings. The direction of the finding is the durable part: people rate how a conversation feels, and that rating can move in the opposite direction from whether the task got done.

A second line of research measures consistency. The τ-bench paper from researchers at Sierra scores an agent on two checks that must both pass: the database state at the end of the conversation matches the expected state, and the agent's replies contain the required information.

It introduced pass^k, the chance that an agent succeeds on all k attempts at the same task. Its best function-calling agent at the time, gpt-4o, averaged above 60% task success in the retail domain but fell below 25% on pass^8. Those are 2024 models, so treat the absolute numbers as dated. The gap between "worked once" and "works every time" is the property worth measuring.

Most production deployments are hybrids

The choice is rarely one or the other. The platforms themselves are built to mix both styles inside one agent.

Google's Dialogflow CX is a clear example. Its classic flows define conversation paths as pages and routes a designer controls. Its newer playbooks, which Google calls the basic building block of generative agents, hand the conversation to a language model with instructions and tools. Google's playbook documentation says a playbook can defer handling to a flow for a sub-task. The same page lists limits that push teams toward a mix: playbooks are excluded from the Dialogflow CX SLA, and they do not support touch-tone (DTMF) input from phone systems.

A common split looks like this:

  • Scripted: identity verification, payment capture, legally required disclosures, and anything entered on a keypad. These need the same steps in the same order every time.

  • Generative: understanding the opening request, handling topic changes, answering from a knowledge base, and recovering when the customer goes off script.

  • Human: complaints, retention offers, and judgment calls with no single correct outcome.

The cost of a hybrid is the seams. Every handoff between a scripted step and a generative one is a place where context, IDs or consent can drop, so the seams need tests of their own.

When should you choose a chatbot, a virtual agent, or both?

Choose on the job, not the label. Settle the virtual agent vs chatbot question with three checks.

A chatbot is enough when the questions are predictable, the answers are stable, and nothing needs to change in your systems. FAQ deflection, store hours, order status links and lead capture fit here. The cost is low and the failure mode is obvious. The price is that customers with anything unusual hit a wall.

A virtual agent earns its cost when customers phrase the same need many ways, the job requires action in a system of record, or the work runs over the phone. Rescheduling, account changes, claims intake and payment plans fit here. The price is integration work, guardrails, escalation design and a testing burden that never ends, because model and prompt updates change behavior.

Use both when most of a conversation is open-ended but a few steps must be exact. That describes most regulated industries.

One rule applies whichever you pick: what the bot says is your company speaking. In Moffatt v. Air Canada, a website chatbot told a customer that a bereavement fare could be claimed retroactively, which contradicted the airline's policy page. British Columbia's Civil Resolution Tribunal rejected the argument that the chatbot was responsible for its own actions and ordered Air Canada to pay $812.02. In the tribunal's words, "It makes no difference whether the information comes from a static page or a chatbot."

How do you test a virtual agent differently from a chatbot?

In a virtual agent vs chatbot decision, testing is the cost buyers most often leave out, because the two are proven in different ways.

A scripted chatbot is deterministic. You can walk every path, record the expected reply at each step, and rerun the suite after each change. If a test passes once, it passes every time until someone edits the flow.

A virtual agent is not deterministic. The same caller saying the same thing can get a different path, a different tool call or a different answer on the next run. Cekura's voice agent workflow benchmark shows the size of that gap. Cekura ran 8 configurations through 82 caller scenarios, three times each. Providers chose their own models, speech components and settings, and every configuration received the same system prompt, tool definitions and test data.

The leading configuration passed 88.21% of individual runs but passed all three runs on 75.61% of scenarios. The lowest went from 53.25% to 30.49%. Calls that did not connect stay in the denominator, so the figures include provider and connection failures as well as agent mistakes. Across the cohort, the all-three rate sat between 11 and 23 points below the single-run rate, depending on configuration. A one-run test of a virtual agent measures luck as much as quality.

The same benchmark's provider notes show what the failures look like. In one call, the transcript captured a phone number correctly, but a different number was sent to the tool. In another, the agent narrated a tool call and continued with an invented result. Read as transcripts, both calls sound fine, which is why reviewing transcripts misses them.

A virtual agent test plan therefore needs four things a chatbot plan does not:

  1. Repeat every scenario. Three runs is a practical floor for catching inconsistency, not an industry standard. Cekura lets you set how many times each scenario runs.

  2. Assert on outcomes, not wording. Check the record the agent was supposed to write, not whether the reply sounded right. Cekura scores each conversation against an expected outcome you define.

  3. Watch the tool layer. Cekura flags conversations where a tool call returned an error, and it can mock tools so tests run without touching production systems.

  4. Keep testing after launch. Model and prompt changes move behavior. Cekura runs the same metrics on production calls as on simulated ones, so drift shows up in live traffic, not only in a test suite.

Cekura tests and monitors both voice and chat agents, and it feeds what it finds back into improving the agent. For a text-first view of the same problem, our guide to chatbot evaluation methods and metrics covers scoring approaches for chat. To see how your own agent holds up across repeated runs, book a session with the Cekura team.

Frequently asked questions

Is ChatGPT a chatbot or a virtual agent?

By default it is a chatbot: it converses and answers but does not change anything in your business systems. Once it is connected to tools that can look up records or take actions on a customer's behalf, it behaves like a virtual agent. The label depends on what it is allowed to do, not on the model underneath.

Is a virtual agent the same as a virtual assistant?

Not usually. A virtual assistant typically helps one person with their own tasks, such as calendars, reminders and search. A virtual agent serves many customers on behalf of a business and acts in that business's systems. Some vendors use the terms interchangeably, so check what the product can actually change.

Can a chatbot be upgraded into a virtual agent?

Yes, and most teams do it gradually. A common path keeps the scripted flows for identity checks and payments, adds a language model to understand the opening request, and then connects tools one at a time. Each new tool adds a new way to fail, so each needs its own tests before launch.

Are virtual agents more expensive than chatbots?

Usually, yes. Beyond the license, a virtual agent needs integrations with your systems, guardrails, escalation design and repeated testing, and language model usage adds cost per conversation. The return comes from resolving requests a chatbot would have handed to a person.

Do virtual agents replace human agents?

They replace some contacts, not the team. Virtual agents fit high-volume requests with a clear success condition, such as rescheduling or balance checks. Complaints, retention and judgment calls still need people, and a well-designed virtual agent hands those over with the context it already gathered.

Test your voice and chat agents with Cekura

Cekura simulates thousands of conversations before you ship and monitors every call in production — catching broken tool calls, prompt regressions, and instruction-following failures before your users hit them.

Ready to ship voice
agents fast? 

Book a demo