From Hype to Reality: Talking Real-Time Voice Design with the Experts

Last week, I wrote about the rise of Real-Time Voice Agents in Dynamics 365 Contact Center and why they represent such a significant shift in customer experience.

The technology is exciting. The demos are impressive. And let's be honest, we've all had that moment where we've thought: "This is amazing... but how would I actually design one?"

Because the real challenge isn't turning the technology on; the challenge is deciding:

  • When should a voice agent handle the interaction?

  • When should a human take over?

  • How much freedom should the AI have?

  • What happens when customers don't behave as expected? (Spoiler: they won't.)

So this week, I decided to go straight to the source. This week, I sat down with Sree& Rob, TTEC Digital's Contact Center and AI specialists, to talk about real-world design decisions, implementation considerations, and the questions every organisation should be asking before launching a real-time voice agent.

So this means i have a good old fashioned Q&A with them both below

Grab a coffee…or a Matcha!. Let's get into it.

‍ ‍

Q1.How do you decide whether a process is suitable for a Real-Time Voice Agent?

ROB:Real-time voice agents use AI to decide what actions to take, so this come down to picking use cases that can follow responsible AI principles. I say 'can' instead of ‘are’ because sometimes you'll need to introduce extra steps into an existing process to achieve this rather than just reproducing the current human process. It's mostly about ensuring these agents are subject to human oversight for all important decisions and that you can explain why those decisions were made. The Copilot Studio foundations these voice agents in D365CC are built on makes this easy to achieve because there's a catalog of easy to set up approval mechanisms and data connectors - allowing decisions to be sent up to humans - as well as built in AI transparency - so you can always see what the AI chose to do and why.

SREE:Real-Time Voice Agent suitability by plotting processes on a spectrum between highly process-driven interactions and those requiring significant human judgement. On one side are scenarios that are decision-based, data-driven, and context-aware, where the agent can identify intent, retrieve information, execute actions, and follow well-defined business rules. For example, "I want to update information on my policy" is a strong candidate for automation because the intent is clear and the outcome is predictable. On the other side are scenarios requiring empathy, multiple exception paths, negotiation, or significant interpretation, such as "My claim was rejected and I need help to resolve it because it's making my life harder." These interactions often need human understanding and adaptive decision-making. A key factor in this assessment is intent identification. If the customer's intent can be confidently understood and mapped to a structured outcome, a Real-Time Voice Agent can usually handle it effectively. If the intent is emotionally driven, ambiguous, or requires complex judgement, the interaction is better suited for advisor engagement or a hybrid AI-to-human handoff model.


Q2.What’s the biggest design mistake organisations make?

ROB:Designing these voice agents as though you are building a last-gen IVR is the biggest mistake. If you design with exact sequences of steps of prompts and responses that you expect to see when testing, this won’t realise the benefits of the much more flexible non-linear capabilities that agentic AI brings.Designing for these new agents is instead about defining the goal and the boundaries - What should and shouldn’t it do, but not the exact steps to get there, you then need to think about how you will measure the results you’re getting. Just like humans, no AI will follow all instructions perfectly, so you need to define a catalog of ‘evaluation’ cases that you can run at scale to measure and report on the results - given this scenario the bot should do X, Y and Z. That needs to be run regularly so you can decide if and where to target efforts to improve the results.

SREE:The biggest design mistake organisations make is focusing on automating individual contact centre processes rather than looking at the end-to-end customer journey and broader business transformation. In my experience, many teams spend significant time designing deflection scenarios, such as handling customer chaser calls, without stepping back to ask why customers are calling to chase in the first place. If customers are contacting the organisation because they have not received an update, missed an expected service date, or lack visibility of progress, then the real opportunity is to fix the underlying process and proactively communicate with customers. By addressing the root cause, those calls never reach the contact centre, which delivers far greater value than simply automating them. A second common mistake is designing only for the ideal happy path and not considering real-world customer behaviours and exceptions. Organisations need to think about scenarios such as vulnerable customers, emotional situations, customers calling from noisy environments, poor audio quality, language barriers, or cases where the customer's intent is unclear. A successful Real-Time Voice Agent design should not only focus on automation rates but also on how the solution adapts to these situations and when it should intelligently transition to another channel or a human advisor. The most successful implementations are those that treat Voice AI as part of a customer experience transformation rather than as a standalone automation tool.‍ ‍


Q3.What separates a good voice agent from a great voice agent?

SREE:I like to think of it in the same way we differentiate between a good advisor and a great advisor in contact centre. A good advisor can answer questions and follow processes. A great advisor knows when to bring in others, seeks specialist help when needed, and focuses on getting the best outcome for the customer rather than trying to do everything themselves. The same applies to Voice Agents. A good Voice Agent can recognise intent and automate tasks. A great Voice Agent acts as a digital advisor, helping customers navigate their journey while recognising its own limits. It always keeps the human in the loop, knowing when a specialist, advisor, or back-office team should be involved. Rather than measuring success by how many interactions it contains, it measures success by how effectively it collaborates with humans to achieve the right customer outcome.In my view, the best Voice Agents are not designed to replace advisors. They are designed to augment them, handling what they can, gathering context, reducing effort, and ensuring that when human expertise is needed, the customer is connected to the right person at the right time with the right information. ‍ ‍

After reading Sree's response, it got me thinking about how we often define success when it comes to AI. So much of the conversation focuses on containment rates, automation percentages, and how many interactions a Voice Agent can handle without human intervention. But is that really the right measure? When we think about the best customer experiences we've had ourselves, they rarely come from being prevented from speaking to a person. They come from getting the right help, at the right time, with as little effort as possible.Perhaps the future of Voice Agents isn't about replacing people or keeping customers away from them. Perhaps it's about creating a partnership between human and digital colleagues, where each plays to their strengths to deliver the best possible outcome.
Which leaves me with a question for organisations to consider: are we building AI to maximise automation or are we building it to maximise customer success?


Q4.How much freedom should we give AI?

SREE:I think of it like giving a bird space to fly. If you put it in a tiny cage, it can't do what it was designed to do. But if you release it into the entire world with no boundaries, it may go somewhere you never intended. The right approach is to give it a safe environment to operate in, with enough freedom to be effective but within known boundaries. The same applies to AI and Real-Time Voice Agents. We shouldn't give AI unrestricted access to all the knowledge it possesses. Instead, we should define the area in which it can operate confidently using approved knowledge, business processes, and integrations. For example, if a Voice Agent is supporting a council contact centre, it should be focused on council services and customer enquiries. It doesn't need to know about space exploration or random topics unrelated to its role. A successful Voice Agent is not one that knows everything, but one that knows the right things. Start with controlled boundaries, allow natural conversations within those boundaries, and expand responsibly as confidence and trust grow.‍ ‍


Q5.Just because AI can do something, doesn't mean it should. How do you decide what decisions a Real-Time Voice Agent is allowed to make on its own?

SREE:My approach is to treat a Real-Time Voice Agent like a new advisor joining the contact centre. On day one, we don't give a new advisor every process, every exception scenario, and unrestricted access to every system. We start with a controlled scope, let them build experience, monitor performance, and gradually increase responsibility. The same principle applies to AI. Rather than trying to automate every scenario from the start, I prefer an iterative approach. Begin with a small set of well-understood, low-risk journeys, observe how the Voice Agent handles customer intent, where it succeeds, where it struggles, and what exceptions emerge. As confidence grows, you can expand its knowledge, processes, and decision-making authority.

The other critical factor is remembering that a Voice Agent is only as good as the data, knowledge, and processes we provide to it. Many people focus on the AI model itself, but the real foundation is the quality of the business data, knowledge articles, and process design behind it. So, when deciding what a Voice Agent should be allowed to do, I focus less on what the AI is capable of and more on whether we have the right knowledge, controls, and data to support that decision. Start small, measure, learn, improve, and expand over time. That's typically how you build trust in the solution and create a Voice Agent that delivers consistent customer outcomes.‍ ‍

Sree's answer got me thinking about customer expectations. As consumers, we're generally far more forgiving when a new human advisor takes time to learn than we are when technology gets something wrong. We expect people to improve through experience, but often expect AI to be perfect from day one.Perhaps that's why trust is such an important part of any Voice Agent strategy. Customers don't necessarily care how advanced the underlying AI model is. They care whether it understands them, whether it resolves their issue without unnecessary effort, and whether they feel confident in the outcome.The organisations that succeed with Voice Agents may not be the ones that deploy the most sophisticated AI first. They may be the ones that take the time to build customer trust, proving value journey by journey, interaction by interaction, before expanding further.

Which raises an interesting question: in the race to deploy AI, are we spending enough time thinking about what customers need in order to trust it?


Q6. If you could connect a Real-Time Voice Agent to only three data sources, which would you choose and why?

SREE:
Operational Guardrails & Guidance Framework
A controlled set of instructions, policies, escalation rules, and behavioural guidelines that govern how the Voice Agent responds to exceptional or sensitive situations. This ensures the agent can recognise and appropriately handle scenarios such as customer distress, safeguarding concerns, off-topic conversations, or conversations that fall outside its intended domain, while knowing when to escalate to a human advisor.

CRM / System of Record
Provides customer context such as profile, history, active cases, policies, and previous interactions, enabling the Voice Agent to deliver personalised and relevant responses.

Customer-Facing Knowledge Base
A curated repository of FAQs, guidance articles, policies, and service information that enables the Voice Agent to provide accurate, consistent answers. I would intentionally exclude internal-only documents to avoid exposing information not intended for customers.
 ‍ ‍


Q7. What early warning signs tell you a Voice Agent design is likely to fail before it even goes live?

SREE: In my experience, a Real-Time Voice Agent is only as good as the information provided to it. While AI models can use their general knowledge, I generally don't recommend relying on that for customer service scenarios when a strong, curated knowledge base exists. If the knowledge is incomplete, outdated, or inconsistent, the Voice Agent will struggle to provide reliable outcomes. Strong data foundations, trusted customer-facing content, and clear business guidance are often better indicators of success than the AI technology itself. Another early warning sign is inconsistent behaviour during testing. If the Voice Agent gives different answers to the same question, struggles with slight variations in customer wording, or performs well in some scenarios but poorly in others, it often points to underfitting or overfitting in the knowledge, instructions, or training approach. A successful Voice Agent should be predictable, consistent, and resilient across a range of customer interactions before it goes live. 


Q8. If you could give a customer just one piece of advice before they start designing a Real-Time Voice Agent, what would it be?

SREE: In my experience, organisations spend a lot of time evaluating models and features, but the success of a Real-Time Voice Agent is usually determined by the quality of the knowledge, business processes, and customer outcomes behind it. Start with a small number of high-volume, well-understood journeys, ensure your customer-facing knowledge is accurate and maintained, and design for exceptions from day one.Also, ask yourself not only whether your knowledge base is good, but whether it is AI-ready. Many organisations have comprehensive knowledge articles, but they are written for human advisors who can interpret context, navigate multiple documents, and fill in gaps. A Voice Agent needs content that is structured, clear, consistent, and consumable by AI. If the knowledge cannot be understood and applied effectively by the AI, even the best model will struggle to deliver the right outcome.

Sree's advice also made me reflect on my own role in pre-sales. When demonstrating Voice Agents, it's easy to focus on the art of the possible. We showcase the natural conversations, the integrations, the speed, and the intelligence. But perhaps one of the most important conversations isn't about what the AI can do today, it's about what needs to be in place for it to succeed tomorrow. Increasingly, I find myself asking customers different questions. How well maintained is your knowledge? Are your processes documented and consistent? Is your content structured in a way that both humans and AI can understand? Because the reality is that a Voice Agent can only be as effective as the foundation it is built upon.

The most successful demonstrations aren't the ones that leave a customer saying, "That's impressive." They're the ones that leave a customer thinking, "Are we ready for this?"

Perhaps, as pre-sales professionals, our role isn't simply to demonstrate the technology. Perhaps it's to help organisations understand what it takes to turn that technology into a great customer experience.
After all, is the real value of a demo showing what AI can do, or helping customers understand what they need to do to make it successful? or maybe it is both!


Q9. Any Final Thoughts

SREE: Real-Time Voice Agents are not an AI project, they're a customer experience transformation supported by AI. From my experience, success comes from focusing on the fundamentals: understanding why customers contact you, fixing the underlying processes, providing trusted and AI-ready knowledge, and starting with a controlled scope before expanding capabilities. The AI model itself is rarely the limiting factor. The quality of the customer data, knowledge base, business rules, and operational guidance will determine the outcome.Treat the Voice Agent like a new advisor. Give it clear instructions, trusted knowledge, and well-defined responsibilities. Let it learn through an iterative approach, keep humans in the loop for complex scenarios, and gradually increase autonomy as confidence grows. Most importantly, don't measure success by how many calls are contained. Measure it by how effectively you reduce customer effort, resolve customer needs, and improve the overall customer journey. When designed this way, a Voice Agent becomes more than an automation tool; it becomes a valuable digital team member that works alongside advisors to deliver better customer outcomes..


After talking to Sree and Rob, I've come away thinking differently about Real-Time Voice Agents.

Going into this blog, I expected the conversation to focus on AI models, automation capabilities, and technical design decisions. Instead, the recurring theme was something much more fundamental: customer experience.

What struck me most was that neither of them measured success by how many calls a Voice Agent could contain or how many tasks it could automate. They talked about customer outcomes, trust, human oversight, knowledge quality, and designing experiences that make life easier for customers and advisors alike.

As someone who spends a lot of time demonstrating this technology to organisations, it's also made me reflect on my own role. Perhaps the most important part of a demo isn't showing what the AI can do. It's helping customers understand what they need to do to make it successful. Do they have clear processes? Trusted knowledge? The right guardrails? Are they designing around customer needs or simply trying to automate existing problems?

If there's one thing I've learned from this conversation, it's that great Voice Agents aren't built by starting with AI. They're built by starting with the customer.

The technology may be what gets people's attention, but customer experience is what ultimately determines success.

And perhaps that's the question we should all be asking ourselves as we explore this new era of AI-powered service:

Are we designing Voice Agents to automate conversations, or are we designing better experiences for the people having them?

Next
Next

Understanding Real-Time Voice Agents in Dynamics 365 Contact Center