Verizon
Building an AI Sandbox to evaluate the future of customer support
Verizon wanted to understand how generative AI could improve customer support without committing to a specific AI model or architecture. I led the experience design of an AI-powered sandbox that enabled teams to prototype, compare and measure AI assistants using synthetic customers, realistic support journeys and evaluation metrics.
Project activities
The project was delivered through a series of five key work streams and sub work streams - all with a UCD approach.
My role
-
Led experience design
-
Facilitated stakeholder workshops
-
Defined testing methodology
-
Designed end-to-end sandbox
-
Created customer journeys
-
Designed reporting framework
Key project challenges and successes
-
AI model landscape changing rapidly
-
No objective way to compare LLM performance
-
Needed to prove value before implementation
-
Personalisation required realistic customer data
-
Solution needed to remain model agnostic
01.
Designed an LLM-agnostic platform
02.
Created synthetic customer ecosystem based on real data
03.
Developed measurable AI evaluation metrics
The objective
Choosing an AI model wasn’t the problem, proving it was the right one was.
Generative AI is evolving at an extraordinary pace. Every few months, new models promise better performance, lower costs and improved reasoning. For Verizon, selecting the wrong model could mean committing to an expensive customer support architecture that quickly becomes outdated.
Rather than recommending a single AI model, we proposed a different approach: build a controlled sandbox where Verizon could safely evaluate AI assistants before making strategic implementation decisions.

Understanding Verizon’s current experience
Before considering the future, we needed to understand the experience Verizon already had.
Before exploring AI, we needed to understand how customers currently experienced Verizon support. We audited existing chatbot and IVR journeys, reviewed visual and sonic branding, and tested common customer scenarios to identify friction, inconsistencies and opportunities for improvement.
-
Reviewed existing chatbot and IVR journeys
-
Audited visual and sonic brand expression
-
Tested common customer support scenarios
-
Identified experience gaps and recommendations

Defining how AI should behave
Creating a consistent AI personality
An effective AI assistant is more than accurate answers it should feel like Verizon. Working closely with stakeholders, we defined the conversational principles, tone of voice and behavioural guidelines that every AI interaction should follow, regardless of channel.
-
Defined conversational principles
-
Established AI tone of voice
-
Created reusable system persona guidelines

Defining personalisation
Verizon knew any LLM generated experience had to match the level of personalisation their service already performed, but they didn't have a baseline.

Personalisation is often discussed in broad terms, but rarely defined consistently. We facilitated workshops with Verizon to establish what a personalised customer support experience should look like, what customer data should influence AI behaviour, and which support journeys offered the greatest opportunity for improvement.
This work established the foundation for every experience tested within the sandbox.

Designing AI support journeys
Customer journeys became controlled experiments for testing AI.
Rather than designing ideal future-state journeys, we created realistic troubleshooting scenarios based on common Verizon support issues. Each journey combined customer context, service history and behavioural traits to test how different AI assistants responded under identical conditions.
These journeys enabled us to compare not only issue resolution, but also how effectively each model personalised conversations, retained context and adapted its communication style.

Building the sandbox
A platform for testing AI before deployment
Rather than building a single chatbot, we designed a sandbox where Verizon could safely test AI assistants under controlled conditions. The platform allows teams to create synthetic customers, compare AI models, evaluate prompts and measure customer outcomes before implementation.
-
Created synthetic customer profiles
-
Designed structured testing workflows
-
Enabled AI vs AI conversation testing
-
Built prompt and model comparison tools




Turning complexity into a usable workflow
As the platform evolved, I iteratively refined key workflows to reduce cognitive load, improve discoverability and make complex AI evaluation tasks intuitive for the Verizon staff.
While the MVP successfully enabled AI evaluation, testing quickly revealed that the experience itself needed refinement.
Because we didn't have a large amount of time to roll this sandbox out to the Verizon team, we needed to ensure the platform was an intuitive as we could.

The original experience made it difficult to understand what each prompt contained and how different configurations varied. I explored a more transparent selection flow that allowed researchers to compare prompt structures, inspect individual instructions and select configurations without leaving the task.
Building an AI evaluation framework
The first step towards measuring AI performance consistently.


Measuring the quality of AI conversations is inherently difficult, and many of today’s evaluation approaches are still evolving. Rather than trying to create a definitive scoring system, we focused on establishing the first iteration of a repeatable evaluation framework.
This provides Verizon with a foundation to compare models consistently, while continuing to refine the metrics, testing methodology and reporting as the sandbox matures.
Looking ahead
From MVP to enterprise capability
The sandbox was intentionally designed as an MVP, providing a foundation for future expansion. While the first iteration focuses on chatbot experiences using synthetic customer data, the architecture is designed to scale across additional support channels, richer datasets and more advanced evaluation capabilities.
-
Integrate real Verizon customer data
-
Expand into My Verizon, IVR and agent assist
-
Extend testing across more customer journeys
-
Continue evolving evaluation metrics
Outcome
The project delivered a working AI sandbox that enables Verizon to evaluate customer assistants before investing in production systems. More than a prototype, it established a repeatable framework for comparing AI models, testing prompt strategies and measuring customer experience, helping Verizon make evidence-based decisions as AI technology continues to evolve.