When only engineering is familiar with operating AI systems, other disciplines don't feel comfortable contributing. Teams knew quality mattered, but didn't have a shared place to define it or a clear way to act on it. So the work defaulted to whoever was closest to the system.

The fix wasn't better engineering. It was designing a structure where the right expertise had a clear way into the process. That’s what led to a framework that gives every discipline a defined role in AI quality:

  • The Success Metric: Business requirements define the “what” (the standards). - The Testing Narrative: Quality assurance defines the “who, when and where” (the context). - The Enablers: Engineering defines the “how” (the tools).

This structure doesn't redistribute burden, it brings in expertise that was always needed but had no way to contribute.

We put this framework into practice evaluating NBCU's OLI, an AI assistant built to help U.S. viewers navigate the thousands of hours of competition of the 2026 Milan Cortina Winter Games.

For OLI, we built a custom service named Osiris on top of Google’s Vertex AI Evaluation API. Hosted on Google Cloud Platform, Osiris made the API's powerful features accessible to the entire team.

Here is how the team contributed to the framework.

1. The “What”: Metrics for measurement.

This is an exercise that leverages Product’s insights of the feature’s purpose and formats them in a way an AI can be evaluated programmatically.

The system we built makes this shift explicit: Product translates business requirements directly into evaluation metrics. No guesswork or interpretation gap.

A Metric Manager was created to allow anyone to take their business requirements  applying them to AI evaluations. Here’s how we structured it for OLI:

  • Criteria: What specific behavior are we measuring? For OLI, this meant rules like:
  • “the response speaks to dates and times naturally” - “the chatbot does not compare athletes” - Metrics: Group-related criteria that allows to see patterns in performance. OLI’s metrics included categories that allowed the team to say, for instance, “OLI is strong on scheduling equality, but weak on athlete disambiguation.” Examples of these categories included: