Under NDAAI SafetyComplianceContent SafetyTrust

Making enterprise AI safe to adopt

Making enterprise AI safe to adopt

A configurable guardrails framework, benchmarked for accuracy and built to self-serve, that lets enterprise clients run generative AI without gambling on unsafe output.

The problem

As AI+ Studio scaled past 1,400 enterprise clients running GenAI use cases, customer care bots, AI agents, social workflows, across three model providers, the next frontier was safety at the platform level: content moderation and prompt injection defence built in by design. This wasn't a niche ask. It was one of the most consistent requests from enterprise clients, and a defining shift happening across the AI industry at the same time. The mandate: define and build an entire guardrails framework, from taxonomy to architecture, from scratch.

The approach

Research first

Studied how other platforms handled content safety before building our own approach.

Right method for the job

Simple rules for clear cases, machine learning only where judgment was truly needed.

Move fast, build right

Launched with an existing model while a custom one was trained in parallel.

Design ownership

Also designed how clients configure and review every guardrail.

Tools & stack

FigmaJIRAConfluenceHuggingFace modelsPrompt GuardBenchmarking framework

From pilot to GA

  1. Stage 01

    Framework & first build

    Defined the five-guardrail taxonomy and shipped v1.

  2. Stage 02

    Early access

    Rolled out to the enterprise partners who'd specifically asked for it.

  3. Stage 03

    Evaluate & iterate

    Ran it over a full release, tuning against real usage.

  4. Stage 04

    General availability

    Took it to GA for all 1,400+ clients, across every model provider.

Guardrails in action

Setting up guardrails

Blocking malicious content in real time

Impact

1,400+

enterprise clients now on the framework, GA across Azure OpenAI, Google Vertex, and Amazon Bedrock

92-93%

benchmarked accuracy, measured independently per guardrail category

~55ms

added latency end to end, safety with no meaningful performance cost

5

guardrail categories, each independently configurable: self-serve for control, ready to use out of the box

Under NDA

This overview stays high-level by design. The architecture and trade-offs underneath sit under NDA, I'm happy to walk you through them directly.