Technical architecture
AI.Guard is a programmable reverse proxy that speaks OpenAI- and Anthropic-compatible APIs upstream of any model. No changes to your application logic — you point the existing SDK at the gateway's base URL.
Request / response flow
Integrates via API: AI.Guard sits between the application and the internal or external model.
Every user request passes through AI.Guard first.
The system analyzes and filters it — masking confidential data, blocking anything dangerous or disallowed.
The request only reaches the LLM once it clears every check (or access is denied).
The LLM's response is checked too, before it reaches the user, unmasking data as needed.
Every step is recorded in the audit log.
The model itself may be grounded on company data, including confidential records — which is exactly why the gateway inspects both directions. Every step is recorded in an immutable audit log.
Hybrid inspection
Deterministic structural checks (template integrity, role boundaries, schema validation on tool calls) run alongside a fine-tuned classifier ensemble for prompt-injection and jailbreak scoring.
Streaming-aware
SSE and chunked streaming flow through end-to-end. Inspection runs incrementally on partial output; violations can interrupt mid-stream with a typed error your SDK already understands.
Signed audit
Every request emits a signed audit envelope: caller identity, policy version, scores, redaction set, upstream provider, latency. Exported to your SIEM as it happens.
Latency budget
Inspection runs in the same hop as the proxy. During pilots, non-streaming chat completions measured approximately 11 ms p50 added latency and less than 17 ms p95. Your numbers are validated on your own traffic.
| Stage | Budget | Notes |
|---|---|---|
| mTLS edge + auth | ~1.2 ms | Token verification, per-tenant rate-limit check |
| Policy resolution | ~0.4 ms | Cached compiled rules, scoped by app + user |
| Prompt inspection | ~4–7 ms | Heuristic + ML classifier ensemble, parallelized |
| DLP & tokenization | ~2–3 ms | Regex + NER, reversible token issuance |
| Upstream call | passthrough | Streaming preserved end-to-end |
| Response inspection | ~3–5 ms | Output DLP, toxicity, tool-arg validation |
| Audit emit | async | Signed, queued to SIEM exporter |
Deployment topologies
Pick the topology that matches your data-handling posture. The capability surface is the same; what differs is where prompts and responses physically travel.
Managed SaaS
Multi-tenant gateway in MOAI Cloud, US/EU regions. Fastest to deploy, fully managed updates.
Prompts and responses traverse MOAI infrastructure for inspection, never persisted beyond audit metadata.
Regional cluster (single-tenant)
Isolated gateway cluster in MOAI Cloud or in your cloud account. Per-region pinning for data residency.
Prompts and responses inspected inside your tenant boundary. Only signed audit summaries flow out, if at all.
Sidecar in your VPC
Gateway runs alongside your apps inside your Kubernetes cluster. Outbound calls go from the sidecar to upstream LLMs.
Only the call you would have made anyway leaves the cluster — to the chosen LLM provider. Nothing leaves to MOAI.
Air-gapped on-prem
Full stack including the classifier model deployed inside your network. Self-hosted model required upstream.
Nothing leaves your perimeter. Update bundles are pulled in manually via your secure artifact channel.
Supported protocols
No changes to your application logic — you point the existing SDK at the gateway's base URL.

