AI.GuardAI.Guard

Technical architecture

AI.Guard is a programmable reverse proxy that speaks OpenAI- and Anthropic-compatible APIs upstream of any model. No changes to your application logic — you point the existing SDK at the gateway's base URL.

Request / response flow

1

Integrates via API: AI.Guard sits between the application and the internal or external model.

2

Every user request passes through AI.Guard first.

3

The system analyzes and filters it — masking confidential data, blocking anything dangerous or disallowed.

4

The request only reaches the LLM once it clears every check (or access is denied).

5

The LLM's response is checked too, before it reaches the user, unmasking data as needed.

6

Every step is recorded in the audit log.

AI.Guard — proxy modeA numbered request travels from the user through AI.Guard to an LLM, then returns through AI.Guard for inspection and unmasking. Access management and an inspected RAG knowledge base inform policy, and every step creates a signed audit record.AI.Guard — proxy modeNumbered path of one request. Solid = request, dashed = response.1 User / applicationunchanged client SDKAI.GuardgatewayLLMinternal and / or external2 original prompt3 masked prompt4 model response5 unmasked response or denial noticeInspection chain, in order• identity + access rights• policy resolution• DLP → reversible masking• prompt-injection scoring • response filtering + unmaskingAccess managementAD / IAM — who may see whataccess rightsKnowledge base / RAGcompany data, incl. confidentialRAG context, inspected6 Signed audit recordevery step → SIEMSample payload — the same request at three points2 user → gateway“Show the financial records of John Dawson”3 gateway → LLM“Show the financial records of [CLIENT_1]”5 gateway → user“Here are the records of John Dawson: …” or “You do not have access…”

The model itself may be grounded on company data, including confidential records — which is exactly why the gateway inspects both directions. Every step is recorded in an immutable audit log.

Hybrid inspection

Deterministic structural checks (template integrity, role boundaries, schema validation on tool calls) run alongside a fine-tuned classifier ensemble for prompt-injection and jailbreak scoring.

Streaming-aware

SSE and chunked streaming flow through end-to-end. Inspection runs incrementally on partial output; violations can interrupt mid-stream with a typed error your SDK already understands.

Signed audit

Every request emits a signed audit envelope: caller identity, policy version, scores, redaction set, upstream provider, latency. Exported to your SIEM as it happens.

Latency budget

Inspection runs in the same hop as the proxy. During pilots, non-streaming chat completions measured approximately 11 ms p50 added latency and less than 17 ms p95. Your numbers are validated on your own traffic.

StageBudgetNotes
mTLS edge + auth~1.2 msToken verification, per-tenant rate-limit check
Policy resolution~0.4 msCached compiled rules, scoped by app + user
Prompt inspection~4–7 msHeuristic + ML classifier ensemble, parallelized
DLP & tokenization~2–3 msRegex + NER, reversible token issuance
Upstream callpassthroughStreaming preserved end-to-end
Response inspection~3–5 msOutput DLP, toxicity, tool-arg validation
Audit emitasyncSigned, queued to SIEM exporter

Deployment topologies

Pick the topology that matches your data-handling posture. The capability surface is the same; what differs is where prompts and responses physically travel.

Managed SaaS

Multi-tenant gateway in MOAI Cloud, US/EU regions. Fastest to deploy, fully managed updates.

What leaves your perimeter

Prompts and responses traverse MOAI infrastructure for inspection, never persisted beyond audit metadata.

Regional cluster (single-tenant)

Isolated gateway cluster in MOAI Cloud or in your cloud account. Per-region pinning for data residency.

What leaves your perimeter

Prompts and responses inspected inside your tenant boundary. Only signed audit summaries flow out, if at all.

Sidecar in your VPC

Gateway runs alongside your apps inside your Kubernetes cluster. Outbound calls go from the sidecar to upstream LLMs.

What leaves your perimeter

Only the call you would have made anyway leaves the cluster — to the chosen LLM provider. Nothing leaves to MOAI.

Air-gapped on-prem

Full stack including the classifier model deployed inside your network. Self-hosted model required upstream.

What leaves your perimeter

Nothing leaves your perimeter. Update bundles are pulled in manually via your secure artifact channel.

Supported protocols

No changes to your application logic — you point the existing SDK at the gateway's base URL.

OpenAI Chat Completions
OpenAI Responses
Anthropic Messages
AWS Bedrock InvokeModel
Azure OpenAI
Google Vertex AI (generateContent)
Mistral
Self-hosted vLLM / TGI
Raw HTTP passthrough