AI.GuardAI.Guard

Enforcement catalog

Each inspection layer is independently configurable. Below is what they do — and a live mock of inline redaction in flight.

Inline Inspection — /v1/chat/completions
received
Original — as received
“Summarize this support ticket from Jane Roe (jane.roe@example.com, SSN 000-00-0000) and draft a reply.”
Original values remain visible only inside the protected boundary.
Masked — forwarded upstream
“Summarize this support ticket from [PII_NAME_001] ([PII_EMAIL_001], SSN [PII_SSN_001]) and draft a reply.”
3 reversible tokens · 8.2 ms
Unmasked — returned to caller
“Ticket from Jane Roe (jane.roe@example.com) — suggested reply: …”
Response tokens are rehydrated before the application receives the answer.
policy: pii-strict-v3test data: example.com · never-issued SSNreversible: yes

Prompt-injection & jailbreak defense

  • Classifier ensembleMultiple fine-tuned classifiers vote on injection score; ensemble reduces single-model evasion.
  • Full attack taxonomyMulti-turn conversations, impersonation, code words and obfuscation, instructions embedded in data, formatting tricks (markdown, HTML), and injection through structured fields.
  • Indirect injection coverageRAG context is analyzed; injected instructions inside retrieved documents are scored separately from user intent.
  • Multilingual coverageDetection trained on attacks in 18 languages including low-resource ones used to bypass single-language classifiers.

DLP & reversible tokenization

  • Detailed masking categoriesPII (names, addresses, IDs), Financial (cards, transactions, debt), Medical (diagnoses, prescriptions), Business (salaries, supplier data), and Documents (SSN, real-estate).
  • Reversible tokensSensitive values are replaced with deterministic tokens upstream. The original value never leaves your perimeter.
  • Loss-less rehydrationTokens are reversed in the response on the way back so the application sees the original data, transparently.
  • Flexible policiesInstant on/off rules, client-specific policies per team/environment, and Natural-Language detection rules (no query language required).

Output filtering & guardrails

  • Toxicity, bias, regulated contentConfigurable thresholds per class. Defaults aligned to OWASP LLM Top 10 and NIST AI RMF categories.
  • Hallucination guardrails for toolsTool / function-call arguments are validated against your declared JSON schema before being executed.
  • Citation enforcementFor RAG applications, responses can be required to cite from the retrieved set; uncited claims are flagged.
  • Outbound exfil detectionEncoded data, suspicious markdown image beacons, and base64-blobs in outputs are detected and stripped.

Policy-as-code

  • Versioned, reviewed policiesPolicies live in your repo as YAML/Rego. Changes go through PR review; old versions remain queryable in audit.
  • Per-app, per-user scopingBind policies to apps, environments, user roles, or upstream provider. Hot-reload propagates in seconds.
  • Shadow modeTest new rules against live traffic without enforcing. Observe hit-rate and false-positive cost before promotion.
  • Exception workflowJustified bypass requests are logged with caller identity and routed to the AppSec owner for retroactive review.

Quotas & rate limiting

  • Per-tenant, per-app, per-userIndependent token, request, and cost budgets with sliding windows.
  • Model-aware accountingDifferent budgets per upstream model — cap expensive frontier models per role.
  • Burst & smoothingToken-bucket with configurable burst. Friendly 429s with retry-after honored downstream.
  • Denial-of-wallet protectionAnomaly detection on call patterns blocks cost-abuse loops before they exhaust budgets.

Content provenance

  • Watermark detectionWhere upstream providers emit content watermarks, the gateway can require them on response and flag missing signals.
  • Outbound provenanceOptional C2PA-style provenance signing on application-generated content for downstream verification.
  • Model & policy attestationEvery response carries a signed header noting model id, policy version, and gateway build for downstream attribution.
  • Drift signalingUpstream model identifier changes are surfaced; pin to a specific version and get notified on silent rollouts.