DEPLOYMENT & INTEGRATIONS

Fit the boundary to your environment.

Plan how AI traffic reaches Milgram, how the platform runs, and who owns the operational controls around it.

CONNECT THE TRAFFIC PATH

Use your AI clients.
Keep your provider choice.

Milgram sits between a supported client workflow and the model provider. Integration begins with routing that workflow through the proxy.

Milgram’s provider handlers include OpenAI Chat Completions and Responses, Anthropic, Google, OpenRouter, and Moonshot. This describes implemented API paths; it is not a partnership claim or a promise that every provider feature is covered.

Client configuration, authentication, streaming, tool-call formats, and modality support can differ. Validate the specific API endpoint and client version you intend to use before expanding the rollout.

Investigation uses a separate MCP connection. An authorized client connects to your deployment’s MCP endpoint, completes the supported OAuth flow, and receives tools constrained by token scopes and user capabilities.

Understand MCP permissions

CUSTOMER-MANAGED DEPLOYMENT

Choose the runtime your team can operate.

PathBest suited toOperational considerations
Docker ComposeSingle-machine evaluations and compact deployments.Runs the application services with bundled PostgreSQL and Redis, or uses managed services. Plan persistent storage, backup, upgrades, and host monitoring.
Kubernetes / HelmTeams with an existing Kubernetes operating model.Uses signed, digest-pinned releases and runtime Secrets. CPU is the default; GPU workers require the appropriate NVIDIA infrastructure.
Air-gapped deliveryEnvironments requiring offline installation and controlled artifact transfer.Signed bundles include runtime images and model artifacts. Verify signatures and checksums before import, and plan an update process.
Offline detection and model inference are separate. Offline-complete detector artifacts remove runtime model downloads for detection. Requests to an external AI provider still need an approved network path.

Access is currently invite-only beta. Confirm the available release, integration scope, and delivery arrangements with Milgram before selecting a deployment.

A CONTROLLED ONBOARDING SEQUENCE

Begin narrow.
Expand with evidence.

  1. Map the boundary

    Identify AI clients, model endpoints, authentication, data classes, and tool workflows. Assign an owner for policy and an owner for operations.

  2. Validate compatibility

    Test representative requests, streaming responses, tool loops, error responses, and timeouts in an isolated evaluation environment.

  3. Establish a shadow baseline

    Confirm session completeness, review detections, and measure false positives before enabling policies that change or stop traffic.

  4. Introduce controls and optimization

    Enable selected masking, blocking, or compression mechanisms. Check business-task outcomes and practice rollback before increasing traffic.

Plan the operating responsibilities.

Capacity and availability

Size the database, worker, ingress, and storage against representative traffic. The current Helm design uses one stateful worker replica with read-write-once volumes for compiled engines and model promotion.

Backend and dashboard scaling should be evaluated alongside shared dependencies. Establish throughput and recovery evidence for your topology rather than assuming a generic availability or latency figure.

Updates and recovery

Customer releases use immutable image digests. Release procedures include signed artifacts, vulnerability scanning, and software bills of materials.

Agree on backup and restore procedures, secrets management, ownership of release verification, and rollback of an application or detection update. Request delivery-specific evidence during technical evaluation.

Is a self-service hosted plan available?

Milgram currently offers invite-only beta access. Contact the team to discuss a hosted evaluation or customer-managed deployment and confirm availability, operating boundaries, and commercial terms.

Is this a drop-in integration for every AI application?

No universal compatibility claim is made. Many provider-shaped workflows can use a proxy configuration, but managed applications, subscription clients, endpoint versions, and modalities can impose constraints. Bring the exact workflow to evaluation.

INVITE-ONLY BETA

Map your deployment with Milgram.

Bring your clients, provider endpoints, and operating requirements to a technical evaluation.

Talk to Milgram