Skip to content

Deployment

The docker target supports Phase 1 only. Phase 2 and above require --target aws-ec2. For first-time machine and credential preparation, start with Getting started.

Target scope

Target Phases Lifecycle Inbound Status
docker 1 only Docker Compose localhost ports Runnable
aws-ec2 1 – 3 Terraform → EC2 → Compose 80/443 only Running (--profile phase-3)
k8s — Helm — Planned and disabled

docker is the right choice for iterating on Phase 1 without needing an AWS account. aws-ec2 is the primary target for demos and all phases beyond Phase 1.

Docker (Phase 1)

Start the Phase 1 profile:

./scripts/stack.sh up --target docker --profile phase-1
./scripts/stack.sh status

Useful operations:

./scripts/stack.sh urls
./scripts/stack.sh logs
./scripts/stack.sh config
./scripts/stack.sh models
./scripts/stack.sh down

status checks only the active profile. urls prints endpoints without making requests.

Published ports

Service Default
Langfuse 3000
LibreChat 3080
LiteLLM 4000
Feedback sidecar 8080
MinIO console 9001
MinIO API 9002

Ports can be remapped for one invocation:

LANGFUSE_PORT=3100 MINIO_API_PORT=9012 ./scripts/stack.sh up

Postgres, Redis, MongoDB, and the Langfuse worker are not published to the host. Trace analytics live in ClickHouse Cloud, so there is no local ClickHouse port.

Data and teardown

./scripts/stack.sh down          # stop containers, keep volumes
./scripts/stack.sh down --purge  # also delete volumes

--purge permanently removes local stack data. Use it only for a disposable environment.

Service credentials initialized into a persistent database or object-store volume do not rotate merely because .env changes. Rotate the account inside the service or recreate disposable volumes.

AWS EC2

The primary deployment target from Phase 2 onwards, and where all three phases run today. A single EC2 instance running the same Docker Compose stack. Region: ap-northeast-2. Instance: t3.xlarge, 100 GiB gp3.

Services are accessible over HTTPS subdomains only (chat.<domain>, langfuse.<domain>, litellm.<domain>, media.<domain>). The security group publishes 80 and 443 and nothing else — configuring a domain is therefore part of provisioning, not an optional extra.

Prerequisites

Requirement Check
AWS CLI v2 configured aws sts get-caller-identity
EC2 key pair in ap-northeast-2 AWS console → EC2 → Key Pairs
Terraform ≥ 1.5 terraform version
Phase 1 credentials set locally ./scripts/stack.sh secrets status --phase 1

1. Push credentials to SSM

The EC2 instance reads Phase 1 credentials from SSM Parameter Store at boot. Push them before provisioning:

./scripts/stack.sh secrets push --target aws-ec2

This writes every set Phase 1 credential as an SSM SecureString parameter under /sais/phase-1/. The EC2 IAM role (provisioned by Terraform) grants read access to that prefix only — no other AWS resources are reachable.

Verify:

aws ssm get-parameters-by-path \
  --region ap-northeast-2 \
  --path /sais/phase-1 \
  --query 'Parameters[*].Name'

2. Configure a domain

Not optional: without it there is no published route to any service, because the direct ports are closed.

a. Add DNS records. In your DNS provider, create A records pointing each subdomain at the EC2 instance's public IP:

Subdomain Points to
chat.<your-domain> EC2 public IP
langfuse.<your-domain> EC2 public IP
litellm.<your-domain> EC2 public IP
media.<your-domain> EC2 public IP

b. Set domain config interactively:

./scripts/stack.sh secrets domain

This prompts for DOMAIN_BASE (e.g. example.com) and DOMAIN_SSL_EMAIL (Let's Encrypt contact) and writes them to .env.

c. Push the domain config to SSM:

./scripts/stack.sh secrets push --target aws-ec2

The bootstrap script reads DOMAIN_BASE from SSM at first boot. If set, it renders a Caddyfile and starts the Caddy reverse proxy automatically. Caddy obtains a TLS certificate from Let's Encrypt without any manual steps.

Certificate lifecycle: Let's Encrypt certificates are valid for 90 days. Caddy renews them automatically (typically at 30 days remaining). As long as the instance is running and reachable on port 80/443, certificates stay current indefinitely.

3. Provision and bootstrap

./scripts/stack.sh up --target aws-ec2 \
    --tf-var key_name=<your-key-pair-name>

Add --tf-var 'ssh_allowed_cidrs=["<your-ip>/32"]' if you want SSH open from the start; otherwise port 22 stays closed and you open it per session (below).

This runs terraform apply, then waits up to 5 minutes for the bootstrap script to complete. The bootstrap script: 1. Installs Docker, yq 2. Clones this repository 3. Pulls all credentials (including DOMAIN_BASE) from SSM 4. Starts the Phase 1 compose stack 5. If DOMAIN_BASE is set: renders Caddyfile and starts the Caddy proxy

Bootstrap deliberately stops at Phase 1 — it is the scope that needs no MCP or RunPod credentials. Move the instance to Phase 2 or 3 with one command once those are pushed to SSM:

./scripts/stack.sh up --profile phase-3 --target aws-ec2

4. Verify

./scripts/stack.sh status --target aws-ec2
./scripts/stack.sh urls --target aws-ec2

If bootstrap is still running, inspect the log:

./scripts/stack.sh ssh --target aws-ec2
# on the instance:
sudo tail -f /var/log/bootstrap-ec2.log

5. Tear down

./scripts/stack.sh down --target aws-ec2

Runs terraform destroy. Removes the EC2 instance and security group. SSM parameters are not deleted automatically — clean them up separately:

aws ssm delete-parameters \
  --region ap-northeast-2 \
  --names $(aws ssm get-parameters-by-path \
    --region ap-northeast-2 \
    --path /sais/phase-1 \
    --query 'Parameters[*].Name' \
    --output text)

Credential rotation

Database credentials written into persistent volumes (Postgres, ClickHouse, MinIO, MongoDB) do not rotate when SSM parameters change. After changing a credential:

  1. Update the SSM value: ./scripts/stack.sh secrets push --target aws-ec2
  2. Either recreate the affected service volume, or run down --purge and reprovision from scratch.

Enabling HTTPS on a running instance

If the instance is already running and you want to add HTTPS:

# 1. Set domain config locally and push to SSM
./scripts/stack.sh secrets domain
./scripts/stack.sh secrets push --target aws-ec2

# 2. SSH into the instance and apply
./scripts/stack.sh ssh --target aws-ec2

On the instance:

cd /opt/llmops-in-a-box

# Pull latest code
git pull origin main

# Write DOMAIN_BASE and DOMAIN_SSL_EMAIL into .env
# (re-fetch from SSM, or set directly)
echo "DOMAIN_BASE=example.com" >> .env
echo "DOMAIN_SSL_EMAIL=you@example.com" >> .env

# Render Caddyfile and start proxy
./scripts/stack.sh render --target aws-ec2 --profile phase-1
docker compose --project-name sais --profile proxy -f docker/docker-compose.yml up -d

Approximate cost

Resource On-demand, ap-northeast-2
t3.xlarge ~$120 / month
100 GiB gp3 EBS ~$8 / month
Data transfer usage-dependent

Stop or terminate the instance when not in use. This is a demo stack, not a production service.

Published ports (EC2)

The security group opens two ports, and only two:

Port CIDR Purpose
80 0.0.0.0/0 HTTP — Caddy redirects to HTTPS
443 0.0.0.0/0 HTTPS — Caddy terminates TLS for every service

Every service is reached through its subdomain, not through a port:

Service URL Container port (internal)
LibreChat https://chat.<domain> 3080
Langfuse https://langfuse.<domain> 3000
LiteLLM https://litellm.<domain> 4000
MinIO (images) https://media.<domain> 9000
mcp-clickhouse — 9100, no public route

The application ports are deliberately not published. They served plain HTTP, and 4000 fronts the gateway's admin API — there is no reason to expose either when Caddy already terminates TLS for the same services. status, urls, and smoke-test follow the HTTPS route whenever DOMAIN_BASE is set.

Without a domain there is no way in

The subdomains are the only inbound path. If DOMAIN_BASE is not configured, provision with --tf-var 'ssh_allowed_cidrs=["<your-ip>/32"]' and reach the services over an SSH tunnel, or add the port rules back deliberately.

SSH access

There is no SSH ingress by default — ssh_allowed_cidrs is an empty list, and Terraform rejects 0.0.0.0/0 for it. Application traffic never needs port 22. Open it for the session that needs it:

MYIP=$(curl -s https://checkip.amazonaws.com)
SG=$(cd terraform && terraform output -raw security_group_id 2>/dev/null || echo "<sg-id>")

aws ec2 authorize-security-group-ingress --group-id "$SG" --region ap-northeast-2 \
  --ip-permissions "IpProtocol=tcp,FromPort=22,ToPort=22,IpRanges=[{CidrIp=$MYIP/32,Description=temp}]"

# ... work ...

aws ec2 revoke-security-group-ingress --group-id "$SG" --region ap-northeast-2 \
  --ip-permissions "IpProtocol=tcp,FromPort=22,ToPort=22,IpRanges=[{CidrIp=$MYIP/32}]"

Or declare it for a provisioning run: --tf-var 'ssh_allowed_cidrs=["1.2.3.4/32"]'.

Never edit the security group's description

AWS treats it as immutable, so Terraform can only change it by replacing the security group — and with it the instance, destroying the root volume that holds Langfuse, MongoDB, and MinIO data. Put explanations in comments instead. For the same reason aws_instance.stack carries lifecycle { ignore_changes = [ami] }: the AMI data source is most_recent, so without it an unrelated apply would replace the instance the next time Amazon publishes an AL2023 image.

Phase 2 — MCP tool layer

Phase 2 adds the mcp-clickhouse service and wires it into the LiteLLM gateway. Requires --target aws-ec2. Port 9100 is internal to the Docker network; no security group change is needed.

Prerequisites

Set Phase 2 credentials and push to SSM:

./scripts/stack.sh secrets setup --phase 2
./scripts/stack.sh secrets push --target aws-ec2

Start Phase 2

./scripts/stack.sh render --profile phase-2 --target aws-ec2
./scripts/stack.sh up --profile phase-2 --target aws-ec2
./scripts/stack.sh status --target aws-ec2

render --profile phase-2 generates docker/litellm_config.yaml with the mcp_servers block pointing at http://mcp-clickhouse:9100/sse. The tools layer starts under the tools compose profile alongside the gateway, observability, and UI layers.

Verify the MCP endpoint

curl https://litellm.<domain>/mcp

A healthy response lists the clickhouse server.

Published ports (Phase 2)

Port 9100 (mcp-clickhouse) is not published to the host. It is accessible only within the Docker network by the LiteLLM container.


Phase 3 — GPU serving on RunPod

vLLM is an externally managed OpenAI-compatible endpoint, not a Compose service: the RunPod Serverless endpoint is created in the RunPod console (Serverless → New Endpoint → vLLM worker template) with Qwen/Qwen2.5-7B-Instruct, and the gateway only needs to be told where it is.

./scripts/stack.sh secrets setup --phase 3   # VLLM_API_BASE, VLLM_API_KEY, RUNPOD_COST_PER_TOKEN
./scripts/stack.sh secrets push --target aws-ec2
./scripts/stack.sh up --profile phase-3 --target aws-ec2

Requirements:

  • VLLM_API_BASE — https://api.runpod.ai/v2/<endpoint_id>/openai/v1, ending in /v1
  • VLLM_API_KEY — a RunPod API key
  • a GPU tier of RTX 3090 or better (~24 GiB VRAM for a 7B bfloat16 model)

Operational notes:

  • qwen-7b carries a 600 s timeout because a serverless cold start can take minutes. Set min_workers=1 in the RunPod console to avoid it, at the cost of a permanently billed worker.
  • qwen-7b falls back to claude-sonnet, so a stopped endpoint degrades to the commercial API instead of failing. Traces tagged fallback:true and scored routing_accuracy=0 are how that shows up.
  • RUNPOD_COST_PER_TOKEN sets the per-token cost used for Langfuse cost attribution. Pod-hour billing is not derived automatically — see the comment in docker/litellm_callbacks.py for the arithmetic.

Do not expose a vLLM endpoint without authentication; a public unauthenticated GPU endpoint can be used by anyone.

Client endpoint

Every model is called through LiteLLM. The base URL is the only thing that differs between targets:

Target base_url
docker http://localhost:4000
aws-ec2 https://litellm.<domain>
from openai import OpenAI

client = OpenAI(
    base_url="https://litellm.example.com",   # or http://localhost:4000
    api_key="<LITELLM_MASTER_KEY>",
)

response = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Hello"}],
)

Changing provider or serving implementation changes the model alias and gateway configuration, not the client protocol.

Troubleshooting

yq not found

Install mikefarah/yq v4. Other programs named yq use incompatible syntax.

A model appears in the picker but requests fail

The catalog is rendered independently of provider-key liveness. Run ./scripts/stack.sh doctor and verify the selected provider credential.

Traces are missing

LiteLLM reads Langfuse keys at startup. Validate the credentials, then run ./scripts/stack.sh up again to reload them.

Langfuse fails to start — ClickHouse connection error

Langfuse now connects to ClickHouse Cloud instead of a local container. Verify that LANGFUSE_CLICKHOUSE_USER and LANGFUSE_CLICKHOUSE_PASSWORD are set and that the user has GRANT ALL ON llmops.* on ClickHouse Cloud. The llmops database must exist before the first boot:

CREATE DATABASE IF NOT EXISTS llmops;
CREATE USER langfuse_writer IDENTIFIED BY '<password>';
GRANT ALL ON llmops.* TO langfuse_writer;
The Kubernetes target refuses to start

Expected in the current repository. That target is a declaration of the intended interface, not a completed deployment artifact.

Langfuse shows no traces — \"Event type not accepted\" in the ingestion response

Langfuse v4.0.0-rc.2 defaults LANGFUSE_MIGRATION_V4_WRITE_MODE to events_only, which rejects the SDK v2 trace-create events that LiteLLM sends. The fix is applied in docker/docker-compose.yml: LANGFUSE_MIGRATION_V4_WRITE_MODE: "dual".

Valid values are "legacy" | "dual" | "events_only" (not "disabled" or "all" — those fail Zod validation and crash the server on startup).

  • "events_only" — default for fresh v4 installs; rejects SDK v2 events.
  • "legacy" — accepts SDK v2 events but crashes the worker when LANGFUSE_MIGRATION_V4_NATIVE_OTEL_BEHAVIOUR=direct (the v4 fresh-install default) is also set.
  • "dual" — writes to both v3 and v4 paths; compatible with both the SDK v2 client (LiteLLM) and the fresh-install worker config. This is the correct value for this stack.

Verify the env var is live in the running container:

docker compose --project-name sais exec langfuse-web env | grep MIGRATION
# expected: LANGFUSE_MIGRATION_V4_WRITE_MODE=dual

Test the ingestion endpoint (include a timestamp field — v4 requires it):

curl -s -X POST http://localhost:3000/api/public/ingestion \
  -H "Content-Type: application/json" \
  -u "${LANGFUSE_PUBLIC_KEY}:${LANGFUSE_SECRET_KEY}" \
  -d '{"batch":[{"id":"t1","timestamp":"2026-01-01T00:00:00Z","type":"trace-create","body":{"id":"test","name":"test","timestamp":"2026-01-01T00:00:00Z"}}]}'

A healthy response is {"successes":[{"id":"t1","status":201}],"errors":[]}.

Important: docker compose restart does not apply env var changes from docker-compose.yml. Use docker compose up -d --no-deps langfuse-web langfuse-worker to recreate the containers and pick up the new value.

Remove the override once LiteLLM upgrades its bundled Langfuse SDK from v2 to v3, which uses the new v4 ingestion path.

No fallback model group found for original model_group=auto

The language-routing callback rewrites auto to qwen-7b (English/CJK) or claude-sonnet (Korean) before the request is dispatched. When that model then fails, LiteLLM looks up the fallback for the original model group (auto), not the rewritten one. Without an entry for auto in the fallback list, no recovery occurs.

The fix is in stack.yaml: auto is listed under layers.gateway.options.routing.fallbacks with claude-sonnet as its target, and scripts/stack.sh includes auto when rendering the LiteLLM fallback table.

If you see this error after editing stack.yaml, verify that auto appears in the fallbacks list, then re-render:

./scripts/stack.sh render
mcp-clickhouse container exits immediately after start

Symptom: The mcp-clickhouse container stops with an error about an unrecognised flag (--transport sse not supported).

Cause: The mcp-clickhouse package version in use may not support --transport sse as a CLI flag. Use the mcp-proxy wrapper approach or pin to a version that supports SSE transport.

Fix: Check the docker/mcp/Dockerfile entrypoint. Use the mcp-proxy wrapper to expose the stdio-only server over SSE and pass environment variables through to the subprocess:

CMD ["mcp-proxy", "--port", "9100", "--host", "0.0.0.0", \
     "--pass-environment", "mcp-clickhouse"]

Then rebuild: docker compose build mcp-clickhouse.

LiteLLM /mcp endpoint returns 404

Cause: The stack was not rendered with --profile phase-2, so mcp_servers is absent from docker/litellm_config.yaml.

Fix:

./scripts/stack.sh render --profile phase-2
docker compose up -d --no-deps litellm

Verify the block is present:

grep -A5 mcp_servers docker/litellm_config.yaml
ClickHouse connection refused inside mcp-clickhouse

Symptom: The container starts but tool calls fail with a connection error. Container logs show Connection refused or authentication failed.

Cause: CLICKHOUSE_HOST, CLICKHOUSE_USER, or CLICKHOUSE_PASSWORD is missing or incorrect, or CLICKHOUSE_SECURE is not set to true for ClickHouse Cloud.

Fix: Verify the values in the running container:

docker compose exec mcp-clickhouse env | grep CLICKHOUSE

If any value is wrong, update credentials and restart:

./scripts/stack.sh secrets setup --phase 2
./scripts/stack.sh secrets write
docker compose up -d --no-deps mcp-clickhouse
Image generation returns an error

The UnifiedRouter callback calls Cloudflare Workers AI directly. Verify that CF_API_TOKEN and CF_ACCOUNT_ID are present in the running container:

docker compose exec litellm env | grep CF_

If any variable is missing, set it with ./scripts/stack.sh secrets setup --phase 1, then restart LiteLLM:

./scripts/stack.sh secrets write
docker compose up -d --no-deps litellm

See Credentials — Image generation (free tier) for token acquisition steps.