Deployment¶
The docker target supports Phase 1 only. Phase 2 and above require
--target aws-ec2. For first-time machine and credential preparation, start
with Getting started.
Target scope¶
| Target | Phases | Lifecycle | Inbound | Status |
|---|---|---|---|---|
docker |
1 only | Docker Compose | localhost ports | Runnable |
aws-ec2 |
1 – 3 | Terraform → EC2 → Compose | 80/443 only | Running (--profile phase-3) |
k8s |
— | Helm | — | Planned and disabled |
docker is the right choice for iterating on Phase 1 without needing an AWS
account. aws-ec2 is the primary target for demos and all phases beyond Phase 1.
Docker (Phase 1)¶
Start the Phase 1 profile:
Useful operations:
./scripts/stack.sh urls
./scripts/stack.sh logs
./scripts/stack.sh config
./scripts/stack.sh models
./scripts/stack.sh down
status checks only the active profile. urls prints endpoints without making
requests.
Published ports¶
| Service | Default |
|---|---|
| Langfuse | 3000 |
| LibreChat | 3080 |
| LiteLLM | 4000 |
| Feedback sidecar | 8080 |
| MinIO console | 9001 |
| MinIO API | 9002 |
Ports can be remapped for one invocation:
Postgres, Redis, MongoDB, and the Langfuse worker are not published to the host. Trace analytics live in ClickHouse Cloud, so there is no local ClickHouse port.
Data and teardown¶
./scripts/stack.sh down # stop containers, keep volumes
./scripts/stack.sh down --purge # also delete volumes
--purge permanently removes local stack data. Use it only for a disposable
environment.
Service credentials initialized into a persistent database or object-store
volume do not rotate merely because .env changes. Rotate the account inside
the service or recreate disposable volumes.
AWS EC2¶
The primary deployment target from Phase 2 onwards, and where all three phases
run today. A single EC2 instance running the same Docker Compose stack. Region:
ap-northeast-2. Instance: t3.xlarge, 100 GiB gp3.
Services are accessible over HTTPS subdomains only (chat.<domain>,
langfuse.<domain>, litellm.<domain>, media.<domain>). The security group
publishes 80 and 443 and nothing else — configuring a domain is therefore part
of provisioning, not an optional extra.
Prerequisites¶
| Requirement | Check |
|---|---|
| AWS CLI v2 configured | aws sts get-caller-identity |
EC2 key pair in ap-northeast-2 |
AWS console → EC2 → Key Pairs |
| Terraform ≥ 1.5 | terraform version |
| Phase 1 credentials set locally | ./scripts/stack.sh secrets status --phase 1 |
1. Push credentials to SSM¶
The EC2 instance reads Phase 1 credentials from SSM Parameter Store at boot. Push them before provisioning:
This writes every set Phase 1 credential as an SSM SecureString parameter
under /sais/phase-1/. The EC2 IAM role (provisioned by Terraform) grants
read access to that prefix only — no other AWS resources are reachable.
Verify:
aws ssm get-parameters-by-path \
--region ap-northeast-2 \
--path /sais/phase-1 \
--query 'Parameters[*].Name'
2. Configure a domain¶
Not optional: without it there is no published route to any service, because the direct ports are closed.
a. Add DNS records. In your DNS provider, create A records pointing each
subdomain at the EC2 instance's public IP:
| Subdomain | Points to |
|---|---|
chat.<your-domain> |
EC2 public IP |
langfuse.<your-domain> |
EC2 public IP |
litellm.<your-domain> |
EC2 public IP |
media.<your-domain> |
EC2 public IP |
b. Set domain config interactively:
This prompts for DOMAIN_BASE (e.g. example.com) and DOMAIN_SSL_EMAIL
(Let's Encrypt contact) and writes them to .env.
c. Push the domain config to SSM:
The bootstrap script reads DOMAIN_BASE from SSM at first boot. If set, it
renders a Caddyfile and starts the Caddy reverse proxy automatically.
Caddy obtains a TLS certificate from Let's Encrypt without any manual steps.
Certificate lifecycle: Let's Encrypt certificates are valid for 90 days. Caddy renews them automatically (typically at 30 days remaining). As long as the instance is running and reachable on port 80/443, certificates stay current indefinitely.
3. Provision and bootstrap¶
Add --tf-var 'ssh_allowed_cidrs=["<your-ip>/32"]' if you want SSH open from
the start; otherwise port 22 stays closed and you open it per session (below).
This runs terraform apply, then waits up to 5 minutes for the bootstrap
script to complete. The bootstrap script:
1. Installs Docker, yq
2. Clones this repository
3. Pulls all credentials (including DOMAIN_BASE) from SSM
4. Starts the Phase 1 compose stack
5. If DOMAIN_BASE is set: renders Caddyfile and starts the Caddy proxy
Bootstrap deliberately stops at Phase 1 — it is the scope that needs no MCP or RunPod credentials. Move the instance to Phase 2 or 3 with one command once those are pushed to SSM:
4. Verify¶
If bootstrap is still running, inspect the log:
5. Tear down¶
Runs terraform destroy. Removes the EC2 instance and security group. SSM
parameters are not deleted automatically — clean them up separately:
aws ssm delete-parameters \
--region ap-northeast-2 \
--names $(aws ssm get-parameters-by-path \
--region ap-northeast-2 \
--path /sais/phase-1 \
--query 'Parameters[*].Name' \
--output text)
Credential rotation¶
Database credentials written into persistent volumes (Postgres, ClickHouse, MinIO, MongoDB) do not rotate when SSM parameters change. After changing a credential:
- Update the SSM value:
./scripts/stack.sh secrets push --target aws-ec2 - Either recreate the affected service volume, or run
down --purgeand reprovision from scratch.
Enabling HTTPS on a running instance¶
If the instance is already running and you want to add HTTPS:
# 1. Set domain config locally and push to SSM
./scripts/stack.sh secrets domain
./scripts/stack.sh secrets push --target aws-ec2
# 2. SSH into the instance and apply
./scripts/stack.sh ssh --target aws-ec2
On the instance:
cd /opt/llmops-in-a-box
# Pull latest code
git pull origin main
# Write DOMAIN_BASE and DOMAIN_SSL_EMAIL into .env
# (re-fetch from SSM, or set directly)
echo "DOMAIN_BASE=example.com" >> .env
echo "DOMAIN_SSL_EMAIL=you@example.com" >> .env
# Render Caddyfile and start proxy
./scripts/stack.sh render --target aws-ec2 --profile phase-1
docker compose --project-name sais --profile proxy -f docker/docker-compose.yml up -d
Approximate cost¶
| Resource | On-demand, ap-northeast-2 |
|---|---|
| t3.xlarge | ~$120 / month |
| 100 GiB gp3 EBS | ~$8 / month |
| Data transfer | usage-dependent |
Stop or terminate the instance when not in use. This is a demo stack, not a production service.
Published ports (EC2)¶
The security group opens two ports, and only two:
| Port | CIDR | Purpose |
|---|---|---|
80 |
0.0.0.0/0 |
HTTP — Caddy redirects to HTTPS |
443 |
0.0.0.0/0 |
HTTPS — Caddy terminates TLS for every service |
Every service is reached through its subdomain, not through a port:
| Service | URL | Container port (internal) |
|---|---|---|
| LibreChat | https://chat.<domain> |
3080 |
| Langfuse | https://langfuse.<domain> |
3000 |
| LiteLLM | https://litellm.<domain> |
4000 |
| MinIO (images) | https://media.<domain> |
9000 |
mcp-clickhouse |
— | 9100, no public route |
The application ports are deliberately not published. They served plain HTTP,
and 4000 fronts the gateway's admin API — there is no reason to expose either
when Caddy already terminates TLS for the same services. status, urls, and
smoke-test follow the HTTPS route whenever DOMAIN_BASE is set.
Without a domain there is no way in
The subdomains are the only inbound path. If DOMAIN_BASE is not
configured, provision with --tf-var 'ssh_allowed_cidrs=["<your-ip>/32"]'
and reach the services over an SSH tunnel, or add the port rules back
deliberately.
SSH access¶
There is no SSH ingress by default — ssh_allowed_cidrs is an empty list,
and Terraform rejects 0.0.0.0/0 for it. Application traffic never needs port
22. Open it for the session that needs it:
MYIP=$(curl -s https://checkip.amazonaws.com)
SG=$(cd terraform && terraform output -raw security_group_id 2>/dev/null || echo "<sg-id>")
aws ec2 authorize-security-group-ingress --group-id "$SG" --region ap-northeast-2 \
--ip-permissions "IpProtocol=tcp,FromPort=22,ToPort=22,IpRanges=[{CidrIp=$MYIP/32,Description=temp}]"
# ... work ...
aws ec2 revoke-security-group-ingress --group-id "$SG" --region ap-northeast-2 \
--ip-permissions "IpProtocol=tcp,FromPort=22,ToPort=22,IpRanges=[{CidrIp=$MYIP/32}]"
Or declare it for a provisioning run:
--tf-var 'ssh_allowed_cidrs=["1.2.3.4/32"]'.
Never edit the security group's description
AWS treats it as immutable, so Terraform can only change it by replacing
the security group — and with it the instance, destroying the root volume
that holds Langfuse, MongoDB, and MinIO data. Put explanations in comments
instead. For the same reason aws_instance.stack carries
lifecycle { ignore_changes = [ami] }: the AMI data source is
most_recent, so without it an unrelated apply would replace the
instance the next time Amazon publishes an AL2023 image.
Phase 2 — MCP tool layer¶
Phase 2 adds the mcp-clickhouse service and wires it into the LiteLLM
gateway. Requires --target aws-ec2. Port 9100 is internal to the Docker
network; no security group change is needed.
Prerequisites¶
Set Phase 2 credentials and push to SSM:
Start Phase 2¶
./scripts/stack.sh render --profile phase-2 --target aws-ec2
./scripts/stack.sh up --profile phase-2 --target aws-ec2
./scripts/stack.sh status --target aws-ec2
render --profile phase-2 generates docker/litellm_config.yaml with the
mcp_servers block pointing at http://mcp-clickhouse:9100/sse. The tools
layer starts under the tools compose profile alongside the gateway,
observability, and UI layers.
Verify the MCP endpoint¶
A healthy response lists the clickhouse server.
Published ports (Phase 2)¶
Port 9100 (mcp-clickhouse) is not published to the host. It is
accessible only within the Docker network by the LiteLLM container.
Phase 3 — GPU serving on RunPod¶
vLLM is an externally managed OpenAI-compatible endpoint, not a Compose service:
the RunPod Serverless endpoint is created in the RunPod console (Serverless → New
Endpoint → vLLM worker template) with Qwen/Qwen2.5-7B-Instruct, and the gateway
only needs to be told where it is.
./scripts/stack.sh secrets setup --phase 3 # VLLM_API_BASE, VLLM_API_KEY, RUNPOD_COST_PER_TOKEN
./scripts/stack.sh secrets push --target aws-ec2
./scripts/stack.sh up --profile phase-3 --target aws-ec2
Requirements:
VLLM_API_BASE—https://api.runpod.ai/v2/<endpoint_id>/openai/v1, ending in/v1VLLM_API_KEY— a RunPod API key- a GPU tier of RTX 3090 or better (~24 GiB VRAM for a 7B bfloat16 model)
Operational notes:
qwen-7bcarries a 600 s timeout because a serverless cold start can take minutes. Setmin_workers=1in the RunPod console to avoid it, at the cost of a permanently billed worker.qwen-7bfalls back toclaude-sonnet, so a stopped endpoint degrades to the commercial API instead of failing. Traces taggedfallback:trueand scoredrouting_accuracy=0are how that shows up.RUNPOD_COST_PER_TOKENsets the per-token cost used for Langfuse cost attribution. Pod-hour billing is not derived automatically — see the comment indocker/litellm_callbacks.pyfor the arithmetic.
Do not expose a vLLM endpoint without authentication; a public unauthenticated GPU endpoint can be used by anyone.
Client endpoint¶
Every model is called through LiteLLM. The base URL is the only thing that differs between targets:
| Target | base_url |
|---|---|
docker |
http://localhost:4000 |
aws-ec2 |
https://litellm.<domain> |
from openai import OpenAI
client = OpenAI(
base_url="https://litellm.example.com", # or http://localhost:4000
api_key="<LITELLM_MASTER_KEY>",
)
response = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Hello"}],
)
Changing provider or serving implementation changes the model alias and gateway configuration, not the client protocol.
Troubleshooting¶
yq not found
Install mikefarah/yq v4. Other programs named yq use incompatible
syntax.
A model appears in the picker but requests fail
The catalog is rendered independently of provider-key liveness. Run
./scripts/stack.sh doctor and verify the selected provider credential.
Traces are missing
LiteLLM reads Langfuse keys at startup. Validate the credentials, then run
./scripts/stack.sh up again to reload them.
Langfuse fails to start — ClickHouse connection error
Langfuse now connects to ClickHouse Cloud instead of a local container.
Verify that LANGFUSE_CLICKHOUSE_USER and LANGFUSE_CLICKHOUSE_PASSWORD
are set and that the user has GRANT ALL ON llmops.* on ClickHouse Cloud.
The llmops database must exist before the first boot:
The Kubernetes target refuses to start
Expected in the current repository. That target is a declaration of the intended interface, not a completed deployment artifact.
Langfuse shows no traces — \"Event type not accepted\" in the ingestion response
Langfuse v4.0.0-rc.2 defaults LANGFUSE_MIGRATION_V4_WRITE_MODE to
events_only, which rejects the SDK v2 trace-create events that LiteLLM
sends. The fix is applied in docker/docker-compose.yml:
LANGFUSE_MIGRATION_V4_WRITE_MODE: "dual".
Valid values are "legacy" | "dual" | "events_only" (not "disabled" or
"all" — those fail Zod validation and crash the server on startup).
"events_only"— default for fresh v4 installs; rejects SDK v2 events."legacy"— accepts SDK v2 events but crashes the worker whenLANGFUSE_MIGRATION_V4_NATIVE_OTEL_BEHAVIOUR=direct(the v4 fresh-install default) is also set."dual"— writes to both v3 and v4 paths; compatible with both the SDK v2 client (LiteLLM) and the fresh-install worker config. This is the correct value for this stack.
Verify the env var is live in the running container:
docker compose --project-name sais exec langfuse-web env | grep MIGRATION
# expected: LANGFUSE_MIGRATION_V4_WRITE_MODE=dual
Test the ingestion endpoint (include a timestamp field — v4 requires it):
curl -s -X POST http://localhost:3000/api/public/ingestion \
-H "Content-Type: application/json" \
-u "${LANGFUSE_PUBLIC_KEY}:${LANGFUSE_SECRET_KEY}" \
-d '{"batch":[{"id":"t1","timestamp":"2026-01-01T00:00:00Z","type":"trace-create","body":{"id":"test","name":"test","timestamp":"2026-01-01T00:00:00Z"}}]}'
A healthy response is {"successes":[{"id":"t1","status":201}],"errors":[]}.
Important: docker compose restart does not apply env var changes from
docker-compose.yml. Use docker compose up -d --no-deps langfuse-web langfuse-worker
to recreate the containers and pick up the new value.
Remove the override once LiteLLM upgrades its bundled Langfuse SDK from v2 to v3, which uses the new v4 ingestion path.
No fallback model group found for original model_group=auto
The language-routing callback rewrites auto to qwen-7b (English/CJK) or
claude-sonnet (Korean) before the request is dispatched.
When that model then fails, LiteLLM looks up the fallback for the
original model group (auto), not the rewritten one. Without an
entry for auto in the fallback list, no recovery occurs.
The fix is in stack.yaml: auto is listed under
layers.gateway.options.routing.fallbacks with claude-sonnet as its
target, and scripts/stack.sh includes auto when rendering the
LiteLLM fallback table.
If you see this error after editing stack.yaml, verify that auto
appears in the fallbacks list, then re-render:
mcp-clickhouse container exits immediately after start
Symptom: The mcp-clickhouse container stops with an error about an
unrecognised flag (--transport sse not supported).
Cause: The mcp-clickhouse package version in use may not support
--transport sse as a CLI flag. Use the mcp-proxy wrapper approach or
pin to a version that supports SSE transport.
Fix: Check the docker/mcp/Dockerfile entrypoint. Use the mcp-proxy
wrapper to expose the stdio-only server over SSE and pass environment
variables through to the subprocess:
Then rebuild: docker compose build mcp-clickhouse.
LiteLLM /mcp endpoint returns 404
Cause: The stack was not rendered with --profile phase-2, so
mcp_servers is absent from docker/litellm_config.yaml.
Fix:
Verify the block is present:
ClickHouse connection refused inside mcp-clickhouse
Symptom: The container starts but tool calls fail with a connection
error. Container logs show Connection refused or authentication failed.
Cause: CLICKHOUSE_HOST, CLICKHOUSE_USER, or
CLICKHOUSE_PASSWORD is missing or incorrect, or CLICKHOUSE_SECURE is
not set to true for ClickHouse Cloud.
Fix: Verify the values in the running container:
If any value is wrong, update credentials and restart:
Image generation returns an error
The UnifiedRouter callback calls Cloudflare Workers AI directly. Verify
that CF_API_TOKEN and CF_ACCOUNT_ID are present in the running container:
If any variable is missing, set it with
./scripts/stack.sh secrets setup --phase 1, then restart LiteLLM:
See Credentials — Image generation (free tier) for token acquisition steps.