Configuration¶
One declarative file describes what the stack is. The script resolves the
selected target and profile. Both the docker and aws-ec2 targets are
currently runnable; k8s remains a declaration of the intended interface.
flowchart TB
SY["<b>stack.yaml</b><br/><small>layers · models · phases<br/>targets · secret names</small>"]
SH["<b>scripts/stack.sh</b>"]
L["resolve layers<br/><small>→ compose profiles</small>"]
R["render model catalog"]
T["pick target"]
LL["docker/litellm_config.yaml"]
LR["docker/librechat.yaml"]
D["docker<br/><small>compose</small>"]
A["aws-ec2<br/><small>terraform</small>"]
CF["docker/caddy/Caddyfile"]
SY --> SH
SH --> L
SH --> R
SH --> T
R --> LL
R --> LR
R --> CF
T --> D
T --> A
classDef src fill:#0969da,stroke:#0969da,color:#fff
classDef gen fill:#bf8700,stroke:#bf8700,color:#fff
class SY,SH src
class LL,LR,CF gen
stack.yaml¶
The single source of truth: layers and their implementations, the model catalog, build-out phases, deployment targets, and the names of every secret.
It never contains a secret value, and it never names a specific host. That is deliberate — the same file describes a laptop deployment and a cloud one.
schema: 1
project:
name: llmops-in-a-box
environment: demo
defaults:
target: docker
profile: phase-1
Layers¶
layers:
gateway:
phase: 1
enabled: true
impl: litellm
managed_by: compose
compose_profile: gateway
port: 4000
requires: []
options:
master_key_env: LITELLM_MASTER_KEY
callbacks:
success: []
failure: []
| Key | Meaning |
|---|---|
phase |
Which build-out phase introduces this layer |
enabled |
false = declared but not wired up in this repo yet |
impl |
The swappable implementation (litellm, vllm, minio, …) |
managed_by |
compose · runpod · external — who owns the lifecycle |
compose_profile |
Compose profile that starts it, or null if not a compose service |
requires |
Other layers this one cannot run without — resolved transitively |
managed_by matters
The serving layer is managed_by: runpod — it is not a compose service. stack.sh won't try to start it; it validates that VLLM_API_BASE points somewhere live instead.
Models¶
The model catalog is the one place models are defined:
models:
- alias: claude-sonnet
phase: 1
enabled: true
tier: commercial
provider: anthropic
litellm_model: anthropic/claude-sonnet-4-5
api_key_env: ANTHROPIC_API_KEY
context_window: 200000
cost_per_1k: { input: 0.003, output: 0.015 }
rate_limit: { rpm: 500, tpm: 400000 }
ui:
label: Claude Sonnet 4.5
description: Anthropic commercial API, format-translated by LiteLLM.
Costs are authored per 1,000 tokens because that is how provider pricing pages read. The renderer converts to LiteLLM's per-token fields, so Langfuse cost attribution stays correct without anyone counting zeros by hand.
scripts/stack.sh¶
Resolves the config for a chosen --target and --profile, then renders the model catalog into the two configs that would otherwise duplicate it.
Why render at all
A model has to be declared twice in a naive setup — once in LiteLLM's model_list so the gateway can route it, and once in LibreChat's models.default so the UI offers it. Those two lists drift, and the failure mode is a model in the picker that 400s on use. Rendering both from one catalog removes the class of bug.
docker-compose.yml stays hand-written and greppable — layer selection maps onto native compose profiles rather than generating YAML. Only the genuinely duplicated part is generated.
Commands¶
| Command | Purpose |
|---|---|
doctor |
Preflight: tooling, secrets, layers, models |
secrets setup |
Technology-grouped credential menu: input, defaults, generation, .env write |
secrets domain |
Set DOMAIN_BASE and DOMAIN_SSL_EMAIL for the aws-ec2 HTTPS proxy |
secrets status \| validate |
Report presence and check formats without printing values |
secrets write \| generate \| audit |
Automation primitives and leak checks |
secrets push |
Push .env values to SSM Parameter Store (aws-ec2 target) |
phases |
Build-out phases, current status, and the next steps beyond them |
config |
Resolved stack for the selected target/profile |
models |
Model table with per-1k costs |
render |
Regenerate litellm_config.yaml, librechat.yaml, and (aws-ec2) Caddyfile |
up |
Render, then deploy to the selected target |
status |
Curl every health check for active layers |
smoke-test |
Send one request end to end and verify the trace reached Langfuse |
urls |
Print the endpoint list |
logs |
Follow compose logs |
down |
Tear down (--purge also drops volumes) |
ssh |
SSH into the EC2 instance (aws-ec2 target only) |
Flags¶
| Flag | Effect |
|---|---|
-t, --target <name> |
Deployment target — default from defaults.target |
-p, --profile <name> |
Stack profile — default from defaults.profile |
--tf-var k=v |
Extra Terraform variable (repeatable) |
--all |
doctor: check every phase's secrets, not just active |
--no-render |
Skip config rendering on up |
--purge |
down: also delete volumes — destructive |
-n, --dry-run |
Print commands instead of running them |
-f, --file <path> |
Alternate stack.yaml |
bash 3.2
stack.sh targets bash 3.2 so macOS system bash works with no upgrade — no associative arrays, no readarray, no ${var^^}. Its only hard dependency is mikefarah/yq v4 (brew install yq).
Inspecting the resolved stack¶
$ ./scripts/stack.sh config --target aws-ec2 --profile phase-3
project llmops-in-a-box (demo)
target aws-ec2 — Primary deployment target for Phase 2 and above. …
profile phase-3 — Phase 3 — Phase 2 plus GPU serving on RunPod
layers
gateway litellm compose gateway
observability langfuse compose obs
ui librechat compose ui
tools mcp compose tools
serving vllm runpod
compose profiles gateway obs ui tools
models qwen-7b claude-sonnet
serving has no compose profile because RunPod owns its lifecycle — the row is
there to show it resolved, not to start a container.
$ ./scripts/stack.sh models --profile phase-3
ALIAS TIER LITELLM_MODEL IN/1k OUT/1k
qwen-7b self_hosted openai/Qwen/Qwen2.5-7B-Instruct 0.0 0.0
claude-sonnet commercial anthropic/claude-sonnet-4-5 0.003 0.015
Adding a model¶
Add one entry to models: in stack.yaml, then re-render. The gateway and the UI picker update together:
- alias: gpt-4o-mini
phase: 1
enabled: true
tier: commercial
provider: openai
litellm_model: openai/gpt-4o-mini
api_key_env: OPENAI_API_KEY
context_window: 128000
cost_per_1k: { input: 0.00015, output: 0.0006 }
ui:
label: GPT-4o mini
description: Cheap, fast baseline for cost comparisons.
Routing overview¶
Phase 1 has two routing paths. Both go through LiteLLM and produce Langfuse traces.
flowchart TD
LC["LibreChat"]
G["LiteLLM Gateway\n(UnifiedRouter callback)"]
LF["Langfuse"]
QW["qwen-7b\n(RunPod)"]
CS["claude-sonnet\n(Anthropic)"]
CF["Cloudflare Workers AI\nFLUX.1-schnell"]
MN["MinIO\n(media.<domain>)"]
LC -- "chat · model=auto" --> G
G -- "image keywords detected" --> CF
CF -- "store" --> MN
G -- "English / CJK" --> QW
G -- "Korean (Hangul)" --> CS
G -- "traces (completion calls only)" --> LF
The chat path and the image path share the same gateway endpoint and the same Langfuse project. No client-side changes are needed to switch providers.
Language routing¶
LiteLLM routes chat requests automatically by the dominant script of the last user message — no client changes required.
| Detected script | Target model |
|---|---|
| Hangul (Korean) | claude-sonnet |
| CJK (Chinese, Japanese) | qwen-7b |
| Latin (English, etc.) | qwen-7b |
Detection is a pure Unicode heuristic over the last user message, checked in priority order: Hangul first (syllables, Jamo, compatibility Jamo), then CJK ideographs and kana, then Latin as the default. A script wins when more than 15 % of the message's characters belong to it, so a single Korean word in an otherwise-English sentence does not flip the route. No extra network call; under 1 ms added to p99.
The router only rewrites the model field when the client sends "auto" or an empty string. Any other value (e.g. "claude-sonnet") is treated as an explicit model choice and left untouched. Non-chat requests (call_type not in {"completion", "acompletion"}) bypass the callback entirely.
Fallback: qwen-7b → claude-sonnet. When the RunPod pod is cold or unavailable, LiteLLM falls back to claude-sonnet automatically. The virtual auto model also falls back to claude-sonnet, because the language-routing callback rewrites auto to a specific model before dispatch — if that model then fails, LiteLLM uses the fallback registered for the original group name (auto).
Langfuse spans: The async_log_success_event callback logs to Langfuse only for completion and acompletion call types — management API calls (/v2/user/info, etc.) are skipped to avoid noise traces. Each trace carries:
| Observation | When |
|---|---|
routing span |
every routed request — detected script and selected model |
<model>/response generation |
the model call itself, with tokens and computed cost |
tool-result/<name> span |
each MCP tool the gateway executed — arguments, truncated result, latency |
<model>/tool-hop-N generation |
each follow-up model call in the tool loop, with its own tokens and cost |
mcp-loop/incomplete span |
the loop ended without prose (hops exhausted or a follow-up failed) |
The hop generations matter for cost: a follow-up call names an explicit model, so async_pre_call_hook returns before setting routed_model and async_log_success_event skips it. Without them the loop's token spend — often several times the first call's — would be missing from the trace entirely. _run_agentic_loop therefore logs each hop itself rather than relying on metadata surviving the round trip through the proxy.
The routing logic lives in docker/litellm_callbacks.py as a CustomLogger pre-call hook and is controlled by stack.yaml:
layers:
gateway:
options:
language_routing:
enabled: true
english_model: qwen-7b
multilingual_model: claude-sonnet # Korean
cjk_model: qwen-7b
threshold: 0.15
routing:
fallbacks:
- from: qwen-7b
to: [claude-sonnet]
- from: auto
to: [claude-sonnet]
When language_routing.enabled is true, render adds callbacks: [callbacks.language_router] to litellm_settings in the generated config, and the callback file is mounted read-only into the litellm container at /app/callbacks.py.
To disable routing and send all requests to a single model, set language_routing.enabled: false and re-render.
Automated scoring¶
Every completion trace receives five scores automatically. No client change is needed.
Score inventory¶
| Score name | Type | Description |
|---|---|---|
routing_accuracy |
BOOLEAN | 1.0 when the request reached the intended model; 0.0 when a fallback was triggered |
language_consistency |
BOOLEAN | 1.0 when the dominant script of the output matches the dominant script of the input |
latency_score |
NUMERIC | Linear decay from 1.0 at 0 s to 0.0 at 30 s — configurable cap |
helpfulness |
NUMERIC | LLM-as-judge score (0.0–1.0): how well the response answers the user's question |
judge_language_match |
BOOLEAN | LLM-as-judge opinion on whether the response language matches the request language |
Score tiers¶
Rule-based (synchronous) — routing_accuracy, language_consistency, and latency_score are computed inside async_log_success_event before the event returns. They add no latency to the gateway response.
LLM-as-judge (asynchronous) — helpfulness and judge_language_match are computed by a direct call to the Anthropic API (claude-haiku-4-5-20251001). The call is dispatched as a fire-and-forget asyncio.Task so it does not block the response path. The judge calls Anthropic directly via httpx rather than through LiteLLM to avoid triggering the scoring callbacks recursively.
User feedback (sidecar) — a lightweight FastAPI service (feedback) runs alongside the gateway. It maintains a SHA-256(content[:500]) → trace-id index. The LiteLLM callback populates the index automatically after each response via POST /register. LibreChat feedback can be correlated to a Langfuse trace by posting the response text to POST /feedback:
curl -X POST http://localhost:8080/feedback \
-H "Content-Type: application/json" \
-d '{"content": "<first 500 chars of response>", "rating": 1}'
The service accepts rating as 1 (thumbs-up) or -1 (thumbs-down).
Implementation notes¶
success_callback: []indocker/litellm_config.yamlis intentional. All Langfuse logging, including scoring, is performed manually inside the callback, so management API calls never produce traces or scores.- The judge model is set by
_JUDGE_MODELindocker/litellm_callbacks.py. Changing this requires rebuilding thelitellmcontainer. - The latency cap is
_LATENCY_CAP_S = 30.0. Adjust inlitellm_callbacks.pyand rebuild. ANTHROPIC_API_KEYmust be set in.envfor judge scores to be computed. If absent, the judge task exits silently and the two judge scores are omitted.
Image generation¶
Image generation is triggered directly from chat — type an image-intent message (e.g. "그려줘", "draw …", "generate an image of …") in the auto chat window and the UnifiedRouter callback detects the intent, calls Cloudflare Workers AI directly, and replaces the LLM response with the generated image before it reaches LibreChat.
flowchart LR
LC["LibreChat\nchat (auto)"]
CB["UnifiedRouter\npre_call_hook"]
CF["Cloudflare Workers AI\nFLUX.1-schnell"]
MN["MinIO\n(media.<domain>)"]
LF["Langfuse trace"]
LC -- "image keywords\nin message" --> CB
CB -- "asyncio.Task" --> CF
CF -- "store image" --> MN
CB -- "1-token LLM call\n(placeholder)" --> LC
CB -- "streaming_hook:\nreplace with\n" --> LC
CB -- "trace" --> LF
The callback calls Cloudflare Workers AI directly via httpx (LiteLLM's built-in image generation endpoint does not correctly support Cloudflare). The generated image is stored in MinIO and served via the media.<domain> subdomain. The async_post_call_streaming_iterator_hook drains the 1-token LLM stream and replaces it with a streaming SSE chunk containing the markdown image link. LibreChat renders the image inline.
Image generation is enabled by default when the credentials are present.
The relevant stack.yaml block:
layers:
gateway:
options:
image_generation:
enabled: true
# dalle_alias is the model name LibreChat requests for image generation.
# The UnifiedRouter callback intercepts chat requests with this alias
# and routes them to Cloudflare — it does NOT call OpenAI DALL-E.
dalle_alias: dall-e-3
timeout: 90
providers:
cloudflare:
model: "@cf/black-forest-labs/flux-1-schnell"
api_key_env: CF_API_TOKEN
account_id_env: CF_ACCOUNT_ID
The credentials (CF_API_TOKEN, CF_ACCOUNT_ID) are optional.
If absent, image generation fails with an API error; the chat path is unaffected.
To disable cleanly, set image_generation.enabled: false and re-render.
See Credentials — Image generation for token acquisition steps.
MCP tool layer (Phase 2)¶
MCP tools are wired at the gateway layer, not at the client. Every application that reaches LiteLLM gains the same tool access automatically — no per-client configuration required.
flowchart LR
LC["LibreChat / apps"]
GW["LiteLLM Gateway\n:4000"]
MCP["mcp-clickhouse\n:9100 (internal)"]
CH["ClickHouse Cloud"]
LC --> GW
GW -- "MCP / SSE" --> MCP
MCP -- "SQL (CLICKHOUSE_SECURE=true)" --> CH
GW -. "tool call traces" .-> LF["Langfuse"]
mcp-clickhouse (official ClickHouse MCP server) connects to ClickHouse Cloud
as a database client and exposes an SSE endpoint on port 9100 inside the Docker
network. mcp-proxy wraps its stdio transport. That port is internal to the
Docker network — no security group change needed for the aws-ec2 target.
Where the tool loop runs¶
The gateway owns the whole loop. UnifiedRouter fetches the tool definitions
from LiteLLM's /mcp-rest/tools/list, injects them into the request, and — when
the model answers with tool_calls — executes each call against
/mcp-rest/tools/call and re-prompts, up to five hops, until the model returns a
plain answer. Buffered streaming means a streaming client such as LibreChat sees
only the final text. Nothing MCP-specific is configured in librechat.yaml.
If the five hops run out, the gateway makes one more call asking the model to answer from the results it already has. A capable model will always find one more query worth running, and the request is not worth abandoning with the data already fetched.
The client never receives unexecuted tool_calls
If the loop cannot produce prose, the gateway replies with a short apology and
records an mcp-loop/incomplete span — it does not pass the model's
tool_calls back. A client with no MCP configuration cannot run them, and
LibreChat renders that as Tool "<name>" not found, which reads like a
missing tool rather than the gateway-side problem it actually is.
Tools are injected for Korean-language requests only
Language routing sends Korean to claude-sonnet and English/CJK to
qwen-7b, and qwen-7b is not reliable enough at function calling to be
handed tool schemas. So UnifiedRouter injects MCP tools only when the last
user message is Hangul-primary, and skips injection for image-intent
messages and for its own follow-up completions. Tools supplied by the
client are handled separately: those route to claude-sonnet regardless of
language.
stack.yaml — tools layer¶
layers:
tools:
phase: 2
enabled: true
impl: mcp
managed_by: compose
compose_profile: tools
servers:
clickhouse:
enabled: true
transport: sse
url_env: MCP_CLICKHOUSE_URL
Enable this layer for --profile phase-2:
Generated litellm_config.yaml¶
render --profile phase-2 appends the mcp_servers block to the generated
docker/litellm_config.yaml:
mcp_servers:
clickhouse:
url: "http://mcp-clickhouse:9100/sse"
transport: "sse"
Deploying Phase 2¶
Set the required credentials first (see Credentials — MCP (ClickHouse Cloud)):
Domain / HTTPS proxy (aws-ec2)¶
The aws-ec2 target runs a Caddy reverse proxy that provides automatic HTTPS
via Let's Encrypt — no certificate management required.
It is not optional in practice: the security group publishes 80 and 443 only, so Caddy's subdomains are the only route to any service. The application ports (3000, 3080, 4000, 9002) are closed — they carried plain HTTP, and 4000 fronts the gateway's admin API.
Configuring a domain¶
Set the domain interactively and push to SSM before provisioning:
./scripts/stack.sh secrets domain # prompts for DOMAIN_BASE and DOMAIN_SSL_EMAIL
./scripts/stack.sh secrets push --target aws-ec2
DNS setup — create A records for each subdomain pointing at the EC2 public IP:
| Subdomain | Service |
|---|---|
chat.<domain> |
LibreChat |
langfuse.<domain> |
Langfuse |
litellm.<domain> |
LiteLLM |
media.<domain> |
MinIO (image storage) |
The root domain is not touched — it can point elsewhere (e.g. a homepage).
How it works¶
At first boot, bootstrap-ec2.sh reads DOMAIN_BASE from SSM. If set:
- Runs
./scripts/stack.sh render --target aws-ec2to generatedocker/caddy/Caddyfile - Starts the
proxycompose profile (caddyservice) - Caddy performs the ACME TLS-ALPN-01 challenge on port 443 and obtains certs from Let's Encrypt
Certificates are stored in a caddy-data named volume and auto-renewed at 30 days before expiry. As long as the instance is reachable on port 443, certs stay current indefinitely.
Access restriction¶
LibreChat allows registration only from permitted email domains:
Anyone can reach the registration form, but accounts are created only for addresses matching the configured domain. Set to an empty value to allow all domains.
Subdomain defaults¶
Subdomain prefixes are declared in stack.yaml under targets.aws-ec2.domain.subdomains
and can be customised:
targets:
aws-ec2:
domain:
base: "" # overridden by DOMAIN_BASE env var
ssl_email: "" # overridden by DOMAIN_SSL_EMAIL env var
subdomains:
litellm: litellm
langfuse: langfuse
librechat: chat
media: media
The same map drives three things, so they cannot drift: the rendered Caddyfile,
the URLs stack.sh urls prints, and the endpoints stack.sh status and
smoke-test probe.
How health checks resolve¶
stack.yaml gives each check both forms:
- name: litellm
layer: gateway
url: http://{{host}}:4000/health/liveliness # docker, or aws-ec2 with no domain
subdomain: litellm # aws-ec2 with a domain
path: /health/liveliness
With DOMAIN_BASE set and --target aws-ec2, status probes
https://litellm.<domain>/health/liveliness. Without it, it falls back to the
direct port and warns — on EC2 that port is closed, so every check would
otherwise report a false outage. A check marked internal: true
(mcp-clickhouse) is skipped against a remote host instead of being reported
as down.
Generated files¶
docker/litellm_config.yaml, docker/librechat.yaml, and docker/caddy/Caddyfile carry a generation banner and must not be hand-edited — render overwrites them, and up renders by default. The Caddyfile is only generated when --target aws-ec2 is used and DOMAIN_BASE is set.
# GENERATED by scripts/stack.sh render — DO NOT EDIT.
# Edit stack.yaml and re-run `./scripts/stack.sh render`.
# profile: phase-3
model_list:
- model_name: qwen-7b
litellm_params:
model: openai/Qwen/Qwen2.5-7B-Instruct
api_base: os.environ/VLLM_API_BASE
api_key: os.environ/VLLM_API_KEY
rpm: 60
tpm: 200000
timeout: 600
model_info:
mode: chat
max_input_tokens: 32768
input_cost_per_token: 0
output_cost_per_token: 0
- model_name: claude-sonnet
litellm_params:
model: anthropic/claude-sonnet-4-5
api_key: os.environ/ANTHROPIC_API_KEY
rpm: 500
tpm: 400000
timeout: 30
model_info:
mode: chat
max_input_tokens: 200000
input_cost_per_token: 3e-06
output_cost_per_token: 1.5e-05
- model_name: auto
litellm_params:
model: qwen-7b
model_info:
mode: chat
The committed copy is rendered with --profile phase-3, which is what the EC2
deployment runs. render --profile phase-1 produces the same file without
qwen-7b and without the mcp_servers block.
They are committed anyway, so a reader can see the gateway config without running anything. Use --no-render for a one-off manual override.