Troubleshooting¶
Collected fixes for issues encountered during development and deployment of the Phase 1 stack.
Terraform / AWS provisioning¶
terraform apply fails: apostrophe in security group description
Symptom:
Cause: AWS security group descriptions do not allow apostrophes. Any
description containing a ' character (e.g. "Let's Encrypt") will fail
AWS validation.
Fix: Remove apostrophes from all description fields in sg.tf, then
re-run terraform apply.
AWS session token invalid when passed inline with line breaks
Symptom:
Cause: Pasting a multi-line terraform apply command in the terminal
(with environment variables inlined) can split the AWS_SESSION_TOKEN value
across shell lines, embedding newlines into the header value.
Fix: export each variable separately first, then run terraform apply
on its own line:
EC2 public IP changes after terraform apply
Symptom: The Terraform output shows a new public_ip after a re-apply
(e.g., when user_data changes). Services were reachable before but now
return certificate errors or fail to connect.
Cause: EC2 instances with dynamic (non-Elastic IP) addresses are
assigned a new public IP when stopped and started. Any user_data change
forces an instance replacement, which triggers this.
Fix:
- Note the new
public_ipfromterraform output. - Update all four DNS
Arecords (chat,langfuse,litellm,media) to the new IP. - Wait for DNS propagation, then restart Caddy so it re-issues TLS certificates against the new IP:
Caddy TLS certificate fails after an IP change
Symptom:
Cause: Let's Encrypt's validation servers still resolve the domain to the old IP because DNS has not yet propagated.
Fix: Wait for DNS propagation (usually a few minutes for TTL-1 records, up to the TTL otherwise), then restart Caddy — it retries ACME automatically:
Do not restart repeatedly before DNS has propagated; Let's Encrypt rate-limits failed certificate requests.
status --target aws-ec2 reports every service down, or a direct port times out
Symptom: ./scripts/stack.sh status --target aws-ec2 shows 000 for
litellm, langfuse, librechat, and minio, while the site works fine in a
browser. Or curl http://<ec2-ip>:4000/... hangs.
Cause: Expected. The security group publishes 80 and 443 only — the application ports are closed. Everything is served through Caddy's HTTPS subdomains.
Fix: Give the command the domain, so it probes the same route a browser does:
./scripts/stack.sh secrets domain # persists DOMAIN_BASE to .env
# or, one-off:
DOMAIN_BASE=example.com ./scripts/stack.sh status --target aws-ec2
status warns when DOMAIN_BASE is unset on the aws-ec2 target for exactly
this reason. To reach a port directly anyway, tunnel over SSH rather than
reopening it:
SSH times out on a host that worked yesterday
Cause: There is no standing SSH ingress — ssh_allowed_cidrs is empty
by default, and a rule opened for an earlier session was closed (or your
address changed).
Fix: Open it for this session, then close it. See Deployment — SSH access.
LiteLLM callback configuration¶
Custom callback not loading — UnifiedRouter never runs
Symptom: LiteLLM startup log shows:
UnifiedRouter is never invoked even though callbacks.py is mounted.
Langfuse traces are absent for completion calls.
Cause: LiteLLM proxy silently ignores the custom_callbacks key in
litellm_settings. Only the callbacks key is honoured.
Fix: Use callbacks: [callbacks.language_router] (a
module.attribute path), not custom_callbacks:
litellm_settings:
callbacks: [callbacks.language_router]
success_callback: []
The callback module is callbacks.py (mounted at /app/callbacks.py) and
language_router is the UnifiedRouter() instance at module level.
success_callback is intentionally empty. Langfuse logging is performed
manually inside UnifiedRouter.async_log_success_event via the Langfuse
SDK, filtered to completion calls only. Adding langfuse to
success_callback causes LiteLLM to log all management API calls
(e.g. /v2/user/info, /videos/{video_id}) as traces with
"default-message-value" placeholder inputs.
async_pre_call_hook never called for LibreChat requests
Symptom: The hook is defined and the module loads successfully, but the hook body never executes for requests sent from LibreChat.
Cause: LibreChat sends requests asynchronously. LiteLLM sets
call_type = 'acompletion', not 'completion'. A guard that checks only
for 'completion' will skip all LibreChat traffic.
Fix: Check for both call types:
Image generation¶
Image generation via /v1/images/generations returns {\"data\":[]}
Symptom: The LiteLLM image generation endpoint returns an empty data
array in ~17 ms. No image is produced and no error is logged.
Cause: LiteLLM proxy does not natively support Cloudflare Workers AI as an image generation backend. It returns an empty response without error.
Fix: Call the provider APIs directly using httpx inside the callback.
Do not route image generation through LiteLLM's /v1/images/generations
endpoint. The UnifiedRouter callback handles this in the
async_pre_call_hook by dispatching an asyncio.Task that calls the
providers directly.
LibreChat shows a blank response for image generation
Symptom: The API returns the correct 
markdown, but LibreChat displays an empty message bubble.
Cause: Setting data["stream"] = False in async_pre_call_hook causes
LiteLLM to return a single non-streaming JSON response. LibreChat's stream
handler expects SSE events and silently discards the JSON body.
Fix: Do not force stream=False in async_pre_call_hook. Instead,
implement async_post_call_streaming_iterator_hook to drain the 1-token
streaming response and yield new SSE chunks containing the image markdown.
The stream stays open from LibreChat's perspective; only the content changes.
LibreChat does not render inline data: URI images
Symptom: A response containing 
appears as a blank image placeholder in LibreChat.
Cause: Browser Content Security Policy blocks inline data: images
embedded in markdown rendered inside LibreChat.
Fix: Upload the generated image to MinIO and return a public https://
URL instead (e.g. https://media.<domain>/images/generated/<id>.jpg). The
UnifiedRouter callback stores the image in MinIO before injecting the
markdown link.
MCP tool layer (Phase 2)¶
mcp-clickhouse container exits immediately — --transport sse not supported
Symptom: The container stops at startup. Logs show an error about an unrecognised flag or unsupported transport.
Cause: The installed mcp-clickhouse package version does not accept
--transport sse as a CLI argument.
Fix: Use the mcp-proxy wrapper in docker/mcp/Dockerfile to expose
the stdio-only server over SSE:
Rebuild the image:
LiteLLM /mcp endpoint returns 404
Symptom: curl http://localhost:4000/mcp returns a 404 or empty
response.
Cause: The stack was rendered without --profile phase-2, so the
mcp_servers block is absent from docker/litellm_config.yaml.
Fix:
Confirm the block is present:
LibreChat cannot reach mcp-clickhouse — SSRF protection blocks internal addresses
Applies only if you wire MCP into LibreChat directly. This stack does
not: the gateway injects the tools and runs the loop, and librechat.yaml
contains no MCP configuration at all. Recorded because the symptom is easy
to hit when reintroducing client-side wiring.
Symptom: MCP tools never appear in LibreChat, or tool calls return a connection error. LibreChat logs show an SSRF-related rejection.
Cause: LibreChat's built-in SSRF protection blocks requests to private
or Docker-internal network addresses. http://mcp-clickhouse:9100/sse is an
internal hostname and is blocked by default.
Fix: Add the MCP server hostname to mcpSettings.allowedAddresses.
librechat.yaml is generated, so the change belongs in the renderer in
scripts/stack.sh, not in the file:
mcpSettings:
allowedAddresses:
- "mcp-clickhouse:9100"
Recreate the LibreChat container to pick up the change:
A Korean question triggers a ClickHouse tool call but the same question in English does not
Not a bug. UnifiedRouter injects MCP tools only when the last user
message is Hangul-primary. Those requests route to claude-sonnet, which is
reliable at function calling; English and CJK route to qwen-7b, which is
not, so it is never handed tool schemas.
The gate is in docker/litellm_callbacks.py — the _pre_script == "hangul"
condition in async_pre_call_hook. Widening it means either serving a
tool-capable self-hosted model or routing tool-shaped requests to
claude-sonnet regardless of language.
Confirm which branch ran from the gateway log:
Tool calls fail — ClickHouse connection refused or authentication error
Symptom: The /mcp endpoint responds but tool calls return a
connection or authentication error. Container logs show Connection refused
or Code: 516. DB::Exception: Authentication failed.
Cause: CLICKHOUSE_HOST, CLICKHOUSE_USER, or CLICKHOUSE_PASSWORD
is wrong, or CLICKHOUSE_SECURE is not set to true (required for
ClickHouse Cloud).
Fix: Check the live environment variables:
Update credentials and restart:
Langfuse traces¶
Traces missing — \"Event type not accepted\" in ingestion response
See Deployment — Troubleshooting for the
full fix. Short version: set LANGFUSE_MIGRATION_V4_WRITE_MODE=dual in
docker-compose.yml for the langfuse-web and langfuse-worker services,
then recreate the containers:
No fallback model group found for original model_group=auto
See Deployment — Troubleshooting for the
full fix. Short version: add auto to the fallback list in stack.yaml
under layers.gateway.options.routing.fallbacks, then re-render.
Traces show default-message-value as the user input
Symptom: Langfuse contains traces with "default-message-value" as the
input. They often correspond to management API paths such as
/v2/user/info or /videos/{video_id} rather than chat completions.
Cause: When success_callback: [langfuse] is set, LiteLLM forwards
every successful call — including internal management API calls — to
Langfuse. Those requests carry no chat messages, so LiteLLM substitutes a
"default-message-value" placeholder.
Fix: Remove langfuse from success_callback and rely on
UnifiedRouter.async_log_success_event to log to Langfuse via the SDK.
The custom handler filters to completion calls only, so management API
traffic is never forwarded:
litellm_settings:
callbacks: [callbacks.language_router]
success_callback: []
Confirm via the LiteLLM startup log:
After the fix, only actual chat completion requests appear in Langfuse.