Configuration¶
llama.cpp is configured differently on the two platforms:
- EC2 — pass native
llama-serverflags as container arguments. A small set of DLC environment variables control authentication and the upstream port. - Amazon SageMaker AI — set
SM_LLAMA_CPP_*environment variables on the container, and the entrypoint translates them intollama-serverflags.
The complete, authoritative flag list is the upstream llama-server documentation — this page covers the DLC-specific surface and the flags most users need.
EC2 Environment Variables¶
| Variable | Default | Description |
|---|---|---|
LLAMA_API_KEY |
(unset) | When set, it is passed to llama-server --api-key, so every request must then carry Authorization: Bearer <key>. When left unset, the endpoint is unauthenticated |
On EC2, everything else is a llama-server argument appended to docker run (the host/port are fixed at 0.0.0.0:8080 by the
entrypoint). See Common llama-server Flags.
SageMaker Environment Variables¶
The Amazon SageMaker AI entrypoint reads a few control variables and then translates every other SM_LLAMA_CPP_* variable into a llama-server flag:
SM_LLAMA_CPP_FOO_BAR=value → --foo-bar value. A value of true becomes a bare flag (--foo-bar), and a value of false is dropped.
| Variable | Default | Description |
|---|---|---|
SM_LLAMA_CPP_MODEL |
(unset) | Explicit GGUF path, passed as --model — skips the /opt/ml/model auto-detection |
SM_LLAMA_CPP_MODEL_DIR |
/opt/ml/model |
Directory scanned for a *.gguf when SM_LLAMA_CPP_MODEL is unset |
SM_LLAMA_CPP_HF_REPO |
(unset) | Mapped to --hf-repo — the HuggingFace repo to download the GGUF from at startup |
SM_LLAMA_CPP_HF_FILE |
(unset) | Mapped to --hf-file — the GGUF filename within the HuggingFace repo |
SM_LLAMA_CPP_CTX_SIZE |
(llama-server default) | Mapped to --ctx-size — the context window in tokens |
SM_LLAMA_CPP_PORT |
8080 |
Public port nginx serves on |
LLAMA_CPP_UPSTREAM_PORT |
8081 |
Loopback port llama-server binds behind nginx |
SM_LLAMA_CPP_<ANY> |
(unset) | Any other suffix maps to --<any> on llama-server |
SM_LLAMA_CPP_MODEL, SM_LLAMA_CPP_MODEL_DIR, and SM_LLAMA_CPP_PORT are consumed by the entrypoint itself and are not forwarded as flags.
Common llama-server Flags¶
These are passed as container arguments on EC2, or as SM_LLAMA_CPP_* variables on Amazon SageMaker AI (e.g. --n-gpu-layers ↔
SM_LLAMA_CPP_N_GPU_LAYERS).
| Flag | Applies to | Description |
|---|---|---|
--model <path> |
all | Path to a local GGUF file |
--hf-repo <repo> / --hf-file <file> |
all | Download a GGUF from HuggingFace at startup |
--ctx-size <n> |
all | Context window size in tokens |
--n-gpu-layers <n> |
GPU only | Layers to offload to the GPU (999 = all) — has no effect on the CPU or Graviton images |
--parallel <n> |
all | Number of parallel request slots |
--threads <n> |
all | CPU threads for generation |
--batch-size <n> |
all | Logical batch size |
--api-key <key> |
all | Require a bearer token (on EC2, prefer the LLAMA_API_KEY env var) |
The DLC sets no defaults for concurrency, threads, batch size, or context size beyond llama-server's own — tune them for your model and instance.
Known Limitations¶
- GGUF only. The server loads GGUF-format models, so other formats must be converted first (see Supported Models).
- No baked-in model. You must supply a GGUF at launch via a local file,
/opt/ml/model, or a HuggingFace download. - Unauthenticated by default on EC2. The endpoint binds
0.0.0.0:8080with no auth unlessLLAMA_API_KEYis set — run it inside a private network. The Amazon SageMaker AI path is gated by SageMaker's own request authentication. - HuggingFace download needs network egress. The
--hf-repo/SM_LLAMA_CPP_HF_REPOpath is not air-gapped. Mount the GGUF locally to run offline. - GPU on Amazon SageMaker AI needs a recent driver AMI. The GPU image ships CUDA 13.0.2, so pin a recent
InferenceAmiVersion— see SageMaker Deployment.