LlamaCppAppEnvironment
Package: flyteplugins.llamacpp
App environment backed by llama.cpp (llama-server) for serving GGUF models.
This environment serves an OpenAI-compatible endpoint (under /v1) plus the llama.cpp
Web UI, with the specified GGUF model and configuration. llama.cpp shines where vLLM and
SGLang don’t fit: quantized GGUF weights, partial CPU offload of models larger than VRAM,
and CPU-only serving.
Parameters
class LlamaCppAppEnvironment(
name: str,
depends_on: List[Environment] = <factory>,
pod_template: Optional[Union[str, PodTemplate]] = None,
description: Optional[str] = None,
secrets: Optional[SecretRequest] = None,
env_vars: Optional[Dict[str, str]] = None,
resources: Optional[Resources] = None,
interruptible: bool = False,
include: Tuple[str, ...] = <factory>,
service_account: Optional[str] = None,
args: Optional[Union[List[str], str]] = None,
command: Optional[Union[List[str], str]] = None,
requires_auth: bool = True,
scaling: Scaling = <factory>,
domain: Domain | None = <factory>,
links: List[Link] = <factory>,
parameters: List[Parameter] = <factory>,
cluster_pool: str = 'default',
cluster: str | None = None,
timeouts: Timeouts = <factory>,
image: str | Image | Literal['auto'] = Image(base_image='ghcr.io/flyteorg/flyte:py3.12-v2.11.0', dockerfile=None, registry=None, name='llama-cpp-app-image', platform=('linux/amd64', 'linux/arm64'), python_version=(3, 12), extendable=True, _is_cloned=True, _ref_name=None, _layers=(AptPackages(git='build-essential'), Commands(wget https://developer.download.nvidia.com/compute/cuda/repos/debian12/x86_64/cuda-keyring_1.1-1_all.deb='dpkg -i cuda-keyring_1.1-1_all.deb'), Commands(wget -q https://nodejs.org/dist/v22.12.0/node-v22.12.0-linux-x64.tar.xz -O /tmp/node.tar.xz='mkdir -p /opt/node && tar -xJf /tmp/node.tar.xz -C /opt/node --strip-components=1 && rm /tmp/node.tar.xz'), Env(env_vars=(('PATH', '/opt/llama.cpp/build/bin:/usr/local/cuda-12.8/bin:$PATH'), ('LLAMA_CACHE', '/tmp/llama.cpp/cache'), ('CUDA_HOME', '/usr/local/cuda-12.8'))), PipPackages(pre=True, packages=('flyteplugins-llamacpp',))), _tag=None, _image_registry_secret=None),
type: str = 'llama.cpp',
port: int | Port = 8080,
extra_args: str | list[str] = '',
model_path: str | RunOutput | ArtifactValue = '',
model_hf_path: str = '',
model_id: str = '',
draft_model_path: str | RunOutput | ArtifactValue = '',
draft_model_hf_path: str = '',
)| Parameter | Type | Description |
|---|---|---|
name |
str |
The name of the application. |
depends_on |
List[Environment] |
|
pod_template |
Optional[Union[str, PodTemplate]] |
|
description |
Optional[str] |
|
secrets |
Optional[SecretRequest] |
Secrets that are requested for application. |
env_vars |
Optional[Dict[str, str]] |
Environment variables to set for the application. |
resources |
Optional[Resources] |
|
interruptible |
bool |
|
include |
Tuple[str, ...] |
|
service_account |
Optional[str] |
|
args |
Optional[Union[List[str], str]] |
|
command |
Optional[Union[List[str], str]] |
|
requires_auth |
bool |
Whether the public URL requires authentication. |
scaling |
Scaling |
Scaling configuration for the app environment. |
domain |
Domain | None |
Domain to use for the app. |
links |
List[Link] |
|
parameters |
List[Parameter] |
|
cluster_pool |
str |
The target cluster_pool where the app should be deployed. |
cluster |
str | None |
|
timeouts |
Timeouts |
|
image |
str | Image | Literal['auto'] |
|
type |
str |
Type of app. |
port |
int | Port |
Port the application listens on. Defaults to 8080. |
extra_args |
str | list[str] |
Extra args to pass to llama-server, e.g. "--ctx-size 32768 --jinja". Run llama-server --help or see https://github.com/ggml-org/llama.cpp/tree/master/tools/server for details. |
model_path |
str | RunOutput | ArtifactValue |
Remote path to the GGUF weights – a directory containing .gguf file(s) or a direct path to one (e.g. s3://bucket/path/to/model), or a RunOutput/ArtifactValue resolved at deploy time. The weights are downloaded into the container and the served .gguf is located at startup (for sharded models, the -00001-of- shard is picked; llama-server finds the rest). |
model_hf_path |
str |
Hugging Face GGUF repo, optionally with a quant tag (e.g. ggml-org/gemma-3-4b-it-GGUF:Q4_K_M). Passed to llama-server as --hf-repo, which downloads the weights at startup. |
model_id |
str |
Model id exposed by the server (llama-server’s --alias). |
draft_model_path |
str | RunOutput | ArtifactValue |
Remote path to the draft model GGUF used for speculative decoding, or a RunOutput/ArtifactValue resolved at deploy time. Downloaded alongside the target model and passed as --model-draft. Tune the speculation via extra_args (--draft-max, --draft-min, --gpu-layers-draft, …). |
draft_model_hf_path |
str |
Hugging Face GGUF repo for the draft model, as an alternative to draft_model_path. Passed as --hf-repo-draft. |
Properties
| Property | Type | Description |
|---|---|---|
endpoint |
str |
Methods
| Method | Description |
|---|---|
add_dependency() |
Add one or more environment dependencies so they are deployed together. |
clone_with() |
|
container_args() |
Return the container arguments for llama.cpp. |
container_cmd() |
|
get_port() |
|
on_shutdown() |
Decorator to define the shutdown function for the app environment. |
on_startup() |
Decorator to define the startup function for the app environment. |
server() |
Decorator to define the server function for the app environment. |
add_dependency()
def add_dependency(
*env: Environment,
)Add one or more environment dependencies so they are deployed together.
When you deploy this environment, any environments added via
add_dependency will also be deployed. This is an alternative to
passing depends_on=[...] at construction time, useful when the
dependency is defined after the environment is created.
Duplicate dependencies are silently ignored. An environment cannot depend on itself.
| Parameter | Type | Description |
|---|---|---|
*env |
Environment |
One or more Environment instances to add as dependencies. |
clone_with()
def clone_with(
name: str,
image: Optional[Union[str, Image, Literal['auto']]] = None,
resources: Optional[Resources] = None,
env_vars: Optional[dict[str, str]] = None,
secrets: Optional[SecretRequest] = None,
depends_on: Optional[list[Environment]] = None,
description: Optional[str] = None,
interruptible: Optional[bool] = None,
**kwargs: Any,
) -> LlamaCppAppEnvironment| Parameter | Type | Description |
|---|---|---|
name |
str |
|
image |
Optional[Union[str, Image, Literal['auto']]] |
|
resources |
Optional[Resources] |
|
env_vars |
Optional[dict[str, str]] |
|
secrets |
Optional[SecretRequest] |
|
depends_on |
Optional[list[Environment]] |
|
description |
Optional[str] |
|
interruptible |
Optional[bool] |
|
**kwargs |
Any |
container_args()
def container_args(
serialization_context: SerializationContext,
) -> list[str]Return the container arguments for llama.cpp.
| Parameter | Type | Description |
|---|---|---|
serialization_context |
SerializationContext |
container_cmd()
def container_cmd(
serialize_context: SerializationContext,
parameter_overrides: list[Parameter] | None = None,
) -> List[str]| Parameter | Type | Description |
|---|---|---|
serialize_context |
SerializationContext |
|
parameter_overrides |
list[Parameter] | None |
get_port()
def get_port()on_shutdown()
def on_shutdown(
fn: F,
) -> FDecorator to define the shutdown function for the app environment.
This function is called after the server function is called.
This decorated function can be a sync or async function, and accepts input parameters based on the Parameters defined in the AppEnvironment definition.
| Parameter | Type | Description |
|---|---|---|
fn |
F |
on_startup()
def on_startup(
fn: F,
) -> FDecorator to define the startup function for the app environment.
This function is called before the server function is called.
The decorated function can be a sync or async function, and accepts input parameters based on the Parameters defined in the AppEnvironment definition.
| Parameter | Type | Description |
|---|---|---|
fn |
F |
server()
def server(
fn: F,
) -> FDecorator to define the server function for the app environment.
This decorated function can be a sync or async function, and accepts input parameters based on the Parameters defined in the AppEnvironment definition.
| Parameter | Type | Description |
|---|---|---|
fn |
F |