llama.cpp
| Class |
Description |
LlamaCppAppEnvironment |
App environment backed by llama.cpp (llama-server) for serving GGUF models. |
| Property |
Type |
Description |
DEFAULT_LLAMA_CPP_IMAGE |
Image |
|
def build_llama_cpp_image(
name: str = 'llama-cpp-app-image',
cuda: bool = True,
cuda_arch: str = '89',
repo: str = 'https://github.com/ggml-org/llama.cpp',
ref: str | None = None,
) -> flyte.Image
Build a Debian image with llama-server compiled from source.
| Parameter |
Type |
Description |
name |
str |
Name of the image. |
cuda |
bool |
Build with CUDA support (GGML_CUDA=ON). Set to False for a CPU-only image. |
cuda_arch |
str |
Target CUDA architecture(s) for the kernel build, as a “;"-separated list of compute capabilities (e.g. “89” for L4/L40S, “80;86;89;90” for a fat binary that also covers A100/A10/H100). Ignored when cuda=False. |
repo |
str |
Git repository to build llama.cpp from. |
ref |
str | None |
Git ref (tag, branch, or commit) to check out. None builds the default branch tip; pin a release tag (e.g. “b6148”) for reproducible builds. |