LL

LLMKube

Kubernetes operator for llama.cpp-native LLM inference with GPU

Local LLM Runners ★ 186 stars Medium setup Apache-2.0

Kubernetes operator for llama.cpp-native LLM inference with GPU scheduling, Apple Silicon Metal support, and OpenAI-compatible API.

Strengths

  • Released under the Apache-2.0 license
  • First-class Docker support for quick deployment
  • Kubernetes-ready with Helm charts available
  • Written in Go

Compare LLMKube

10 head-to-head comparisons.

Similar local llm runners apps