vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
The 5 best Kubernetes local llm runners you can self-host, ranked by community traction.
High-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
Unified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
Run any open LLM as an OpenAI-compatible API endpoint
Replaces OpenAI API
Hugging Face toolkit for production LLM serving
Replaces OpenAI API
Toolkit for compressing and serving large language models
Replaces OpenAI API, Hugging Face Inference Endpoints
No apps match these filters.
Every option here is open-source and self-hostable, tagged Kubernetes within the local llm runners category. Compare them on the individual app pages for setup difficulty and resource requirements.