Serge
Self-hosted LLaMA chat UI with no API keys needed
Serge is a self-hosted chat interface for running LLaMA-family models entirely offline. It bundles llama.cpp with a clean web UI and requires no remote APIs or accounts.
Key features
- No API keys required
- Bundled llama.cpp
- Simple Docker deploy
- Fully offline
Strengths
- Released under the Apache-2.0 license
- Easy to set up — beginner-friendly
- First-class Docker support for quick deployment
- Mature project with 5.7k GitHub stars
Serge replaces
Compare Serge
25 head-to-head comparisons.
- Serge vs Ollama
- Serge vs llama.cpp
- Serge vs vLLM
- Serge vs GPT4All
- Serge vs LiteLLM
- Serge vs exo
- Serge vs New API
- Serge vs Jan
- Serge vs FastChat
- Serge vs One API
- Serge vs Continue
- Serge vs llamafile
- Serge vs MLC LLM
- Serge vs OpenLLM
- Serge vs KoboldCpp
- Serge vs Text Generation Inference
- Serge vs Petals
- Serge vs Page Assist
- Serge vs LMDeploy
- Serge vs Enchanted
- Serge vs LM Studio
- Serge vs Text Embeddings Inference
- Serge vs Harbor LLM Toolkit
- Serge vs ik_llama.cpp
- Serge vs Aphrodite Engine
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API