Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
The 18 best Offline-First local llm runners you can self-host, ranked by community traction.
Run large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
High-performance LLM inference in plain C/C++
Replaces OpenAI API
High-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
Privacy-first desktop chat with local language models
Replaces ChatGPT
Run your own AI cluster across everyday devices
Replaces OpenAI API
Open-source offline ChatGPT alternative for the desktop
Replaces ChatGPT
Open-source AI code assistant for VS Code and JetBrains
Replaces GitHub Copilot, Cursor
Distribute and run LLMs with a single executable file
Replaces OpenAI API, ChatGPT
Run any open LLM as an OpenAI-compatible API endpoint
Replaces OpenAI API
Single-file local LLM runner for text and storytelling
Replaces ChatGPT
Hugging Face toolkit for production LLM serving
Replaces OpenAI API
Browser extension to use local AI models on the web
Replaces ChatGPT
Native Ollama client for iOS and macOS
Replaces ChatGPT
Self-hosted LLaMA chat UI with no API keys needed
Replaces ChatGPT
Desktop app to discover, download, and run local LLMs
Replaces ChatGPT
Containerized LLM toolkit to run a local AI stack with one CLI
Replaces OpenAI Platform
Performance-focused fork of llama.cpp with new quant types
Replaces OpenAI API
Run LLM inference directly in the browser with WebAssembly
Replaces OpenAI API
No apps match these filters.
Every option here is open-source and self-hostable, tagged Offline-First within the local llm runners category. Compare them on the individual app pages for setup difficulty and resource requirements.