Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
The OpenAI API is the paid programmatic interface to GPT-class models hosted on OpenAI's infrastructure. These open-source apps let you replace OpenAI API with software you host and control yourself.
Run large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
State-of-the-art machine learning model library
Replaces OpenAI API
High-performance LLM inference in plain C/C++
Replaces OpenAI API
High-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
Unified OpenAI-compatible API gateway for many models
Replaces OpenAI API
Drop-in OpenAI-compatible API for local inference
Replaces OpenAI API, ElevenLabs
Run your own AI cluster across everyday devices
Replaces OpenAI API
Next-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API
Platform for serving and evaluating large language models
Replaces OpenAI API
Unified OpenAI-compatible gateway for many LLM providers
Replaces OpenAI API, OpenRouter
Fast serving framework for LLMs and vision-language models
Replaces OpenAI API
Distribute and run LLMs with a single executable file
Replaces OpenAI API, ChatGPT
Universal LLM deployment engine for any hardware
Replaces OpenAI API
Constrained generation language for controlling LLMs
Replaces OpenAI API
Run any open LLM as an OpenAI-compatible API endpoint
Replaces OpenAI API
Hugging Face toolkit for production LLM serving
Replaces OpenAI API
Run large language models collaboratively in a swarm
Replaces OpenAI API
Distributed inference framework for LLMs and embeddings
Replaces OpenAI API, Hugging Face Inference Endpoints
Composable API server for building generative AI applications
Replaces OpenAI API
Toolkit for compressing and serving large language models
Replaces OpenAI API, Hugging Face Inference Endpoints
Run and fine-tune LLMs locally on Apple Silicon with MLX
Replaces OpenAI API
Manage GPU clusters for running AI models
Replaces OpenAI API
Fast inference library for quantized LLMs on consumer GPUs
Replaces OpenAI API
Performance-focused fork of llama.cpp with new quant types
Replaces OpenAI API
Memory-efficient inference library for quantized Llama models
Replaces OpenAI API
High-throughput inference engine for large language models
Replaces OpenAI API
Run LLM inference directly in the browser with WebAssembly
Replaces OpenAI API
OpenAI-compatible text-to-speech server using local models
Replaces OpenAI API, ElevenLabs
No apps match these filters.
Moving off OpenAI API means trading a managed subscription for a service you run yourself. The payoff is ownership: no per-seat pricing, no vendor lock-in, and full control of your data. Start with the option above that matches your setup-difficulty comfort level — every app links to a detailed specs page and head-to-head comparisons.