Petals
Run large language models collaboratively in a swarm
Petals lets you run and fine-tune large language models in a distributed, BitTorrent-style network where each participant hosts part of the model. It enables self-hosting models too large for a single machine.
Key features
- Distributed model hosting
- Run models larger than one GPU
- Collaborative inference swarm
- Fine-tuning support
Strengths
- Released under the MIT license
- First-class Docker support for quick deployment
- Mature project with 10.5k GitHub stars
- Written in Python
Petals replaces
Compare Petals
34 head-to-head comparisons.
- Petals vs Ollama
- Petals vs Hugging Face Transformers
- Petals vs llama.cpp
- Petals vs vLLM
- Petals vs GPT4All
- Petals vs GPT4Free
- Petals vs LiteLLM
- Petals vs LocalAI
- Petals vs exo
- Petals vs New API
- Petals vs Jan
- Petals vs FastChat
- Petals vs One API
- Petals vs Continue
- Petals vs SGLang
- Petals vs llamafile
- Petals vs MLC LLM
- Petals vs Guidance
- Petals vs OpenLLM
- Petals vs KoboldCpp
- Petals vs Text Generation Inference
- Petals vs Xinference
- Petals vs Llama Stack
- Petals vs Page Assist
- Petals vs LMDeploy
- Petals vs MLX LM
- Petals vs Enchanted
- Petals vs Serge
- Petals vs GPUStack
- Petals vs LM Studio
- Petals vs Text Embeddings Inference
- Petals vs Harbor LLM Toolkit
- Petals vs ik_llama.cpp
- Petals vs Aphrodite Engine
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API