Petals
Run large language models collaboratively in a swarm
Petals lets you run and fine-tune large language models in a distributed, BitTorrent-style network where each participant hosts part of the model. It enables self-hosting models too large for a single machine.
Key features
- Distributed model hosting
- Run models larger than one GPU
- Collaborative inference swarm
- Fine-tuning support
Strengths
- Released under the MIT license
- First-class Docker support for quick deployment
- Mature project with 10.6k GitHub stars
- Written in Python
Petals replaces
Compare Petals
15 head-to-head comparisons.
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
New API
Local LLM RunnersNext-gen LLM gateway and AI asset management system
Replaces OpenRouter, OpenAI API