Xinference
Distributed inference framework for LLMs and embeddings
Xorbits Inference (Xinference) is a framework for serving language, embedding, image, audio, and rerank models with a single command. It exposes OpenAI-compatible APIs and supports distributed deployment across multiple machines.
Key features
- Serve LLMs, embeddings and images
- OpenAI-compatible API
- Distributed cluster support
- Built-in model registry
Strengths
- Released under the Apache-2.0 license
- First-class Docker support for quick deployment
- Kubernetes-ready with Helm charts available
- Mature project with 9.5k GitHub stars
Xinference replaces
Compare Xinference
30 head-to-head comparisons.
- Xinference vs OpenClaw
- Xinference vs Hermes Agent
- Xinference vs OpenCode
- Xinference vs Ollama
- Xinference vs Hugging Face Transformers
- Xinference vs Langflow
- Xinference vs Dify
- Xinference vs llama.cpp
- Xinference vs vLLM
- Xinference vs PaddleOCR
- Xinference vs RAGFlow
- Xinference vs OpenHands
- Xinference vs screenshot-to-code
- Xinference vs GPT4Free
- Xinference vs LocalAI
- Xinference vs exo
- Xinference vs New API
- Xinference vs FastChat
- Xinference vs One API
- Xinference vs SGLang
- Xinference vs llamafile
- Xinference vs MLC LLM
- Xinference vs Guidance
- Xinference vs OpenLLM
- Xinference vs Text Generation Inference
- Xinference vs Petals
- Xinference vs Llama Stack
- Xinference vs LMDeploy
- Xinference vs MLX LM
- Xinference vs GPUStack
Similar self-hosted ai apps
OpenClaw
Self-Hosted AIThe AI that actually does things
Hermes Agent
Self-Hosted AIThe AI agent that grows with you
OpenCode
Self-Hosted AIThe open source AI coding agent
Hugging Face Transformers
Self-Hosted AIState-of-the-art machine learning model library
Replaces OpenAI API
Langflow
Self-Hosted AIVisual framework for building AI agents and RAG pipelines
Replaces Vertex AI Agent Builder
Dify
Self-Hosted AIOpen-source platform for building production LLM apps
Replaces OpenAI Assistants, Vertex AI Agent Builder