TE

Text Embeddings Inference

Fast inference server for text embedding models

Local LLM Runners ★ 5k stars Medium setup Apache-2.0

Text Embeddings Inference is a toolkit for deploying and serving open-source text embeddings and sequence classification models. It provides high-performance extraction with dynamic batching and an OpenAI-compatible route.

Key features

  • Dynamic batching
  • GPU and CPU support
  • OpenAI-compatible endpoint
  • Low-overhead serving

Pros & cons

Strengths

  • Released under the Apache-2.0 license
  • First-class Docker support for quick deployment
  • Active community (5k GitHub stars)
  • Written in Rust

Trade-offs

  • Embeddings only
  • Model compatibility limits

Text Embeddings Inference replaces

Compare Text Embeddings Inference

25 head-to-head comparisons.

Similar local llm runners apps