TR

Triton Inference Server

High-performance inference serving for any model framework

Self-Hosted AI ★ 10.9k stars Hard setup BSD-3-Clause

Triton Inference Server is an open-source serving system from NVIDIA that runs models from TensorFlow, PyTorch, ONNX, TensorRT, and more. It supports concurrent execution, dynamic batching, and CPU or GPU deployment.

Key features

  • Multi-framework model serving
  • Dynamic request batching
  • Concurrent model execution
  • GPU and CPU support

Strengths

  • Released under the BSD-3-Clause license
  • First-class Docker support for quick deployment
  • Kubernetes-ready with Helm charts available
  • Mature project with 10.9k GitHub stars

Triton Inference Server replaces

Compare Triton Inference Server

11 head-to-head comparisons.

Similar self-hosted ai apps