WL

Wllama

Run LLM inference directly in the browser with WebAssembly

Local LLM Runners ★ 1.2k stars Medium setup MIT

Wllama is an open-source WebAssembly binding for llama.cpp that allows large language models to run entirely inside the browser. It can be hosted as a static site to provide fully client-side AI inference.

Key features

  • Browser-based inference
  • WebAssembly powered
  • No server needed
  • Static site deployable

Strengths

  • Released under the MIT license
  • Active community (1.2k GitHub stars)
  • Written in TypeScript
  • Lightweight — runs in 512 MB RAM

Wllama replaces

Compare Wllama

10 head-to-head comparisons.

Similar local llm runners apps