Wllama
Run LLM inference directly in the browser with WebAssembly
Wllama is an open-source WebAssembly binding for llama.cpp that allows large language models to run entirely inside the browser. It can be hosted as a static site to provide fully client-side AI inference.
Key features
- Browser-based inference
- WebAssembly powered
- No server needed
- Static site deployable
Strengths
- Released under the MIT license
- Active community (1.2k GitHub stars)
- Written in TypeScript
- Lightweight — runs in 512 MB RAM
Wllama replaces
Compare Wllama
10 head-to-head comparisons.
Similar local llm runners apps
Ollama
Local LLM RunnersRun large language models locally with a simple CLI and API
Replaces ChatGPT, OpenAI API
llama.cpp
Local LLM RunnersHigh-performance LLM inference in plain C/C++
Replaces OpenAI API
vLLM
Local LLM RunnersHigh-throughput LLM serving engine with PagedAttention
Replaces OpenAI API
GPT4All
Local LLM RunnersPrivacy-first desktop chat with local language models
Replaces ChatGPT
LiteLLM
Local LLM RunnersUnified proxy and gateway for 100+ LLM APIs
Replaces OpenRouter
exo
Local LLM RunnersRun your own AI cluster across everyday devices
Replaces OpenAI API