Model library

Models

Find, download, and manage the models this browser can run. Downloads pin to a specific version, pick up where they left off, and only install once the checksums match.

Loading your models and unfinished downloads.

Quick start

Starter models

One model we've tested for each runtime. Each one is pinned to a specific Hugging Face version and downloads exactly like any other model.

wllama 3.5.1

Llama 3.2 1B Instruct · Q4_K_M

A small Llama chat model that runs on wllama, on CPU or WebGPU. It starts at an 8K context to keep memory sane; Chat lets you go up to its full 128K.

770.3 MiB · commit 7d1f70022fca

License: llama3.2 · access: public/open

See the files and license on Hugging Face

Transformers.js 4.2.0 · wasm

SmolLM2 135M Instruct · wasm

Runs anywhere via ONNX Runtime on wasm. It only uses threads when the page is cross-origin isolated.

175.6 MiB · commit 12fd25f77366

License: apache-2.0 · access: public/open

See the files and license on Hugging Face

Transformers.js 4.2.0 · webgpu

SmolLM2 135M Instruct · webgpu

Runs on WebGPU with a q4/f16 graph. Needs WebGPU to work inside a worker.

114.3 MiB · commit 12fd25f77366

License: apache-2.0 · access: public/open

See the files and license on Hugging Face

Transformers.js 4.2.0 · webnn

SmolLM2 135M Instruct · webnn

Experimental WebNN GPU path with an fp16 graph. You only find out whether every operator is supported when the session starts.

259.8 MiB · commit 12fd25f77366

License: apache-2.0 · access: public/open

See the files and license on Hugging Face

LiteRT-LM 0.14.0

Gemma 4 E2B · LiteRT-LM WebGPU

Google's early-preview web build. It starts at the pack's measured 2K context and lets you go up to its documented 32K maximum.

1.9 GiB · commit 9262660a1676

License: gemma · access: public/open

See the files and license on Hugging Face

Find models

Browse Hugging Face models

Search Hugging Face for GGUF models. WebAI checks that each one is public and downloadable, and shows its exact size, before anything starts downloading.

Search and filtersAll GGUF models · 2 active filters

From a repository

Paste a Hugging Face model

Enter owner/model, add @revision, or paste a model or file URL. WebAI pins it to a specific version before listing files, and only works with public, ungated repos.

This only reads the file list. Nothing downloads until you pick a file.

From this device

Import model files

Pick a GGUF, ONNX, or LiteRT-LM file, or a whole set of GGUF shards. WebAI checks the stored copy and reads its metadata in the background.

Drop GGUF, ONNX, or LiteRT-LM files here
or choose them from this device.

On this device

Your models

Checking what's in browser storage…