Find, download, and manage the models this browser can run. Downloads pin to a specific version, pick up where they left off, and only install once the checksums match.
Loading your models and unfinished downloads.
Quick start
Starter models
One model we've tested for each runtime. Each one is pinned to a specific Hugging Face version and downloads exactly like any other model.
wllama 3.5.1
Llama 3.2 1B Instruct · Q4_K_M
A small Llama chat model that runs on wllama, on CPU or WebGPU. It starts at an 8K context to keep memory sane; Chat lets you go up to its full 128K.
Search Hugging Face for GGUF models. WebAI checks that each one is public and downloadable, and shows its exact size, before anything starts downloading.
Search and filtersAll GGUF models · 2 active filters
From a repository
Paste a Hugging Face model
Enter owner/model, add @revision, or paste a model or file URL. WebAI pins it to a specific version before listing files, and only works with public, ungated repos.
From this device
Import model files
Pick a GGUF, ONNX, or LiteRT-LM file, or a whole set of GGUF shards. WebAI checks the stored copy and reads its metadata in the background.
Drop GGUF, ONNX, or LiteRT-LM files here or choose them from this device.