Side by side

Compare runtimes

Send one prompt to two or three of your installed models at once and watch the replies and timings side by side.

Setup

Comparison setup

Structured output
What to compare

Up to 3 run at once, which can hold several models in memory at the same time. wllama uses WebGPU when it can and reports which backend each lane actually used.

Pick at least two. Install more from the Models page if you need them.

Adapter evidence

Structured output and speculative decoding

WebAI only calls something supported if the app can turn it on and the difference shows up in a repeatable A/B run. Whatever the browser does internally doesn't count.