Side by side
Compare runtimes
Send one prompt to two or three of your installed models at once and watch the replies and timings side by side.
Setup
Comparison setup
Up to 3 run at once, which can hold several models in memory at the same time. wllama uses WebGPU when it can and reports which backend each lane actually used.
Pick at least two. Install more from the Models page if you need them.
Adapter evidence
Structured output and speculative decoding
WebAI only calls something supported if the app can turn it on and the difference shows up in a repeatable A/B run. Whatever the browser does internally doesn't count.