Workbench

Chat

Talk to a local model and see what it costs: every reply reports the timings, tokens, cache, and context its runtime actually exposes.

Loading your downloaded models.

Conversation

Chat

No session

Load a model to begin

Every reply comes with load time, TTFT, and prefill and decode speed.

Sent with every prompt. Models whose chat template supports it will honor it; others ignore it.

Timings come from wllama, plus TTFT measured on the page.