Workbench
Chat
Talk to a local model and see what it costs: every reply reports the timings, tokens, cache, and context its runtime actually exposes.
Loading your downloaded models.
Conversation
Chat
Load a model to begin
Every reply comes with load time, TTFT, and prefill and decode speed.