What's on your mind today?
Model runs entirely on your device.
How much text the model “sees” at once — your prompt, chat history, and attached files all count toward it.
32K is the default: room for several documents while keeping memory use and prefill time down. Auto uses your device’s maximum (up to 128K tokens); other caps trade memory and speed differently.
1 token ≈ ¾ of a word. 8K ≈ a few pages, 32K ≈ a short report, 128K ≈ a small book.
Controls how creative vs. predictable the model is.
0.2 Precise — deterministic, best for code & facts.
0.7 Balanced — good default for most tasks.
1.0–1.3 Creative — more varied, surprising, but can be less accurate.
Lower = sticks to the most likely words. Higher = explores alternatives. On-device sampling is currently greedy — changes are subtle but noticeable over longer replies.
Model runs entirely on your device.
Kernels are the low-level GPU programs that do the model's actual math — the matrix multiplications, attention, and normalization behind every token. And how well they're optimized can dramatically speed up inference.
gemma4-vision.js) — preprocessing, a 16-layer WGSL transformer, pooling and projection — plus the multimodal extension that injects image features into the LLM were written entirely by DeepSeek V4 Flash as custom WebGPU kernels.