26B A4B or 31B?
Use the model card and your hardware to decide. 26B A4B can be lighter to run because only part of the model is active. 31B usually asks more of memory. Do not compare the numbers as if they describe the same architecture.
Gemma 4 × SillyTavern
Trying Gemma 4 for character chat? First separate the official model names from community quant files. Then check the backend, chat template, tokenizer and thinking markers in that order. This guide gives you a short route to a working reply, plus a browser option when local setup is the part slowing you down.
Verify names at Google DeepMind ↗
Start with the model
Google lists E2B, E4B, 12B, 26B A4B and 31B. The 26B A4B model is mixture-of-experts; 31B is dense. A Q4 or Q6 file is a quantized release, not a new official model.
Use the model card and your hardware to decide. 26B A4B can be lighter to run because only part of the model is active. 31B usually asks more of memory. Do not compare the numbers as if they describe the same architecture.
Ollama, LM Studio, KoboldCpp and a llama.cpp server take different routes through templates and context. An OpenRouter or other API route may be simpler, but the exact model ID and alias belong to that provider.
Character consistency depends on the card, system prompt, sampler and context budget as much as the checkpoint. Community reports mention less repetition with DRY and presets, but those are starting points, not official defaults.
SillyTavern path
Change one layer at a time. That makes a bad reply diagnosable instead of turning every setting into a guess.
Pick the right connector for your backend and confirm that it lists the exact Gemma 4 checkpoint. If you use a provider, copy its model slug rather than typing a remembered alias.
For an API or local server that supports it, start with Chat Completions. Text Completion can work, but it exposes more template details and makes a stale format easier to miss.
Use the Gemma 4 or Gemini tokenizer/template supplied by your backend. Older Gemma templates can put headers and reasoning markers in the wrong place.
Load the character card, then add a restrained system prompt. Try one community preset at a time. Keep the original files with their author and record what you changed.
The names FF, Moonlight and other presets come from community discussions. They are not Google defaults and this page does not republish their files.
When the reply looks wrong
Check the model template, tokenizer and thinking setting first. Some Reddit replies suggest a <|think|> marker in the system prompt, but that depends on the backend and can create literal tokens when the format is wrong.
A provider may expose reasoning differently from a local build. Do not force a marker blindly. Compare one short request with thinking on and off, then read the provider or model-card instructions.
Lower the context noise before changing temperature. Trim duplicated card text, check the context limit, then try a community sampler or DRY setting. Reddit users often mention Temperature 1.0, Top-K 64 and Top-P 0.95, but treat these as experiments.
A browser route
Tabbit does not replace SillyTavern cards or lorebooks. It solves a different part of the problem: reading a live wiki, document or reference page while chatting with a model already available in the browser.
Download the Chromium-based browser for macOS or Windows. No local model file or SillyTavern server is required for this route.
Open the model picker after install and select a current model shown there, such as GPT-5.4, GPT-5.2-Chat or Gemini-3-Flash. Availability can change, so the picker is the source of truth.
Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

Pick the right tool
Use the one that matches the bottleneck. A browser shortcut is not a character-card manager, and a local frontend is not automatically a better model.
| SillyTavern | Tabbit | |
|---|---|---|
| Character cards and lorebooks | Built in | Keep the source page open |
| Gemma 4 local checkpoint | Bring your own backend | Use a model shown in the picker |
| Template and sampler control | Detailed | Managed by the selected model |
| Page context | Paste or use extensions | @ the live tab or file |
| First useful reply | Install + endpoint + preset | Install + choose a listed model |

Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

@ the live tab or file
FAQ
No. Google describes 26B A4B as a mixture-of-experts model and 31B as dense. Choose from the model card and hardware you actually have.
There is no universal best file. Match Q4, Q5 or Q6 to available memory, context needs and the publisher’s model card. A quant file is not an official new Gemma name.
The tokenizer, template, backend or provider may disagree about the format. Verify those pieces before adding markers such as <|think|>; a marker in the wrong template can make the output worse.
Do not assume it. Open Tabbit’s current model picker after installation and use a model that is actually listed there. The screenshot on this page does not show Gemma 4.
No. SillyTavern remains the better fit for cards, lorebooks and detailed preset control. Tabbit is for chatting beside live web pages and files.
If you need Gemma 4 locally, use the model card and check the template chain. If you need an answer beside a character wiki now, install Tabbit and choose a model from its current picker.
Available for macOS and Windows. The model list can change.