Gemma 4 31B × SillyTavern

Gemma 4 31B in SillyTavern

The 31B query hides the first trap: Google calls this a dense Gemma 4 model, while 26B A4B is a different MoE model. This guide maps the name, memory, quantization, provider and chat template before you tune a character card.

Official Gemma 4 page
Verify the official naming and template
Tabbit desktop browser with vertical tabs, a centered prompt field and a chat panel on the right.

Name it correctly

31B means dense, not 26B A4B

Google lists Gemma 4 31B as a dense model with about 30.7B parameters and a 256K context window. Gemma 4 26B A4B has about 25.2B total parameters but about 3.8B active parameters because it is MoE.

01

Use the official checkpoint

For the instruction variant, start from a provider or repository that identifies it as Gemma 4 31B IT. GGUF names from Unsloth, bartowski or LM Studio Community are quantized distributions, not new Google model families.

02

Budget for dense inference

A Q4 file is smaller than full precision, but the model still has 31B active parameters. Community reports put Q4_K_M around 18 GB for weights, before context, runtime overhead and any multimodal files. Treat that as a planning figure, not a guarantee.

03

Base and IT are different tests

SillyTavern users report strong creative writing from both base and instruct files, but they need different prompts and templates. Decide whether you want instruction following or a base checkpoint before comparing presets.

SillyTavern path

A clean connection order

Change one layer at a time. That makes a bad reply diagnosable instead of turning every setting into a guess.

  1. 01

    1. Confirm the endpoint

    Pick the right connector for your backend and confirm that it lists the exact Gemma 4 checkpoint. If you use a provider, copy its model slug rather than typing a remembered alias.

  2. 02

    2. Choose Chat Completions

    For an API or local server that supports it, start with Chat Completions. Text Completion can work, but it exposes more template details and makes a stale format easier to miss.

  3. 03

    3. Match the tokenizer and template

    Use the Gemma 4 or Gemini tokenizer/template supplied by your backend. Older Gemma templates can put headers and reasoning markers in the wrong place.

  4. 04

    4. Add your card and preset

    Load the character card, then add a restrained system prompt. Try one community preset at a time. Keep the original files with their author and record what you changed.

The names FF, Moonlight and other presets come from community discussions. They are not Google defaults and this page does not republish their files.

When the reply looks wrong

Fix the layer that failed

CASE 01

Thinking text will not parse

Check the model template, tokenizer and thinking setting first. Some Reddit replies suggest a <|think|> marker in the system prompt, but that depends on the backend and can create literal tokens when the format is wrong.

CASE 02

It thinks only sometimes

A provider may expose reasoning differently from a local build. Do not force a marker blindly. Compare one short request with thinking on and off, then read the provider or model-card instructions.

CASE 03

The character repeats or drifts

Lower the context noise before changing temperature. Trim duplicated card text, check the context limit, then try a community sampler or DRY setting. Reddit users often mention Temperature 1.0, Top-K 64 and Top-P 0.95, but treat these as experiments.

A browser route

Keep the character page open and ask beside it

Tabbit does not replace SillyTavern cards or lorebooks. It solves a different part of the problem: reading a live wiki, document or reference page while chatting with a model already available in the browser.

  1. 1

    Install Tabbit

    Download the Chromium-based browser for macOS or Windows. No local model file or SillyTavern server is required for this route.

  2. 2

    Choose what is actually listed

    Open the model picker after install and select a current model shown there, such as GPT-5.4, GPT-5.2-Chat or Gemini-3-Flash. Availability can change, so the picker is the source of truth.

  3. 3

    Reference the live page

    Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

Tabbit new-tab model picker showing GPT-5.4, GPT-5.2-Chat, Gemini-3.1-Pro, Gemini-3-Flash and Claude-Sonnet-4.6. Gemma 4 is not visible in this screenshot.

Pick the right tool

SillyTavern or Tabbit?

Use the one that matches the bottleneck. A browser shortcut is not a character-card manager, and a local frontend is not automatically a better model.

SillyTavernTabbit
Character cards and lorebooksBuilt inKeep the source page open
Gemma 4 local checkpointBring your own backendUse a model shown in the picker
Template and sampler controlDetailedManaged by the selected model
Page contextPaste or use extensions@ the live tab or file
First useful replyInstall + endpoint + presetInstall + choose a listed model
Tabbit new-tab model picker showing GPT-5.4, GPT-5.2-Chat, Gemini-3.1-Pro, Gemini-3-Flash and Claude-Sonnet-4.6. Gemma 4 is not visible in this screenshot.

Reference the live page

Keep the character wiki or setting document in a tab. Use @ to reference the page, a screenshot or a file, then ask for a summary, scene notes or a second take.

Tabbit new-tab model picker showing GPT-5.4, GPT-5.2-Chat, Gemini-3.1-Pro, Gemini-3-Flash and Claude-Sonnet-4.6. Gemma 4 is not visible in this screenshot.

Page context

@ the live tab or file

FAQ

Gemma 4 and SillyTavern, answered

Is Gemma 4 26B the same as 31B?+

No. Google describes 26B A4B as a mixture-of-experts model and 31B as dense. Choose from the model card and hardware you actually have.

Which Gemma 4 quantization should I download?+

There is no universal best file. Match Q4, Q5 or Q6 to available memory, context needs and the publisher’s model card. A quant file is not an official new Gemma name.

Why do thinking tokens show as plain text?+

The tokenizer, template, backend or provider may disagree about the format. Verify those pieces before adding markers such as <|think|>; a marker in the wrong template can make the output worse.

Can Tabbit run Gemma 4?+

Do not assume it. Open Tabbit’s current model picker after installation and use a model that is actually listed there. The screenshot on this page does not show Gemma 4.

Does Tabbit replace SillyTavern?+

No. SillyTavern remains the better fit for cards, lorebooks and detailed preset control. Tabbit is for chatting beside live web pages and files.

Stop debugging the wrong layer

If you need Gemma 4 locally, use the model card and check the template chain. If you need an answer beside a character wiki now, install Tabbit and choose a model from its current picker.

Available for macOS and Windows. The model list can change.

© 2026 Tabbit Browser. The AI-native browser that understands your context.