Gate: the smallest device with an EXCELLENT fit at 32K context. Seat: the measured GGUF file.
Pick of the week — Week 38, 2026 · selected by rule, written by people
Twenty-six billion parameters, four at a time
Gemma 4 26B-A4B is this week's pick by rule: the largest counted Heat Score rise among tracked models that run comfortably on a consumer card, from 51 to 75 since the evening of 12 September. Its trending score climbed back from 5 to 20 in six days, its measured Q4_K_M file is 15.8 GB, and it fits a 24 GB card at 8K context with 3.5 GB to spare.
Measured signals as of 2026-09-18 18:00 UTC · Hugging Face + Vast.ai · method on /methodology
Where it runs
Fit engine · measured GGUF + computed KV cache · usable memory after marginsOr rent an RTX 4090 from $0.53/h →Check it against your own machine →
What it is for
This is the mixture-of-experts member of Google's Gemma 4 family: 25.8 billion parameters in total, about four billion of them active for any given token, which is what the A4B in the name means. The repository dates from 11 March 2026, the weights are open, Unsloth publishes the GGUF builds we measure, and Ollama carries it in its library. On 18 September it stood at 9.8 million downloads over the last 30 days, up 9 % week on week, with 1,526 likes on Hugging Face and an API price of $0.09 per million input tokens on OpenRouter.
The rise is a recovery, not a launch. The trending score had slid from 20 on 5 September to 5 on the 12th, then turned: 8, 9, 14, 20, 18 and 20 on the following evenings, while downloads went from 9.05 to 9.79 million inside the window. That took the Heat Score from 51 to 75. The runners-up were Qwen3 8B with a rise of 15 (62 to 77), Qwen3 0.6B with 10 and gpt-oss 20B with 9. Qwen3.8 27B carries a higher score at 78 but fell 14 points in the window, so its move did not count.
Memory is the whole question with a mixture-of-experts model: only a fraction of the parameters works on each token, but all of them have to sit in memory. At Q4_K_M that is 15.8 GB of weights, plus a context cache that the model's published configuration keeps small, 0.5 GB at 8K tokens and 1.5 GB at 32K, for 17.6 GB in total at 8K including runtime overhead. A 24 GB card such as the RTX 3090 or RTX 4090 runs it with room at 8K and gets tight at 32K, 19.3 GB against 21.1 GB usable; the RTX 5090 and any 64 GB Mac run it comfortably at both, and the model accepts up to 262,144 tokens of context on machines with the memory for it. Renting instead of buying: an RTX 4090 cost $0.53 per hour at the median of verified Vast.ai offers on 18 September at 18:00 UTC.
- Maker
- Parameters
- 25.8B MoE
- Measured GGUF
- Q4_K_M · 15.8 GB
- Max context
- 256K tokens
- Curated for
- general, research
- Repo created
- 2026-03-11
- Trending score
- 20
- 7-day downloads
- +9%
Model pageHugging Face repo (External link)GGUF file (External link)
Downloads, daily readingsdaily last · UTC
2026-09-05 · 8.3M2026-09-18 · 9.8M
Watch out
- Not a 12 GB or 16 GB card model: 17.6 GB at 8K means everything below 24 GB has to offload part of the weights into system RAM. It runs, but not in the way a fit verdict means it.
- The rise starts from a trough. A Heat Score of 51 on 12 September was this model's low point in our record; the trend recovered, it did not break out.
- "Four billion active" changes nothing about the download or the memory: the file is 15.8 GB and all of it is loaded.
- 64K and 128K context need 21.6 GB and 26.1 GB, beyond any 24 GB card.
- The rental price is one marketplace at one hour; it moves during the day.
How we chose it
Pick of the week is the model with the largest Heat Score rise over the last seven days among tracked models that run comfortably (GOOD or EXCELLENT) on at least one consumer card of up to 24 GB at 8K context. A rise counts only when it is at least three points, the model's Heat Score is at least 50 today, and the model's own direction is not falling — a low score that climbs while the models around it fall has only moved because they moved. Ties break by the higher Heat Score, then by 30-day downloads. A model cannot be picked again within four weeks. While the heat history is younger than a week, the rise is measured over the readings that exist and the report says so; when no rise counts, the rule falls back to the highest Heat Score and records that.
This issue was selected on the heat rise. Rise measured since 2026-09-12 18:15 UTC. 45 scored models were considered; the full list with reasons is stored with the issue.
The rule on the methodology page · rule v1.2 · selected 2026-09-18 18:24 UTC
Issues
- Week 38, 2026 · Gemma 4 26B-A4B2026-09-18
- Week 37, 2026 · Gemma 4 12B2026-09-11
A new pick every Friday. The model is computed, the text is written and released by a person, and every number links back to where it was measured.