Gate: the smallest device with an EXCELLENT fit at 32K context. Seat: the measured GGUF file.
Pick of the week — Week 39, 2026 · selected by rule, written by people
Twenty-four billion parameters, fifteen months on
Mistral Small 3.2 is this week's pick by rule: the largest counted Heat Score rise among tracked models that run comfortably on a consumer card, from 27 to 52 since the evening of 19 September. Its 30-day downloads jumped 45 % in a week, its measured Q4_K_M file is 13.3 GB, and it fits a 24 GB card at 8K context with 5 GB to spare.
Measured signals as of 2026-09-25 18:00 UTC · Hugging Face + Vast.ai · method on /methodology
Where it runs
Fit engine · measured GGUF + computed KV cache · usable memory after marginsOr rent an RTX 4090 from $0.59/h →Check it against your own machine →
What it is for
Mistral Small 3.2 is a 24-billion-parameter dense model from Mistral AI with open weights, released on 19 June 2025 and curated here for general and agent work. Unsloth publishes the GGUF build we measure, and Ollama carries it in its library as mistral-small3.2. On 25 September it stood at 198,640 downloads over the last 30 days, up 45 % week on week, with 620 likes on Hugging Face and a trending score of 3.
The move came in one step. Thirty-day downloads sat between 128,000 and 137,000 from 12 to 18 September, then read 185,194 on the evening of the 19th and climbed to 199,655 by the 24th, while the trending score went from 1 and 2 to 3 and 4 over the same days. The Heat Score followed from 27 to 52, and the engine flagged the move as a breakout. The runner-up was its stablemate Devstral Small 2 24B, one point behind with a rise of 24 (32 to 56); then Qwen3 14B with 17, GLM-4.7-Flash with 15, Qwen3 0.6B and LFM2.5 2.6B with 13. gpt-oss 20B and Qwen3.8 27B lost 21 and 36 points inside the window, so their scores did not count.
A dense 24B has no active-parameter trick: all 13.3 GB of its Q4_K_M weights are read for every token. With the context cache computed from the published attention shape and the runtime overhead, the engine puts the total at 15.8 GB at 8K tokens and 20.3 GB at 32K. That is GOOD on the 24 GB cards, RTX 3090, RTX 4090 and RX 7900 XTX, with 5.3 GB to spare at 8K and TIGHT at 32K; EXCELLENT on the RTX 5090 and the Mac Studio M5 Max 36GB at 8K; and comfortable at both context lengths on any 64 GB Mac, the M5 Ultra and the DGX Spark. A 16 GB card is 1.7 GB short at 8K and has to offload part of the weights into system RAM. The model accepts up to 131,072 tokens of context on machines with the memory for it. Renting instead of buying: an RTX 4090 cost $0.59 per hour at the median of verified Vast.ai offers on 25 September at 19:00 UTC.
- Maker
- Mistral AI
- Parameters
- 24B
- Measured GGUF
- Q4_K_M · 13.3 GB
- Max context
- 128K tokens
- Curated for
- general, agents
- Repo created
- 2025-06-19
- Trending score
- 3
- 7-day downloads
- +45.2%
Model pageHugging Face repo (External link)GGUF file (External link)
Downloads, daily readingsdaily last · UTC
2026-09-12 · 131K2026-09-25 · 199K
Watch out
- Not a 16 GB card model: 15.8 GB at 8K against 14.1 GB usable means every 16 GB card, including the new RX 9070 XT and RTX 4070 Ti Super, has to offload. It runs, but not in the way a fit verdict means it.
- The rise starts low and is young: a Heat Score of 27 on 19 September, then one jump of 48,000 downloads in a day and a week of readings since. Real, but one week old.
- The model is fifteen months old. Our readings say people are downloading it again; they do not say why, and the rule does not ask.
- The Mac mini M5 Pro 24GB misses by 0.2 GB: macOS wires only about two thirds of its memory for the GPU, 15.6 GB, against 15.8 GB needed at 8K.
- One measured GGUF, Q4_K_M from unsloth; other quantizations change every number above. No benchmarks and no speed figures: whether it is good at your task is yours to test.
- The rental price is one marketplace at one hour; it moves during the day.
How we chose it
Pick of the week is the model with the largest Heat Score rise over the last seven days among tracked models that run comfortably (GOOD or EXCELLENT) on at least one consumer card of up to 24 GB at 8K context. A rise counts only when it is at least three points, the model's Heat Score is at least 50 today, and the model's own direction is not falling — a low score that climbs while the models around it fall has only moved because they moved. Ties break by the higher Heat Score, then by 30-day downloads. A model cannot be picked again within four weeks. While the heat history is younger than a week, the rise is measured over the readings that exist and the report says so; when no rise counts, the rule falls back to the highest Heat Score and records that.
This issue was selected on the heat rise. Rise measured since 2026-09-19 18:15 UTC. 46 scored models were considered; the full list with reasons is stored with the issue.
The rule on the methodology page · rule v1.2 · selected 2026-09-25 19:30 UTC
Issues
- Week 39, 2026 · Mistral Small 3.22026-09-25
- Week 38, 2026 · Gemma 4 26B-A4B2026-09-18
- Week 37, 2026 · Gemma 4 12B2026-09-11
A new pick every Friday. The model is computed, the text is written and released by a person, and every number links back to where it was measured.