Skip to content
AI Compute Radar
IndexCollected 3 h ago

Gate: the smallest device with an EXCELLENT fit at 32K context. Seat: the measured GGUF file.

Pick of the week — Week 40, 2026 · selected by rule, written by people

Five gigabytes of weights, context at full price

Granite 4.2 8B is this week's pick by rule: the largest counted Heat Score rise among tracked models that run comfortably on a consumer card, from 35 to 57 since the evening of 26 September. Its trending score climbed back from 1 to 8 in six days, its measured Q4_K_M file is 5.0 GB, and it fits a 12 GB card at 8K context with 3.3 GB to spare.

Measured signals as of 2026-10-02 18:00 UTC · Hugging Face + Vast.ai · method on /methodology

Where it runs

Fit engine · measured GGUF + computed KV cache · usable memory after margins
Device8K context32K context
RTX 30708 GB · usable 7 GBOFFLOAD REQUIRED · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 40608 GB · usable 7 GBOFFLOAD REQUIRED · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 50608 GB · usable 7 GBOFFLOAD REQUIRED · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 3060 Ti8 GB · usable 7 GBOFFLOAD REQUIRED · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 308010 GB · usable 8.8 GBGOOD · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 2080 Ti11 GB · usable 9.7 GBGOOD · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 407012 GB · usable 10.6 GBEXCELLENT · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 3060 12GB12 GB · usable 10.6 GBEXCELLENT · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 507012 GB · usable 10.6 GBEXCELLENT · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 4070 Super12 GB · usable 10.6 GBEXCELLENT · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 4070 Ti12 GB · usable 10.6 GBEXCELLENT · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 3080 Ti12 GB · usable 10.6 GBEXCELLENT · 7.3 GBOFFLOAD REQUIRED · 11.8 GB
RTX 408016 GB · usable 14.1 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
RTX 508016 GB · usable 14.1 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
RTX 4060 Ti 16GB16 GB · usable 14.1 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
RTX 5060 Ti 16GB16 GB · usable 14.1 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
RTX 5070 Ti16 GB · usable 14.1 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
Mac mini M6 16GB16 GB · usable 10.4 GBEXCELLENT · 7.3 GBNOT RECOMMENDED · 11.8 GB
RTX 4070 Ti Super16 GB · usable 14.1 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
RX 9070 XT16 GB · usable 14.1 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
RTX 309024 GB · usable 21.1 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
RTX 409024 GB · usable 21.1 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
Mac mini M5 Pro 24GB24 GB · usable 15.6 GBEXCELLENT · 7.3 GBGOOD · 11.8 GB
RX 7900 XTX24 GB · usable 21.1 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
RTX 3090 Ti24 GB · usable 21.1 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
RTX 509032 GB · usable 28.2 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
Mac Studio M5 Max 36GB36 GB · usable 23.4 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
Mac mini M4 Pro 64GB64 GB · usable 44.8 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
Mac Studio M4 Max 64GB64 GB · usable 44.8 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
Mac Studio M3 Ultra 96GB96 GB · usable 67.2 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
Mac Studio M5 Ultra 96GB96 GB · usable 67.2 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB
NVIDIA DGX Spark128 GB · usable 112.6 GBEXCELLENT · 7.3 GBEXCELLENT · 11.8 GB

Or rent an RTX 4090 from $0.46/h →Check it against your own machine →

What it is for

Granite 4.2 8B is an 8.8-billion-parameter model from IBM with open weights, curated here for general, coding and agent work. The repository dates from 7 August 2026, IBM's own Granite organisation publishes the GGUF build we measure, and Ollama carries it in its library as granite4.2. On 2 October it stood at 127,613 downloads over the last 30 days, up 16 % week on week, with 92 likes on Hugging Face and a trending score of 8.

The move is a rebound. The Heat Score had slid from 50 on 20 September to 35 on the 26th while the trending score fell to 1; then it turned: 3, 4, 5, 5, 8 and 8 on the following evenings, with downloads going from 112,993 to 127,613 inside the window. That took the score to 57. The whole family moved in the same days: Granite 4.2 3B rose 12 points to 46, short of the 50 a model needs for its rise to count, and Granite 4.2 30B rose 5 to 39 but fits no consumer card. The runners-up were Qwen3.8 27B and LFM2.5 2.6B with a rise of 11 each. GLM-5.3 gained 13 points to 86 and Gemma 4 31B 8 to 84, but neither runs comfortably on a card of up to 24 GB. Last week's pick, Mistral Small 3.2, sits at 45, seven points below where it was picked.

The weights are the small part. The Q4_K_M file is 5.0 GB, but all 40 layers keep a full attention cache, so context is charged at full price: 1.25 GB for every 8K tokens, computed from the published attention shape. With runtime overhead the engine puts the total at 7.3 GB at 8K and 11.8 GB at 32K, where the cache is as large as the weights. That is EXCELLENT at 8K on every 12 GB card, from the RTX 3060 12GB to the RTX 4070 Ti, with 3.3 GB to spare, and GOOD on the RTX 3080 and the RTX 2080 Ti. At 32K the 12 GB cards are 1.2 GB short and have to offload; the 16 GB cards carry it with 2.3 GB to spare; the 24 GB cards, the RTX 5090, the DGX Spark and every Mac from 24 GB up are comfortable at both lengths. Of the 32 machines we track, 28 run it comfortably at 8K and 19 at 32K. The model accepts up to 131,072 tokens of context, which would take 29.8 GB. Renting instead of buying: an RTX 4090 cost $0.46 per hour at the median of verified Vast.ai offers on 2 October at 18:00 UTC.

Maker
IBM
Parameters
8.8B
Measured GGUF
Q4_K_M · 5 GB
Max context
128K tokens
Curated for
general, coding, agents
Repo created
2026-08-07
Trending score
8
7-day downloads
+15.7%

Model pageHugging Face repo (External link)GGUF file (External link)

Downloads, daily readingsdaily last · UTC

2026-09-19 · 91K2026-10-02 · 128K

Watch out

  • Not an 8 GB card model at this quantization: 7.3 GB at 8K against 7.0 GB usable means the RTX 3070, RTX 3060 Ti, RTX 4060 and RTX 5060 have to offload part of the weights into system RAM. It runs, but not in the way a fit verdict means it.
  • Long context is where the memory goes: 32K takes every 12 GB card and the Mac mini M6 16GB out of the comfortable range, 64K needs 17.8 GB and 128K needs 29.8 GB.
  • The rise starts from a trough. The Heat Score stood at 51 on 7 September, peaked at 65 on the 11th and slid to 35 by the 26th; this week brought it back to 57, not to new ground.
  • The numbers behind the move are small: a trending score that went from 1 to 8, and 92 likes. All three Granite 4.2 models rose in the same days; our readings do not say why, and the rule does not ask.
  • One measured GGUF, Q4_K_M from IBM; other quantizations change every number above. No benchmarks and no speed figures: whether it is good at your task is yours to test.
  • The rental price is one marketplace at one hour; it moves during the day.

How we chose it

Pick of the week is the model with the largest Heat Score rise over the last seven days among tracked models that run comfortably (GOOD or EXCELLENT) on at least one consumer card of up to 24 GB at 8K context. A rise counts only when it is at least three points, the model's Heat Score is at least 50 today, and the model's own direction is not falling — a low score that climbs while the models around it fall has only moved because they moved. Ties break by the higher Heat Score, then by 30-day downloads. A model cannot be picked again within four weeks. While the heat history is younger than a week, the rise is measured over the readings that exist and the report says so; when no rise counts, the rule falls back to the highest Heat Score and records that.

This issue was selected on the heat rise. Rise measured since 2026-09-26 18:15 UTC. 46 scored models were considered; the full list with reasons is stored with the issue.

The rule on the methodology page · rule v1.2 · selected 2026-10-02 18:28 UTC

Issues

A new pick every Friday. The model is computed, the text is written and released by a person, and every number links back to where it was measured.

RSS feed of every issue · JSON: /api/v1/pick