Skip to content
AI Compute Radar
IndexCollected 23 min ago

01 / Stack builder

What are you building?

Five questions. A complete stack with a reason behind every slot — assembled from rules, not by an LLM, so the same answers always give the same result.

Saved only in this browser (localStorage), never sent to us. Untick to forget.

Assembled deterministically from the same measured data as the rest of the site: local fit from your hardware and published model sizes, prices from OpenRouter, rental rates from the Vast.ai marketplace. No LLM decides this ranking.

Hybrid stack

Cloud budget: $50/mo buys about 234M input tokens/mo at Qwen3.8 27B

Hybrid stack. Main brain: Qwen3.8 27B. Fast + cheap: LFM2.5 2.6B. Local model: GLM-4.7-Flash. Runtime: llama.cpp. Cloud router: OpenRouter.

Main brain

Medium cost

Qwen3.8 27B

$0.21/M input tokens

The most-adopted model for this use case among those with a published price. We publish no quality benchmarks, so this reflects usage, not a capability ranking.

OpenRouter ↗ (External link)

Fast + cheap

No per-token cost

LFM2.5 2.6B

$0.00/M input tokens

Cheaper than your main brain, for classification, routing and other high-volume calls that do not need the strongest model.

OpenRouter ↗ (External link)

Local model

No per-token cost

GLM-4.7-Flash

Q4_K_M · ~18.9 GBMemory fit 73/100

Runs on hardware you already own, so everyday work costs nothing per token and never leaves your machine.

Hugging Face ↗ (External link)

Runtime

No per-token cost

llama.cpp

GGUF weights run under llama.cpp without a packaged runtime.

Cloud router

No cost

OpenRouter

One API across providers, so swapping a model does not become a migration. The prices on this site come from its public catalogue.

OpenRouter ↗ (External link)