Best local models

The best open-weights models in each GPT class — what they cost in the cloud, and what they demand in RAM

← All models

Prices are live from the OpenRouter API, USD per 1M tokens. Local-run estimates assume 4-bit quantization (≈0.6 GB per billion parameters), Macs can dedicate ~75% of unified memory and GPUs ~90% of VRAM; GPT-equivalent tiers are an editorial call, not a benchmark.