Find the Right AI Model or GPU
Plan local inference for downloadable open-weight LLMs. Start with your VRAM or choose a model from Hugging Face.
✓Popular AI models✓Consumer & pro GPUs✓Clear recommendations
How this estimate worksView formulas and assumptions
How are the numbers calculated?
The estimator uses GiB internally and combines model weights, architecture-aware cache or recurrent state, context-aware runtime allowance, multimodal workspace when applicable, and a safety reserve for each GPU.
This calculator includes selected downloadable open-weight text, text-to-text, and multimodal language models from trusted Hugging Face organizations. API-only models, adapters, embeddings, rerankers, and non-generative models are excluded.
WEIGHT MEMORY
Parameters × precision bytes × metadata factorREQUIRED VRAM
Weights + KV Cache + RuntimeAVAILABLE VRAM
GPU Count × (VRAM − Reserve)MoE models are sized by total parameters for weight memory; active parameters mainly affect compute.