Newegg
FOR LLM INFERENCE

Find the Right AI Model or GPU

Plan local inference for downloadable open-weight LLMs. Start with your VRAM or choose a model from Hugging Face.

Popular AI modelsConsumer & pro GPUsClear recommendations
STEP 1 · YOUR INPUTS

Configure Your GPU Memory

Advanced Settings· optional Adjust model quality, context length, and users

Assuming FP8 / INT8 · 8K context · 1 user.

How this estimate worksView formulas and assumptions

How are the numbers calculated?

The estimator uses GiB internally and combines model weights, architecture-aware cache or recurrent state, context-aware runtime allowance, multimodal workspace when applicable, and a safety reserve for each GPU.

This calculator includes selected downloadable open-weight text, text-to-text, and multimodal language models from trusted Hugging Face organizations. API-only models, adapters, embeddings, rerankers, and non-generative models are excluded.

WEIGHT MEMORYParameters × precision bytes × metadata factor
REQUIRED VRAMWeights + KV Cache + Runtime
AVAILABLE VRAMGPU Count × (VRAM − Reserve)

MoE models are sized by total parameters for weight memory; active parameters mainly affect compute.