Enter your graphics card memory and system RAM, pick a precision, and the calculator estimates which open-weight models from the open-weight model register fit. It covers the model weights only, so treat the result as a first check before you download hundreds of gigabytes.
Loading models…
How the estimate works
Weights. The memory for the weights is the total parameter count times the bytes per parameter: 2 bytes in BF16, 1 byte at 8 bits, and about 0.56 bytes for common 4-bit formats, which store a scale alongside every small block of weights. A 70B model therefore needs about 140 GB in BF16, 70 GB at 8 bits and about 39 GB at 4 bits. The calculator adds 10% for runtime buffers.
Mixture-of-experts models. A mixture-of-experts (MoE) model uses only some of its parameters for each token, but all of them must still be in memory. The calculator counts the total parameters. The active parameters listed in the register affect speed, not memory.
Fits on the GPU means the estimate is no larger than your GPU memory, which is the fast case. Fits with offloading means it fits only when part of the model sits in system RAM; runtimes such as llama.cpp can do this, but generation is much slower. The calculator keeps 4 GB of RAM for the operating system. With unified memory it counts 75% of RAM as available to the model, a conservative assumption rather than a fixed limit.
What it leaves out. The context cache grows with the length of the conversation and can take several more gigabytes at long context lengths. Some models, including several DeepSeek and Kimi releases, are published in FP8 or 4-bit formats rather than BF16; the "as published" option uses the size of those files. The calculator does not check whether a quantized build of a given model exists. For the trade-offs of running models on your own machine, read Running AI locally: what your laptop can handle.