LocalNodeOps

Synced from Hugging Face — 2026-09-14

How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need?

Exact on-disk weight sizes below are pulled directly from Hugging Face. Total VRAM figures also include an estimated KV cache and runtime overhead — see the breakdown per quantization, and the calculator below to adjust context length.

By quantization

Meta-Llama-3.1-70B-Instruct-IQ1_M

15.60 GB exact, on disk

~18.54 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-IQ2_M

22.46 GB exact, on disk

~26.08 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-IQ2_S

20.71 GB exact, on disk

~24.16 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-IQ2_XS

19.69 GB exact, on disk

~23.03 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-IQ2_XXS

17.79 GB exact, on disk

~20.94 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-IQ3_M

29.74 GB exact, on disk

~34.09 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-IQ3_XS

27.29 GB exact, on disk

~31.39 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-IQ4_XS

35.30 GB exact, on disk

~40.20 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q2_K

24.56 GB exact, on disk

~28.39 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q2_K_L

25.52 GB exact, on disk

~29.45 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q3_K_L

34.59 GB exact, on disk

~39.42 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q3_K_M

31.91 GB exact, on disk

~36.48 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q3_K_S

28.79 GB exact, on disk

~33.04 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q3_K_XL

35.45 GB exact, on disk

~40.37 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q4_K_L

40.33 GB exact, on disk

~45.74 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q4_K_M

39.60 GB exact, on disk

~44.94 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q4_K_S

37.58 GB exact, on disk

~42.71 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00001-of-00002

37.26 GB exact, on disk

~42.36 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00002-of-00002

9.87 GB exact, on disk

~12.23 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00001-of-00002

37.14 GB exact, on disk

~42.23 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00002-of-00002

9.38 GB exact, on disk

~11.69 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q5_K_S

45.32 GB exact, on disk

~51.23 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00001-of-00002

37.13 GB exact, on disk

~42.22 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00002-of-00002

16.79 GB exact, on disk

~19.84 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00001-of-00002

37.18 GB exact, on disk

~42.27 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00002-of-00002

17.20 GB exact, on disk

~20.29 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00001-of-00002

37.07 GB exact, on disk

~42.15 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00002-of-00002

32.75 GB exact, on disk

~37.40 GB total estimated VRAM at 4,096 tokens context (KV cache + overhead estimated, not exact)

VRAM calculator

Estimate total VRAM for a given model size, quantization, and context length. Estimates only — see the note at the bottom.

7B13B70B120B
2K4K8K16K32K64K128K
Model weights
KV cache
Runtime overhead (10%)
Total VRAM

Fits on:

    Live changes

    • No changes yet — adjust a control above.

    Generic estimate: weight size is bits-per-weight × parameter count; KV cache uses a reference architecture for the selected size bucket (7B/13B/70B: real published models; 120B: extrapolated, no real model exists at that exact size). Model-specific (HF-synced): weight size is the exact .gguf file size fetched from Hugging Face — no estimation. KV cache is still estimated: this schema doesn't carry layer count, head count, or head dimension, so the calculator parses an approximate parameter count from the model's title (e.g. "8B") and reuses the nearest size bucket's reference architecture for the KV math, same caveats as the generic mode. If no parameter count can be parsed from the title, it falls back to the 7B architecture and says so next to the quantization dropdown. Batch size is fixed at 1 in both modes. Treat all figures here as a starting estimate, not a guarantee. Full formula and known limitations: /methodology.

    FAQ

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ1_M?

    The Meta-Llama-3.1-70B-Instruct-IQ1_M quantization is 15.60 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 18.54 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_M?

    The Meta-Llama-3.1-70B-Instruct-IQ2_M quantization is 22.46 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 26.08 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_S?

    The Meta-Llama-3.1-70B-Instruct-IQ2_S quantization is 20.71 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 24.16 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_XS?

    The Meta-Llama-3.1-70B-Instruct-IQ2_XS quantization is 19.69 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 23.03 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ2_XXS?

    The Meta-Llama-3.1-70B-Instruct-IQ2_XXS quantization is 17.79 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 20.94 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ3_M?

    The Meta-Llama-3.1-70B-Instruct-IQ3_M quantization is 29.74 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 34.09 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ3_XS?

    The Meta-Llama-3.1-70B-Instruct-IQ3_XS quantization is 27.29 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 31.39 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-IQ4_XS?

    The Meta-Llama-3.1-70B-Instruct-IQ4_XS quantization is 35.30 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 40.20 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q2_K?

    The Meta-Llama-3.1-70B-Instruct-Q2_K quantization is 24.56 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 28.39 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q2_K_L?

    The Meta-Llama-3.1-70B-Instruct-Q2_K_L quantization is 25.52 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 29.45 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_L?

    The Meta-Llama-3.1-70B-Instruct-Q3_K_L quantization is 34.59 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 39.42 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_M?

    The Meta-Llama-3.1-70B-Instruct-Q3_K_M quantization is 31.91 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 36.48 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_S?

    The Meta-Llama-3.1-70B-Instruct-Q3_K_S quantization is 28.79 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 33.04 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q3_K_XL?

    The Meta-Llama-3.1-70B-Instruct-Q3_K_XL quantization is 35.45 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 40.37 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q4_K_L?

    The Meta-Llama-3.1-70B-Instruct-Q4_K_L quantization is 40.33 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 45.74 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q4_K_M?

    The Meta-Llama-3.1-70B-Instruct-Q4_K_M quantization is 39.60 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 44.94 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q4_K_S?

    The Meta-Llama-3.1-70B-Instruct-Q4_K_S quantization is 37.58 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.71 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00001-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00001-of-00002 quantization is 37.26 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.36 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00002-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q5_K_L/Meta-Llama-3.1-70B-Instruct-Q5_K_L-00002-of-00002 quantization is 9.87 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 12.23 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00001-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00001-of-00002 quantization is 37.14 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.23 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00002-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q5_K_M/Meta-Llama-3.1-70B-Instruct-Q5_K_M-00002-of-00002 quantization is 9.38 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 11.69 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q5_K_S?

    The Meta-Llama-3.1-70B-Instruct-Q5_K_S quantization is 45.32 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 51.23 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00001-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00001-of-00002 quantization is 37.13 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.22 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00002-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q6_K/Meta-Llama-3.1-70B-Instruct-Q6_K-00002-of-00002 quantization is 16.79 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 19.84 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00001-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00001-of-00002 quantization is 37.18 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.27 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00002-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q6_K_L/Meta-Llama-3.1-70B-Instruct-Q6_K_L-00002-of-00002 quantization is 17.20 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 20.29 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00001-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00001-of-00002 quantization is 37.07 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 42.15 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.

    How much VRAM does Meta-Llama-3.1-70B-Instruct-GGUF need at Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00002-of-00002?

    The Meta-Llama-3.1-70B-Instruct-Q8_0/Meta-Llama-3.1-70B-Instruct-Q8_0-00002-of-00002 quantization is 32.75 GB on disk (exact, synced from Hugging Face). At a 4,096-token context, using the 70B reference architecture for the KV cache estimate, total estimated VRAM is approximately 37.40 GB. This total includes a KV cache estimate and a 10% runtime overhead margin — only the weight size is exact.