LocalNodeOps

GPU

RTX 4090 24GB

24GB VRAM handles 70B models at Q4 quantization with room for an 8K context window.

18.2 Tokens/sec (Llama 3 70B, Q4)
Verified

Full benchmark writeup pending — placeholder entry for homepage layout.