
VRAMFit: LLM Calculator
Will that LLM fit your GPU?
0 ratings
Free
About
Before you download gigabytes of model weights, know whether they'll actually fit.
VRAMFit tells you the largest local large language model you can run on your hardware — whether that's an NVIDIA GPU's VRAM or the unified memory on an Apple Silicon Mac. Set your memory, pick a quantization, and get a clear, instant answer.
Built for people who run models locally: llama.cpp, Ollama, LM Studio, and anyone deciding what GPU to buy next.
WHAT IT DOES
• Enter your VRAM (or Apple Silicon RAM) and instantly see what fits
• Accounts for quantization — from full precision down to 4-bit and FP4
• Factors in context length and KV cache, the memory costs people usually forget
• Switches between discrete GPU and Apple Silicon unified-memory math
• Ships with a 2026 catalog of popular open models, sized for you
FREE
• The full calculator, with quantization, context, and FP4
• A starter set of 5 popular models
• GPU and Apple Silicon modes
VRAMFit Pro (one-time upgrade)
• The complete 20-model 2026 catalog
• Per-quantization memory breakdown for every model
• Advanced tools: KV-cache compression and multi-token prediction estimates
PRIVATE BY DESIGN
VRAMFit runs entirely on your device. No account, no sign-in, no tracking, and nothing about you leaves your iPhone. The numbers you enter stay with you.
Stop guessing whether a model will fit. Check first, download once.
Show more
What's New in VRAMFit
1.0
July 21, 2026







