Skip to main content
HighVRAM Hardware Watch

VRAM vs unified memory: what matters for local AI

A local AI model needs room for its weights, the conversation it is processing, and the software running it. The useful question is how much memory your chosen setup can use for that whole workload. This guide focuses on running language models, also called inference.

Use Start Here for the broader buying process. Here, we will separate the memory numbers and explain why a model can load successfully but still struggle with a longer conversation.

VRAM, system RAM, and unified memory

  • Dedicated VRAM is memory on a discrete graphics card. The GPU uses it for model data and working buffers. System RAM is a separate pool.
  • System RAM supports the operating system, applications, and CPU-side work. A computer with 64GB RAM and a 24GB graphics card has two separate capacities; adding them does not create an 88GB dedicated GPU.
  • Apple unified memory is a shared pool accessible to the CPU and GPU. Apple silicon can share data without copying it between separate CPU RAM and graphics-card VRAM. Apple explains the architecture here.

A 64GB unified-memory Mac still has 64GB for the whole system. macOS and other apps use part of it, and the runtime must allocate its own working memory. Apple also exposes a recommended GPU working-set size. Treat the advertised capacity as a starting point, then check what your software can actually allocate. See our Mac Studio listings and unified-memory notes.

The words “shared memory” alone tell you little about performance. Check the actual processor, memory architecture, and supported runtime before comparing a shared-memory computer with a discrete GPU.

Build the memory budget in four parts

1. Model weights. These are the learned numerical values stored in the model. Quantization stores them at lower precision to reduce size, with possible quality tradeoffs. File labels such as Q4_K_M describe a format; they are not a promise that every parameter occupies exactly four bits. llama.cpp documents the formats and their differing sizes.

Illustrative calculation, not a benchmark: exactly 8 billion weights stored at exactly 4 bits each would occupy 8 billion × 4 ÷ 8 = 4 billion bytes, or 4GB (about 3.73GiB). At 16 bits, the same raw calculation gives 16GB. Real model files can add quantization metadata and use mixed precision. Neither calculation includes the context cache or runtime buffers, so it cannot establish which machine will run the model.

2. Context and KV cache. Context is the material available during a request, including prompts and responses. For many transformer models, the key/value, or KV, cache holds attention data from processed tokens. Increasing the configured context length can increase memory requirements even when the model file stays the same. Ollama documents that memory cost; LM Studio documents context and cache settings.

3. Concurrent work. Serving several requests at once needs additional context capacity. Keeping multiple models loaded also consumes memory. In Ollama, parallel requests increase context-related allocation; its concurrency documentation explains the relationship. A single-user chat test does not establish capacity for a multi-user server.

4. Runtime and other applications. Inference also needs temporary computation and output buffers. A llama.cpp maintainer explains these separate allocations. System memory must also cover other apps and the operating system; Apple’s Activity Monitor guide distinguishes those uses. Leave headroom based on your measured workload rather than applying a universal reserve percentage.

What happens when the GPU is too small?

Some runtimes can split a model between the CPU and GPU. llama.cpp supports hybrid inference, and LM Studio provides GPU-offload controls. This can make a larger model usable, but CPU work and transfers can reduce performance. The result depends on the hardware, model, and split. Ollama recommends avoiding CPU offload for best performance.

Compare that compromise with running a smaller model fully on the GPU. Test both on your actual prompts. Extra memory can make a different model or quantization possible; judge that choice on the answers it gives you.

Capacity, speed, and quality are separate checks

Capacity measures how much memory is available. Memory bandwidth measures how quickly data can be read or written, while compute capability affects how quickly operations run. Apple explains these distinct performance limits. The capacity number alone cannot predict response speed. Compare prompt-processing time and generation speed at the same model settings, and assess answer quality with tasks you care about.

Check the configuration you will actually use

  • Record the exact model, quantization, runtime version, context length, and expected simultaneous requests.
  • Use the runtime’s estimate before loading. LM Studio’s memory estimator accounts for settings including context length and GPU offload. Treat its output as an estimate.
  • Run a representative long prompt and generate a useful-length response. Watch memory use, errors, responsiveness, and answer quality. Include the other apps you normally keep open.
  • Check where the model loaded. Ollama’s ollama ps output reports the CPU/GPU split. Recheck after changing context or concurrency.

For help deciding what to test first, use our local AI workload guide. Start with the job you need done, then choose enough usable memory for a configuration that performs well at that job.