Skip to main content
HighVRAM Hardware Watch

Test your local AI workload before you buy

Before buying hardware, run a small, repeatable check of the work you want it to do. The result should tell you what needs improving: answer quality, waiting time, memory capacity or software setup.

This guide covers local language-model tasks such as chat, coding and document questions. Image generation, video and training need tests specific to those applications. You can repeat this check after changing a model or updating your software.

1. Decide what counts as a pass

Choose one everyday task and one demanding task you genuinely need, such as your longest document question. Write down what a useful answer must contain and how long you are willing to wait for the first response and the finished answer.

Use non-sensitive sample material. Confirm the selected model runs on your machine and check whether connected tools use online services. Installing a local app does not make the whole workflow offline. Ollama explains its local and cloud modes.

2. Keep a test record

Copy these fields into your own notes. Use the same inputs and settings when comparing machines or changes:

  • Machine: exact GPU or Mac chip; GPU VRAM and system RAM separately, or total unified memory for an Apple silicon Mac.
  • Software: app/runtime version, operating system and GPU driver version where applicable.
  • Model: exact model tag or file, quantization and any custom settings.
  • Workload: saved prompt and sample input, configured context length, output limit and number of simultaneous requests.
  • Result: answer passes or fails; time to first response and completion; first run versus later runs.
  • Hardware use: GPU/CPU placement, observed memory use, other open apps and any errors.

3. Run the work you actually need

Run the everyday task, then the demanding task, with your normal background apps open. Start a fresh conversation for each. Repeat the tests and keep first-run and later-run timings separate; loading and caching can affect the result.

Check the context setting rather than assuming the app uses the model’s advertised maximum. A larger context requires more memory, so use the length your task needs. Ollama’s context settings and checks.

If you need several requests at once, test that load too and note whether requests run together or queue. A single successful chat does not establish multi-user capacity. How Ollama handles parallel requests.

4. Check what the hardware is doing

If you already use Ollama

While your local model is loaded, open a second terminal and run this read-only check:

ollama ps

Find your model. PROCESSOR shows whether it is loaded on GPU, CPU or both; CONTEXT, when shown, reports the allocated context. “100% GPU” describes model placement, not GPU utilization or spare memory. Read the placement output.

Unexpected CPU use is a reason to check hardware support and GPU detection before shopping. Placement alone does not explain why a model is slow. With another runtime, use its equivalent load-status view.

For an NVIDIA discrete GPU

If NVIDIA’s driver tools are already available, this read-only command refreshes the GPU summary every second. Run it during the tests and press Ctrl+C to stop:

nvidia-smi -l 1

Note reported memory use before and during the task. Samples can miss a brief peak, and some fields may be unavailable on your system. Keep the units shown. NVIDIA’s monitoring reference.

For an Apple silicon Mac

Open Activity Monitor → Memory. Note Memory Pressure and changes in Swap Used before and during the test, alongside any slowdown. A nonzero swap figure alone is not a pass/fail test. Apple explains these readings.

Apple silicon’s CPU and GPU share unified memory. The advertised capacity is total system memory, not dedicated GPU VRAM or guaranteed space for a model. macOS and other apps use that pool too. Apple’s unified-memory explanation.

5. Turn the result into a buying decision

  • Both tasks pass: keep using this setup. Save the record as a baseline for future updates.
  • Quality fails but the app runs reliably: reassess the model, prompt and workflow first. More memory does not guarantee a better answer.
  • Loading fails or memory runs short: record the error, check compatibility, then change one factor at a time, such as context length, model size or parallel load. If the only settings that fit fail your original task, seek evidence for a higher-capacity setup.
  • The model fits but the wait is too long: compare measured response times for the same model, quantization, context and workload on candidate hardware. VRAM capacity alone is not a speed ranking.
  • You cannot test the intended model: look for a reproducible test with those same settings before committing to a purchase. Treat a different model or a short-prompt demo as incomplete evidence.

For a small, checkable quality task, try our fictional invoice extraction practice, which includes original sample documents, an answer key and an offline CSV validator.

A screened listing still needs to meet your workload, compatibility and total-budget checks. If the available offers do not, waiting is a valid outcome. Take your test record to the hardware buying guide and then compare current offers.

Source documentation checked 6 October 2026. App behavior and hardware support can change.