HighVRAM Hardware Watch

Your first local AI chat

Steps checked against official documentation October 6, 2026. This guide has not been independently tested end to end.

Try a short conversation on the computer you already own before buying hardware. This graphical LM Studio walkthrough covers supported Windows PCs and Apple Silicon Macs. The goal is one locally generated answer, followed by a simple check that the same chat can work offline.

1. Check your computer

Write down your processor or Mac chip, operating system, installed memory and free storage. On Windows, Settings → System → About shows basic specifications. On a Mac, About This Mac identifies the chip. Also note your GPU and its dedicated VRAM if your PC has one.

  • Apple Silicon Mac: LM Studio currently requires macOS 14 or newer. It recommends at least 16GB RAM; some 8GB Macs may run smaller models with modest context. Intel Macs are not supported.
  • Windows: x64 processors need AVX2 support. The requirements also list ARM support for Snapdragon X Elite. LM Studio recommends at least 16GB RAM and 4GB dedicated GPU memory. These recommendations do not guarantee that a particular model will fit.

Compare your exact machine with the current system requirements. If you are unsure about AVX2 or your GPU, check its manufacturer specifications before proceeding.

2. Get the desktop app

Open the official LM Studio download page and choose the “Download LM Studio” desktop-app section for your operating system and processor. The page also offers Bionic and the headless llmster service; this walkthrough follows the LM Studio chat app. Run its installer and open the app. If an unexpected security warning appears, stop and verify the download rather than bypassing it.

3. Download one modest chat model

Open the model discovery area, called Discover in the official getting-started guide. Look for a small text-chat model in the curated options, often described as Chat or Instruct because it is designed to answer requests. Read its model card for the intended use, language support and license. A model being available to download does not mean every use is permitted.

Choose a format your app supports: GGUF works through llama.cpp; Apple Silicon also supports MLX. Start with one small download, check the file size against free disk space, and wait for it to finish. The download guide explains quantized variants: these reduce size by storing model data at lower precision, with possible quality trade-offs. It suggests 4-bit or higher when your machine can handle it.

Treat fit labels and memory estimates as guidance. The model file is only part of the running memory requirement. Longer context, which is the material available to the model during a conversation, can need more memory. A smaller model is a sensible first experiment, with no promise about answer quality.

4. Load it and ask a short question

Go to Chat, open the model loader and select the model you downloaded. Review the memory estimate and loading warnings. Keep the initial settings conservative rather than increasing context to its maximum. If the app warns about insufficient resources, choose a smaller model or lower context; leave its safeguards enabled.

If the app needs a runtime download, let it finish while you are online. Wait until loading finishes. Create a fresh chat and send this non-sensitive prompt: “Explain the difference between RAM and storage in three short sentences.” Once an answer arrives, ask a follow-up. The chat documentation describes creating and managing conversations. A plausible reply shows generation worked, not that its claims are correct.

5. Check that inference is on this computer

Confirm the selected model is the downloaded model running on this machine. Avoid cloud choices and linked remote devices: LM Link can run models on another computer. A familiar model name alone does not identify where it runs.

For an optional practical check, first finish the model and any required runtime downloads and get one answer. When disconnecting will not interrupt other work, temporarily disconnect this computer from all networks, including Ethernet or a hotspot, then send a new, harmless prompt. Reconnect afterward. A new answer while fully disconnected is strong evidence that this attempt ran on-device; it is not a complete privacy audit. LM Studio documents offline chat with downloaded models.

6. Understand the privacy boundary

For ordinary local chat, LM Studio says prompts stay on your device. Model searches, downloads, runtime downloads and update checks use the internet. Its privacy policy also describes cloud processing when you choose optional cloud models or web search. Connected MCP tools may access files and the network. Leave tools and remote features out of this first test. Local operation also leaves chat files on your computer, so protect the device and be careful with backups and shared accounts.

If something goes wrong

  • Download fails: check connectivity and free disk space, then retry through the app.
  • Model will not load: save the exact error, check format support and required runtime updates, and try a smaller model. For memory warnings, reduce context and close unneeded apps.
  • Offline chat fails: reconnect and read the error. Check that the selected model and its runtime are downloaded before repeating the test.
  • Very slow or poor answers: try a short prompt in a fresh chat. Record the model, settings and app version before changing anything. Do not assume a hardware purchase is the fix.

Next, test your own workload and read VRAM vs unified memory: what matters for local AI. Return to Start Here when you have evidence of what needs improving.