Local AI, Simply

Setting up a DGX Spark for local LLMs: what it's good at, and where it bit me

Plumbing Human + AI

I run my whole local setup on a DGX Spark now - a small NVIDIA box that sits on a desk and holds models a normal graphics card could never fit. This is the honest account of setting it up: what it is genuinely good at, and the one edge that reached out and bit me.

If you take one thing from this post: the whole box lives on a single idea, unified memory, and that idea is both why it is special and how it will reboot itself if you get greedy.

What it actually is

It is an NVIDIA GB10 machine (the “Spark” class). Three things about it are not like a normal desktop:

Setting it up, the honest short version

You are not wiring this from scratch. NVIDIA ships “playbooks” (a public GitHub repo of self-contained guides), and most capabilities run as containers you start with one flag. The core I stood up:

That is the whole serving platform. Image generation, fine-tuning, faster engines - all of it is install-on-demand from the same playbooks, added the day you actually need it. The value you add is not the wiring. It is deciding what to run, and how much memory to hand it. Which is where the trouble started.

The edge that bit hardest

Unified memory is the gift and the trap, and they are the same feature.

Because the processor and the GPU share one pool, a greedy setting does not just fill “the GPU.” It eats the memory the operating system itself needs to stay alive.

I told vLLM it could take 85% of memory, and loaded a 122-billion model. It asked for more physical memory than the machine has. The box ran out, panicked, and hard-rebooted itself. Then it got worse: I had set the container to restart on boot, so the machine came back up and immediately loaded the same too-big model again. A reboot loop, built out of one number I chose.

The fix was dull and correct. Give each model a right-sized memory budget instead of a flat percentage, and do not auto-load a model on boot. Now the box starts clean with nothing loaded, and I load a model by hand when I want one.

The lesson: on unified memory, “how much memory does this model get” is a safety setting, not a performance dial. Set it too high and you do not go faster. You take the whole machine down.

The other honest limits

If you want a quiet, offline box to experiment with genuinely large models - run several tools side by side, even try training your own - this is rare and good at its size. Almost nothing else this small lets you hold a 100B model in your own room with the network cable unplugged.

If you want a no-thought replacement for the cloud, or you depend on software with no arm64 build, or you were picturing 128GB of spare graphics memory to fill however you like - you will spend your first week fighting it.

It is not a magic cloud in a box. It is a real, capable, slightly sharp-edged experimentation machine. Respect the memory, and it becomes a quietly remarkable thing to have sitting on a desk.

Ingredients: human + AI.