AI/ML

·

·

7 min read

The DGX Spark is the most underutilized AI hardware on the market

Alex Hatfield

CEO, Juno Innovations

A compact desktop AI system on a lab desk next to a monitor showing a molecular model

ON THIS PAGE

NVIDIA built something really impressive with the DGX Spark. It has a GB10 Grace Blackwell chip and 128GB of unified memory shared between the CPU and GPU on a single die. It can run models that would need several server-class GPUs on traditional hardware, in a box that fits on a desk and runs quietly enough to sit in an office.

The hardware is not the problem.

We have spent a lot of time with Spark deployments at research institutions, creative studios, and development teams, and the pattern is always the same. The Spark arrives, the one technical person on the team gets it set up and running, and it becomes that person’s workstation. Everyone else works around it.

The gap between “powerful hardware” and “hardware the whole team gets something out of” is a software problem. More specifically, it is a Linux problem.

Why the Spark underdelivers in practice

Setting up a Spark for a real AI workload takes serious command-line skill. You install the CUDA toolkit, configure the NVIDIA container runtime, manage container images, set up storage paths, juggle environment variables, and untangle library version conflicts between workloads. Every step has documentation, and the documentation assumes you know Linux. Most of the people who need GPU compute do not.

A researcher who wants to run protein structure prediction should not have to understand container orchestration first. A designer who wants local image generation should not have to debug CUDA environment variables. A developer who wants a local LLM for code help should not have to tune inference server parameters.

Without a software layer that hides all of that, they do have to. And most of them will not. So the Spark sits on one desk and does one person’s work, while everyone else pays for API access to tools running on somebody else’s hardware.

What we built, and why

Juno started in VFX. We built compute infrastructure for people who had to render films with almost no hardware budget. We could not afford to waste anything. Every GPU had to be shared intelligently, and every workload had to just work without a setup session.

That constraint is why Orion exists, and it is why Orion fits the Spark so naturally.

Orion installs on the Spark and shows a catalog of workloads in a browser. Think of it like a menu. Anyone on the team with access opens the URL, sees the applications they are allowed to run, clicks what they need, and has a running environment in about sixty seconds. There is no command line and nothing to configure, and workloads do not step on each other, because each one runs in its own isolated container with its own dependencies.

An administrator decides what goes on the menu and what resources each workload gets. Researchers get the AI tools they need. Developers get local LLM access for code help. Creative teams get image generation environments. IT stays in control without having to hand-provision every single request.

What the Spark can do once it is shared

When Orion runs the Spark as a shared resource instead of a personal workstation, the utilization picture changes completely.

Several people have workloads running at the same time. The Spark’s 128GB of unified memory serves a handful of users instead of one. When one person’s workload goes idle, that capacity is there for someone else. When people go home, automated work like CI agents, analysis pipelines, and monitoring fills the gap.

The inference endpoint on the Spark serves the whole team. We run a 35-billion parameter model on a single Spark that handles several users and agents at once, with no per-request bill, no rate limits, and no data leaving the infrastructure. The team’s AI assistant is always there because it runs on hardware they control.

For teams with several Sparks, Orion clusters them into one unified compute plane. A second Spark joins the cluster in minutes and workloads schedule across both devices. For institutions where departments own their own Sparks, Orion can federate them so idle capacity goes into a shared pool.

The fair comparison

A capable API-based AI service charges per request. A team that leans on AI all day, for code review, research analysis, content generation, and agent automation, runs up a real monthly bill. Because the bill grows with every request, teams get hesitant about using AI as freely as they should.

A Spark running Orion costs electricity once you own the hardware. The team uses it as much as they want without counting requests. For teams doing serious AI work, the economics tip toward local compute faster than most people expect.

The data governance argument is separate, and on its own it often decides the question. Research data, proprietary code, patient information, and unreleased work should never cross an external API. In those cases, local inference on a Spark is the only path that meets the requirement.

To be fair to the API side, a Spark is still one box with 128GB of memory. If your team needs the very largest models, or a lot of them at once, you will outgrow one device. That is exactly when the clustering above starts to matter.

What comes next

The Spark is the right foundation for team-scale AI infrastructure. The form factor, the memory architecture, and the performance per watt are all right. What has been missing is the software layer that makes it usable by everyone on the team, not just the person who set it up.

That is what we are excited about with Orion on the Spark. It closes the gap between powerful hardware and the people who need it.

If you are thinking about what a Spark deployment looks like for your team, we are happy to work through it. Get a Demo

Alex Hatfield is the CEO and co-founder of Juno Innovations. Juno builds Orion, the customer-hosted unified compute plane. Orion runs inside your own environment, air-gapped by design, and gives your people one place to use the compute you already have, from GPUs and CPUs to VMs and bare metal.


WRITTEN BY

Alex Hatfield

Alex co-founded Juno to fix how enterprise compute gets done. He leads product vision and customer strategy, and spends most of his time working directly with infrastructure and research teams pushing the limits of what their hardware can do.