HPC

·

·

7 min read

Building a research supercluster from whatever you have lying around

Alex Hatfield

CEO, Juno Innovations

Workbench with a gaming PC, single-board computers and cables wired together

ON THIS PAGE

We started with a Minecraft server.

There was no render farm and no GPU cluster. It was a single server running Minecraft in Docker Compose, with no failover. If that one machine went down, the whole game went down. It was a humble start, and it showed us right away the problem we would spend the next four years on. How do you make a pile of mismatched machines behave like one system?

We are not the only ones asking. Walk into almost any research computing environment outside a well-funded R1 and you will find the same thing. The cluster is built from hardware rescued from retirement, handed down by departments after upgrades, and bought a piece at a time across several grants. It runs, and researchers depend on it. But every new piece of hardware is a maintenance event. You debug the driver conflict, reconfigure the network overlay, patch the node labels, and sort out whatever the new box and the old boxes disagree about.

A mixed cluster is what you get when you build real infrastructure on a real budget. We built Juno inside that constraint, and what we learned is why Orion works the way it does.

What our cluster looks like

Today our development cluster runs on six nodes that have nothing in common except the software holding them together.

The anchor is an NVIDIA DGX Spark. It has a GB10 Grace Blackwell chip with 128GB of unified CPU and GPU memory on a single die, 3.7TB of NVMe storage, and an ARM architecture. It is a desktop box about the size of a thick book.

Next to it sit an AMD Ryzen 7 workstation with an RTX 2080 Super and four Intel NUC machines with i5 processors from 2014 and 2016 and 16GB of RAM each. The NUCs run on older SATA drives, the Spark runs on NVMe, and the AMD box is somewhere in between.

In a traditional environment, that combination is a headache. You have different CPU architectures, different GPU vendors, different OS versions, and different driver requirements. The ARM Spark alone would usually need its own management setup just to deal with the architecture difference.

With Orion, all six nodes show up as one unified compute plane. Workloads schedule across whatever resources are free. A researcher asking for GPU lands on the Spark. A researcher opening a web terminal lands on whichever NUC has room. Neither of them thinks about where their workload is running, which is exactly the point.

What changed when the Spark showed up

Before Orion, adding a node to a Kubernetes cluster meant a manual checklist. You configured the network overlay for the new architecture, debugged any storage path differences, patched the GPU device plugin by hand, and labeled the node correctly so GPU workloads would actually land on it.

When we plugged in the Spark, we did none of that.

The GPU Operator deploys as a single Orion Apps plugin. It detected the GB10 Blackwell architecture on its own, applied the right CUDA drivers, configured the NVIDIA container toolkit, and labeled the node as GPU-capable. The Spark went from plugged in to scheduling workloads in about ten minutes. Nobody injected drivers by hand, wrote architecture-specific config, or sat there debugging.

The bigger payoff is for institutions that add hardware all the time, like donated workstations, grant-funded GPU nodes, and department machines being pulled into a shared pool. Each of those used to be a maintenance event. With the GPU Operator plugin, it stops being one.

What this means for your institution

Most research computing teams have hardware they are not fully using. A department upgraded its workstations and the old ones are sitting in a closet. A grant bought GPU nodes that sit mostly idle between experiments. A faculty member has a Spark on their desk doing one researcher’s work when it could be doing everyone’s.

Orion turns that hardware into a cluster. The same software that tied a Minecraft server to a Blackwell GPU works on whatever you have, because Orion is built on Kubernetes and sits above the hardware.

Funny enough, the Kubernetes scheduler never cared that one node is ARM and another is x86. It cares about free resources and node labels. Orion handles the labels, the scheduler handles placement, and the researcher clicks a button.

For institutions looking at DGX Sparks through Cambridge Computer, the Spark is the right anchor for a research AI cluster. Its unified memory lets you run models that would need several separate GPUs on traditional hardware. But a Spark is worth more as part of a managed pool than as one researcher’s workstation. Orion is what turns it into a shared resource, and the people using it never have to feel like they are sharing.

This is one in a series on building research computing infrastructure from mixed hardware. A related post covers what happens when several people share that cluster, and why file permissions are the first thing to break.

See how Orion handles mixed hardware in your environment. Book a demo

Alex Hatfield is the CEO and co-founder of Juno Innovations. Juno builds Orion, the customer-hosted unified compute plane. Orion runs inside your own environment, air-gapped by design, and gives your people one place to use the compute you already have, from GPUs and CPUs to VMs and bare metal.


WRITTEN BY

Alex Hatfield

Alex co-founded Juno to fix how enterprise compute gets done. He leads product vision and customer strategy, and spends most of his time working directly with infrastructure and research teams pushing the limits of what their hardware can do.