We hear this a lot. “Can’t we just add more nodes?”
And yes, you can. That logic built the VFX industry. You need more throughput, you buy more hardware, you add it to the pool. It is simple math and it worked for decades. I am not going to tell you it failed, because it didn’t. It got studios through some of the most complex productions ever made.
But the work changed, and the infrastructure didn’t.
We came from VFX. We originally built Juno because we kept hitting the same problem. You have 100 servers and 100 jobs queued up, and every server is running at half capacity. Another 50 jobs come in and they all wait. Meanwhile there is idle compute everywhere you look. The math should work. It doesn’t, because nobody taught the farm to share.
That is where this started for us. And the more studios we have worked with since, the more we see the same problem in different forms.
The real issue with how farms are built
Traditional renderfarms were designed for batch compute. You submit a job, the farm processes it, you get output. The queue is clean and so is the utilization. The GPU either has a job or it is waiting for one.
Today’s pipeline is a lot messier. You have batch jobs, and right next to them you have interactive workstations. Artists preview complex simulations in real time, creative directors review renders before they finish, and teams in London and Vancouver hand off work mid-frame. Compute demand is no longer a queue you can schedule cleanly. It is predictable jobs and unpredictable people hitting the same hardware at the same time.
Run those on separate dedicated infrastructure and the economics fall apart fast. The farm sits idle between submissions. The workstation pool has GPUs assigned 1:1 to artists who are on a call or reading notes. Both pools look busy from the outside. Measured by utilization, they are not.
Enterprise GPU utilization across on-premises infrastructure typically sits at 10-15% on a fleet-wide, time-averaged basis. The first time studios see that number, it surprises them, and I get why. The individual nodes look active and the dashboard shows something. But the share of total available GPU-hours doing real work is much lower than it looks, and that gap is where the money goes.
What we did about it
Starting in VFX meant solving this on a tight budget. We were not a hyperscaler and could not just throw hardware at it. We had to get clever about how compute was handed out, and that constraint ended up being the whole product.
GPU slicing is not new. NVIDIA has had partitioning for years through time slicing, MIG, and vGPU. We do not do the slicing ourselves, NVIDIA does. What we built is the orchestration on top, so an artist’s interactive workstation and a batch render job can share the same physical GPU, based on what each one actually needs at that moment.
We built this into Orion because we kept seeing the same thing at studios. The renderfarm had idle capacity while artists waited on their workstations. The workstations had idle GPUs while the farm was peaking. Nobody had a way to let the two pools talk to each other, so they stayed siloed.
When you connect them, things change quickly. At R3D Studios, moving from 1:1 GPU-per-artist provisioning to a shared fleet got them to 2:1 GPU density, with 10 artists running at once across 5 GPU instances. R3D Studios reported up to ~40% compute cost reduction. The artists did not notice a performance difference because the scheduling layer handled when each workload needed priority.
That is one studio at one scale, and other setups will have different numbers. But the direction holds. When idle capacity in one pool flows to the workloads that need it, you get more output from the same hardware. That is really where this shines.
The workstation side of it
The other piece is how workstations get provisioned, and it gets overlooked a lot.
Traditional workstation infrastructure is a fixed image on fixed hardware. You provision it, maintain it, and replace it, and the environment is tied to the machine. If an artist in Vancouver sets up their workstation perfectly for a project and a supervisor in London needs to pick up where they left off, the supervisor starts from scratch or syncs files by hand and hopes it works.
Container-native workstations fix that. The whole workstation environment, meaning the OS, the application stack, and the project state, lives in a container. The hardware is whatever is free. You can run the same workstation on a local node, on a cloud instance, or on a node that just finished a batch job. The workstation follows the artist.
We built this because VFX artists do not want to think about infrastructure. They want to click something and have it work. What started as “let’s make it easy for an artist to get a Blender workstation” turned out to apply directly to researchers who want a Jupyter notebook, developers who want a sandbox, and anyone who just needs a workload without a fight.
The 60-second launch time came out of the same pressure. Studios doing real-time creative reviews, where a director’s feedback loop is measured in minutes, cannot wait 20 minutes for a workstation to spin up. Container-native provisioning is what makes the fast launch possible at scale.
What the renderfarm actually becomes
We are not saying renderfarms go away. They don’t. Studios with steady, predictable batch workloads will keep running them. For pure throughput on known job profiles, a well-tuned farm is still the right tool.
What changes is how the farm relates to everything around it. When the farm is part of one unified compute pool instead of its own silo, idle farm capacity can flow to interactive workloads, and idle workstation GPUs can flow to batch jobs off-peak. The same hardware does more without buying more of it.
The studios we have worked with that made this shift kept their farms at the same size. They run the same farms more efficiently and add interactive capacity on top without a matching hardware bill.
What ends is the renderfarm as a closed system.
Why this matters right now
Buying GPU hardware is hard right now. Lead times are long, and cloud GPU availability tightens at peak demand. If your studio is planning capacity for a bigger slate, the first question is whether the fleet you already have is actually being used.
In most cases we have seen, it is not, and nobody made a bad decision. The infrastructure was built for a workload profile that has since moved on. The studio that did all batch rendering five years ago now has a lot of interactive workstation demand on top, on infrastructure that was never designed to share across both.
Fixing that does not always mean buying more hardware. Sometimes it means connecting the hardware you already have.
We are happy to work through what this looks like for your setup. Book a time with our team
Alex Hatfield is the CEO and co-founder of Juno Innovations. Juno builds Orion, the customer-hosted unified compute plane. Orion runs inside your own environment, air-gapped by design, and gives your people one place to use the compute you already have, from GPUs and CPUs to VMs and bare metal.
WRITTEN BY
Alex Hatfield
Alex co-founded Juno to fix how enterprise compute gets done. He leads product vision and customer strategy, and spends most of his time working directly with infrastructure and research teams pushing the limits of what their hardware can do.
