“We could build this ourselves.”
We hear it from almost every DevOps team we talk to. I ran pipelines and render farms before Juno, so I've said it myself. It's almost always true. A good Kubernetes team can build any single piece of a workload platform, whether that's the scheduler, the template system, the auth layer, the storage provisioning or the workload lifecycle manager. None of it is black magic. These are engineering problems with known solutions.
You can build it. The question is whether you should, and what you're signing up for if you do.
This post walks through what the build actually involves, where the complexity hides, and how I'd think about the decision.
What “build it yourself” actually means
When a DevOps team says they can build a workload platform, they usually mean the visible parts, like a deployment UI, a Helm wrapper, an auth integration and a storage provisioning script. That's real work. It's also the easy part.
The hard part is the orchestration around it, and most teams underestimate it on the first pass.
User identity resolution is one. On a multi-tenant cluster with shared storage, POSIX UID and GID mapping matters. File permissions on NFS or GPFS break in specific, maddening ways when user IDs don't match across workloads. Injecting the right UID and GID into every workload at the Kubernetes layer, for every user, every time, takes an architectural pattern that isn't obvious from the outside. Most teams find out they need it when users start complaining that their files have the wrong owner.
Storage topology is another. The mount map has to answer which volumes each project can see, which are shared across teams and which belong to one user, and how plugins know what to mount without someone configuring every install by hand. A storage topology API that workloads can query at deploy time is not a weekend project. It's several months of careful design and edge cases.
Licensing enforcement is a third. Tracking workloads at the API layer, by project and workload type, with expiry and enforcement, is real implementation work. Most internal platforms skip it and manage access with RBAC. That works until you need granular tracking for chargeback.
None of this is impossible. It's just harder than the first estimate, every time.
The real cost question
The usual “build vs. buy” comparison puts a one-time build cost next to an ongoing subscription. That leaves out maintenance, which is where the money goes.
Internal infrastructure tooling is never done. Kubernetes ships a release every quarter. Your storage backend gets upgraded. A new GPU architecture lands and your workload templates need updating. Dependencies pick up security vulnerabilities. Your team grows and the system has to scale with it. A new cloud provider comes online and you have to integrate it.
Every integration point is ongoing maintenance. As a rough rule, every major dependency you own costs two to three weeks of engineering time a year just to stay current. Ten major dependencies works out to about six months of a senior engineer's year, every year, spent on maintenance instead of product.
The better question is “what does it cost to run, maintain, and evolve over three years?”
Most internal platform projects look cheap until year two.
What DevOps teams actually care about
The “build it ourselves” instinct is often more about control and job security than cost.
Control is a fair concern. DevOps teams have been burned by vendor lock-in. They've watched their employers get held hostage by VMware pricing. They've inherited platforms built on tools that were later deprecated or acquired. Wanting to own the stack makes sense after years of that.
Job security is real too, and worth saying out loud. A team that builds and maintains the internal platform has a clear reason to exist. A team that buys one and configures it has a weaker case. This shows up in most platform evaluations, and pretending it doesn't helps nobody.
Both concerns are fair. For control, pick tools built on open standards with no proprietary lock-in at the data or workload layer. Kubernetes workloads are portable. If Orion disappeared tomorrow, your Helm charts would still work. Job security is harder, because technology doesn't settle it. That one comes down to team culture and leadership.
What “building it yourself” does not get you
The case for building in-house usually goes like this: it fits our needs exactly, we avoid vendor dependency, and we build internal expertise.
What that case leaves out: the 50-plus official plugins that took years to build for common workload types across VFX, research, government, and enterprise environments, the edge cases found across production deployments, the POSIX identity implementation that three different teams attempted before getting right, and the cluster-aware template fields that fill in from live cluster state so template authors never type a wrong resource name.
None of that is magic either. It's engineering that piled up across a lot of production environments. Your internal build would get there eventually. The question is whether you want to pay that discovery cost in-house or start from where someone else already got to.
When building it yourself is actually the right answer
Sometimes it is. Here's when:
Your requirements are narrow and stable. If you need exactly one workload type, in one environment, on one storage backend, and that won't change, a general-purpose platform is more than you need.
Your team has the capacity, and not just the capability. Those are different things. A capable team with no free cycles will build a half-maintained platform that creates more problems than it solves.
The available tools don't fit your constraints. Some environments have requirements no platform handles, and building in-house may be the only option.
The available tools cost more over three years than building and maintaining your own. Then build your own.
The math can come out either way.
What to actually build versus buy
A more useful split than build-or-buy is which parts of the problem are commodity and which parts are specific to you.
The cluster management layer is commodity. Rancher, k9s, kubectl and ArgoCD are mature tools with large communities behind them. Nobody should be building these.
The workload abstraction layer is specific to you. It covers the templates that define what your users can deploy, the identity model that maps to your storage permissions, and the cost model that fits your billing or chargeback needs. The real decision is whether you build that layer yourself or configure a platform that handles it.
Most “build it ourselves” conversations end up being about the commodity layer, because that's the part people can see. The specific layer is what makes a platform worth having.
The teams that decide well do the actual math. They work out what it costs to build, maintain, and evolve over three years, and put that next to the alternatives.
If you want to run that math for your environment, we'll do it with you. Book a time with our team
WRITTEN BY
Tony Como
Tony runs go-to-market at Juno, from first outreach to production deployment. He works across sales, partnerships, and operations to help technical teams move faster without adding infrastructure overhead.
