Paul was hitting a wall that anyone who has run a multi-user Kubernetes cluster knows well.
He needed a new AI agent and a web IDE for a project. In a standard Kubernetes environment, you write the deployment manifest for that. You spell out the container image, resource requests, storage class, PVC definition, service, ingress, and Nginx configuration. For a single workload with GPU access and shared storage, you are looking at several hundred lines of YAML before anything runs. Then you do it all again for the IDE.
Paul is not a Kubernetes engineer, and he should not have to be. The cluster exists to run his work.
We had already built the cluster and already fixed the storage permissions problem. The gap left was between “the infrastructure works” and “the researcher can use it without knowing anything about Kubernetes.” That gap is where most research computing infrastructure lets down the people it is supposed to serve.
A lot of people build products that make managing Kubernetes easier. Almost nobody builds the part that makes it easy for the end user to get at what Kubernetes is running. That is the part we cared about.
What self-service actually has to get right
Self-service compute sounds easy until you list what has to go right every single time a researcher clicks a button.
The workload has to land on the right hardware. A GPU environment has to schedule on a node with a GPU. A web IDE can land anywhere with free CPU. The scheduler handles placement, but only if the nodes are labeled correctly and the resource requests are accurate, so the templates have to encode those requirements correctly.
The storage has to attach with the right permissions. As we covered in the post on multi-tenant storage permissions, that means the researcher’s POSIX identity gets injected before the container starts. The template has to point at the right storage class and the mount has to use the right path. None of it can ask the researcher to configure anything.
The network has to route correctly. The workload needs an ingress route so the researcher can reach it in a browser. That route has to be generated, never hand-configured. The Nginx sidecar has to be configured for the specific application the workload runs.
And all of that has to happen in about sixty seconds, because a researcher who clicks a button and waits five minutes goes back to filing tickets.
How the pipeline works
When a researcher clicks a template card in Orion Workspace, three things happen in order.
Orion Workspace sends the launch request to kuiper, the orchestration service. It generates a kuiper-config, which is a Kubernetes ConfigMap with roughly fifteen fields covering everything about the workload, including the container image, resource requests, storage class, user identity, workload type, and GPU requirement. That ConfigMap is the tracking record of exactly what was requested and by whom.
Next, kuiper reads the kuiper-config and generates every Kubernetes resource the workload needs. That means a Deployment with the right image, resource requests, and a security context that titan has already filled in with the researcher’s UID; a ClusterIP Service; an Nginx sidecar container with config generated for that specific application; a PersistentVolumeClaim with the correct storage class; and an Ingress route with the right path prefix.
The researcher writes zero lines of configuration for any of that. They watch a progress indicator over WebSocket, and sixty seconds later their environment opens in Orion Workspace’s full-screen viewer.
The template itself was written once, by an IT administrator in Orion Admin. They set the resource limits, storage class, GPU requirements, and access controls. The template form fills its options from live cluster state, so the storage classes, priority classes, service accounts, and storage volumes it offers are the ones that actually exist for the project. Template authors pick from real cluster resources instead of typing names that may or may not exist, so nobody gets a deploy-time error from a mistyped storage class.
I always explain this part like a restaurant. The administrators are the kitchen and decide what is on the menu. The researchers order off the menu and never have to walk into the kitchen.
What the catalog looks like today
The official Orion Apps plugin catalog has more than 50 plugins covering most standard research workloads. That includes Jupyter notebooks, a VS Code web IDE, a web terminal, AI coding agents, n8n workflow automation, local LLM inference through vLLM and Ollama, KubeVirt virtual machines, Helios VFX workstations, and infrastructure plugins like GPU Operator, Longhorn, Prometheus, and cert-manager.
For workloads that are not in the official catalog, an Orion Apps author packages a private Helm chart with the right labels and annotations and adds it to a project-specific catalog. The researcher still sees a card and still clicks it. The packaging is a one-time job for someone who knows Helm. Every deployment after that is automatic.
The template system is where IT keeps control. Administrators decide what resources each template can use, which storage it can reach, and which users can deploy it. Researchers deploy inside those limits without ever seeing them. The IT team writes templates once and stops provisioning individual environments.
What this changes for research institutions
The ticket queue in research computing comes from the access model, and more staff will not fix it. When the only way to get an environment is to ask IT to build one, IT is the bottleneck no matter how many people are on the team.
When researchers can launch their own environments from pre-approved templates in sixty seconds, the ticket queue goes away. IT’s job moves from provisioning one environment at a time to maintaining the template catalog. One template definition replaces hundreds of provisioning tickets. The IT team does the same amount of work, and now that work scales.
For institutions that have struggled to justify Kubernetes to non-technical stakeholders, this is the interface. The 65-page deployment process some large studios use to provision workloads is a symptom of the real problem. Kubernetes tooling assumes the person running the cluster is the same person doing the work. In research computing, those are two different people. Orion is built for both of them.
To be clear about the limit, Orion still needs Kubernetes underneath. Kubernetes is still there doing its job. The researcher just never has to look at it.
This is one in a series on building research computing infrastructure. A related post covers what happens when a researcher needs to run a Windows-only tool on the same cluster as a Linux AI agent, and why that is less of a problem than it sounds.
See how Orion handles self-service research computing. Book a demo
Alex Hatfield is the CEO and co-founder of Juno Innovations. Juno builds Orion, the customer-hosted unified compute plane. Orion runs inside your own environment, air-gapped by design, and gives your people one place to use the compute you already have, from GPUs and CPUs to VMs and bare metal.
WRITTEN BY
Alex Hatfield
Alex co-founded Juno to fix how enterprise compute gets done. He leads product vision and customer strategy, and spends most of his time working directly with infrastructure and research teams pushing the limits of what their hardware can do.
