Tony was the first person to really feel it.
He jumped into the cluster to spin up a new AI agent and opened a VS Code IDE to work on the same codebase. The files he needed were there. He could see them. He could not edit them, because the permissions were wrong. The files were owned by a UID from a different container that had created them earlier, and his container was running as a different UID. Standard Linux permissions did the rest.
We had a shared filesystem. What we did not have was a shared workspace.
This is one of those problems that does not exist until more than one person uses a cluster, and then it shows up constantly. Every research computing team running multi-tenant Kubernetes on shared NFS storage hits some version of it. The files are there, the permissions are wrong, someone files a ticket, IT changes the ownership, and the next person hits the same wall.
It happens because Kubernetes containers run as whatever UID the container image says. That UID is baked into the image at build time and has nothing to do with the person who launched the container. On a single-user system, nobody notices. On a multi-tenant research cluster with shared NFS storage, every workload writes files as a different owner, and the Linux permission model enforces that correctly in the most annoying way possible.
What a real fix looks like
The fix is to make the container run as the researcher’s actual POSIX UID instead of the UID the image picked.
That sounds simple. To do it, the platform has to know who is launching the workload, look up their UID and GID from the directory service, and inject both into the container’s security context before it starts. Then every file the container creates belongs to the right person from the very first write.
In Orion this is handled by a service called titan, part of Orion Admin. When a researcher launches a workload through Orion Workspace, titan catches the launch request, resolves their POSIX UID and GID from the logged-in session, and injects both into the Helm values that kuiper uses to generate the Kubernetes resources. The container’s security context is set to the researcher’s real identity before the pod spec ever reaches the cluster.
So a Jupyter notebook, a VS Code IDE, an AI agent, and a web terminal launched by the same researcher all create files with the same owner. A collaborator added to the project group can read and edit those files from any workload, on any node, without anyone fixing permissions by hand. The NFS server sees the researcher’s actual campus UID. Standard Linux group permissions handle access control. IT never gets the ticket.
What the workspace looks like in production
Our shared workspace mounts from a central NFS server to every node in the cluster through a ReadWriteMany storage class. Every workload in a project gets the same workspace at the same path, with the researcher’s correct identity applied.
Now Tony launches an AI agent from the catalog and the workspace is already mounted with his files. He opens a VS Code IDE, and it has the same workspace, files, and permissions. He opens a terminal and lands in the same workspace. He launches a second agent for a different task, and it is the same workspace again. None of those workloads had to be told where the files were. The storage layout is defined once in the project, and Orion’s mount map makes sure every workload in that project gets the right storage attached.
When a new team member joins a project group, the system maps their UID to the project group and applies the correct mount options. Their first workload launches into the same shared workspace with the right permissions already in place. Nobody files an onboarding ticket or fixes permissions after the fact.
Why this matters for research computing
Most research computing documentation treats storage as an infrastructure problem. You configure the storage class, create the PVC, and mount it in the pod spec. That is correct at the Kubernetes level, and it falls well short in a multi-tenant research environment.
The people using the cluster have identities the cluster does not know about, and those identities have to map correctly to the files those people create. That mapping has to happen before the container starts. Think of it like name tags at the door. If you hand them out after everyone is already inside, you spend the whole night figuring out who is who.
For institutions running LDAP or Active Directory, Orion’s POSIX identity layer maps directory identities to container UIDs. Researchers log in the same way they do everywhere else on campus. The cluster knows who they are, and every workload they launch runs as them.
For institutions running NFS, GPFS, Lustre, or VAST storage, the backend stays the same and so does the network. The only thing Orion adds is the identity layer that makes shared storage behave correctly for everyone using the cluster.
This is one in a series on building research computing infrastructure. A related post covers the other side of the problem. Once the infrastructure works, how do you let researchers use it without knowing anything about Kubernetes?
See how Orion handles multi-tenant research storage. Book a demo
Alex Hatfield is the CEO and co-founder of Juno Innovations. Juno builds Orion, the customer-hosted unified compute plane. Orion runs inside your own environment, air-gapped by design, and gives your people one place to use the compute you already have, from GPUs and CPUs to VMs and bare metal.
WRITTEN BY
Alex Hatfield
Alex co-founded Juno to fix how enterprise compute gets done. He leads product vision and customer strategy, and spends most of his time working directly with infrastructure and research teams pushing the limits of what their hardware can do.
