AI help is now normal in software development, writing, and most knowledge work. In regulated research, meaning life sciences, pharmaceutical development, and clinical research, the same tools other teams use every day are often off the table.
Data governance is what keeps them off.
When your research data includes patient information, unreleased clinical trial results, proprietary compound structures, or anything under HIPAA, GxP, or institutional data governance rules, sending it to an external API is a compliance question. In most regulated environments the answer is no.
The way around that is local inference. The model runs on hardware you control, on a network you control, and the data never leaves your environment. Here is what that looks like when it is built properly.
What regulated research AI actually needs
“Run it locally” is the start of the requirements list, and the list is longer than that.
Audit trails matter. Every query to the model, every output, and every action an automated agent takes has to be logged in enough detail to reconstruct what happened and why. GxP environments in particular require that computational processes are validated and their outputs are traceable.
Access controls have to be granular. Different researchers have access to different data, and the AI system has to respect that. For example, a clinical researcher should not be able to use an AI assistant to pull data from a study they are not on, even indirectly through a shared model endpoint.
The model has to be version-controlled and validated. If AI is part of a regulated workflow, you need to document the exact model version used for each analysis. In some contexts, a model update in the middle of a study counts as a protocol deviation.
Workloads have to be isolated from each other. A researcher querying sensitive patient data should not share compute with another workload in any way that could leak data.
Standard cloud AI services do not cover most of this, by design, because they are built for general use. Local infrastructure built for regulated environments can cover all of it.
How we lay it out
At the center is a local inference endpoint running on hardware inside the regulated environment. The model runs on the institution’s hardware. Network traffic stays on the institution’s network. Data never crosses an external connection.
Orion manages the workload layer. Each research team or study gets its own project namespace with its own compute quotas, storage access controls, and network policies that block traffic between namespaces. One study’s AI agent cannot reach another study’s data because the network policies will not route it. We are not leaning on application-level checks for that.
The POSIX identity layer makes file permissions follow the researcher. Files an AI agent creates while running in a researcher’s context are owned by that researcher. A project’s storage mounts are only reachable by users in that project’s group. Access control lives at the filesystem level, under the application.
In this setup, every inference request passes through a logging layer before it reaches the model. The log records the timestamp, the user, the project, the prompt, and the model version. It is write-once and kept separate from the main data store. Audit queries against it can reconstruct exactly what the model was asked and what it answered at any point in time.
Model versions are handled the same way as workload templates. A specific model version is pinned in the template configuration. Moving to a new model means creating a new template version instead of editing the old one. The old version stays available, so a study that needs to reproduce results with the original model can relaunch it.
On a normal server you just upgrade the thing. Here, being able to put the old version back exactly is the whole point.
What agents do in this environment
Agents in regulated research run under tighter rules than agents in a normal dev environment. They do not change study data without explicit researcher approval. They do not query across study boundaries. Their outputs are clearly marked as AI-generated and go through human review before anything enters a regulated workflow.
Inside those rules, agents are still very useful for regulated research.
Literature monitoring agents watch publication databases for new papers in a study’s area, summarize what they find, and bring it to the research team so nobody has to watch the feeds by hand.
Data quality agents run automated checks on incoming experimental data and flag anomalies, missing values, or out-of-range values before the data enters the analysis pipeline. Catching problems early makes them cheaper to fix.
Documentation agents watch analysis workflows and notice when results were generated without matching documentation updates. They draft the missing sections for the researcher, who reviews and approves them. The agent never writes to regulated documentation directly.
Protocol deviation agents check whether computational workflows run according to validated protocols and flag deviations for review before results are recorded.
All of these agents run as Orion workloads inside the project namespace of the study they serve. They can only reach the data that study’s access controls allow. Their actions are logged. Their outputs wait for human review before any regulated action happens.
The air-gapped case
Some regulated environments need complete network isolation, with no external connectivity at all. That includes certain government research environments, some clinical trial environments, and high-security research facilities.
Orion runs air-gapped without modification. Nothing reaches outside at runtime. The model weights are stored locally. The container images are pulled once from an external registry during setup and then kept in a local registry, and from then on nothing needs the outside network.
Updates are the hard part in an air-gapped environment. Security patches, model updates, and software updates all have to cross the air gap. The usual pattern is a staging environment with external connectivity where updates are tested and validated, then packaged and moved into the air-gapped environment through approved transfer methods.
Orion’s deployment model fits that pattern. Updates ship as new container images or Helm chart versions. You validate them in staging and carry them into the air-gapped environment as a bundle. The update never needs the air-gapped side to reach out.
What this takes to set up
The technical pieces exist and are production-tested. The harder part is organizational. You have to decide what the access control model should be, which workload types need audit logging and at what level of detail, how model updates get approved and rolled out, and who owns the infrastructure versus who uses it.
Those questions have answers, and the answers are different at every institution. The infrastructure can support several configurations, and the real work is picking the right one for your regulatory context. Start with one study and one workload type, and get the audit trail right there first.
If you are working out what compliant local AI infrastructure looks like for your research environment, we are happy to go through the specifics. Get a Demo
Anthony Genova is a product engineer at Juno Innovations, building the automation stack and agent frameworks that run on top of Orion.
WRITTEN BY
Anthony Genova
Anthony Genova builds the things that make Orion work. As a product engineer at Juno Innovations, he has spent years writing the automation stack, agent frameworks, and infrastructure tooling that run on top of Kubernetes so researchers, developers, and creative teams do not have to think about what is underneath. Before Juno, he worked across software engineering and systems integration. He writes about agentic systems, workload automation, and the infrastructure patterns that make AI actually useful in production.
