Most people assume running a large language model takes a data center and an API contract.
Dan wanted to grow our AI agent setup by adding more workers to our development workflow. Every new agent meant a bigger API bill, more rate limit juggling, and a per-request cost that added up fast once agents were running all day. We were paying for intelligence by the request, and that does not scale.
So we decided to stop paying per request and run the model ourselves.
The hardware that made this realistic was the NVIDIA DGX Spark. Its GB10 Grace Blackwell chip has 128GB of unified memory shared between the CPU and GPU on a single die. That design is the difference between “we need a server rack” and “we need a desktop box.” Models that would need several separate GPUs on traditional hardware fit in the Spark’s unified memory pool.
We loaded a 35-billion parameter Mixture-of-Experts model, quantized to FP8 to save memory, and pointed a vLLM inference server at it. The whole thing runs on one device. Four agents share it at once, with no quota, no rate limits, and no data leaving our infrastructure.
The model, and why the numbers matter
The model is Qwen3-35B-A3B, a Mixture-of-Experts design with 35.6 billion total parameters and roughly 3 billion active parameters per token. MoE is what makes it fast. The model only wakes up a fraction of its parameters for any given token, so each request does far less work than a dense model of similar ability.
FP8 quantization cuts the memory further. The weights load at 8-bit floating point instead of the standard 16-bit or 32-bit. For most tasks the quality loss at FP8 is minimal and the memory savings are big.
On the Spark’s 128GB of unified memory, we give half to the KV cache, which is the context memory that holds the ongoing conversation for each active session. That gives us 838 cache blocks at 1,056 tokens per block, or roughly 885,000 tokens of total cache capacity. The maximum context length per session is 262,144 tokens.
That context window is what really matters for research work. 262,144 tokens holds an entire research paper, a big codebase, or a large experimental dataset in one context. An agent analyzing genomics output does not have to summarize and squeeze things to fit a small window. It works with the full dataset.
The inference stack
vLLM v0.23.0 runs the inference server. It speaks the OpenAI-compatible API, so any tool built for the standard AI APIs connects to the local endpoint by swapping the base URL and nothing else.
We deploy vLLM as an Orion Apps plugin. An IT administrator fills in the model name, a Hugging Face token if the model needs one, the GPU requirement, and the memory utilization target. The plugin does the rest. It pulls the model weights, configures the vLLM server, sets up the Kubernetes resources, and puts the endpoint at a stable internal address. Model weights live on a local-path NVMe PVC for fast access. The whole thing deploys from the Orion catalog the same way any other workload does.
If a department needs a different model, they install a second vLLM plugin instance with a different configuration. Each instance picks up available GPU capacity on its own, and several models can run at the same time across several instances.
What four agents on one model looks like
This is the part I think is really cool.
We run four AI agents at once, all connected to the same vLLM endpoint. Each agent is a persistent container with its own Longhorn storage for state, its own NFS workspace mount with its POSIX identity applied, its own webhook endpoint for outside triggers, and its own WebSocket connection for live status.
These are full automation instances, a lot more than chatbots. An agent can take a webhook from an outside system, run the incoming data through the local model, write results to the shared workspace, trigger a workflow in n8n, wake another agent with a webhook, and coordinate on tasks that would take several back-to-back calls against a cloud API. Because they share the same filesystem with the right permissions, they hand work to each other by writing to shared paths the next agent reads.
n8n runs as a standard Orion catalog item, with a visual workflow builder and 400-plus integrations. Our production workflows include processing meeting transcripts and pushing summaries to Notion, watching pipeline outputs and kicking off analysis agents, coordinating multi-step research workflows across collaborators, and handling routine report generation from experimental data.
MCP (Model Context Protocol) gives agents access to outside tools and data through a standard integration layer. An agent that needs to query a database, pull from an instrumentation API, or talk to an institutional system can do it through MCP without custom integration work.
Why this matters past our own setup
The research computing case for local inference has two parts, and they stand on their own.
The first is cost. An institution running agents all day, or offering AI tools to a large research community, is looking at API bills most budgets cannot carry for long. Local inference on a Spark, once you own the hardware, costs electricity. Where you break even against API pricing depends on how much you use it, but for workloads that run all day it comes fast.
The second is data governance, and for a lot of institutions it is not optional. Genomics data, clinical research data, unpublished results, and proprietary datasets cannot leave the institution’s infrastructure. The moment inference runs on an external API, the data has left. Local inference removes that problem completely. The model runs on your hardware and the data stays in your environment.
For institutions in air-gapped environments or with strict data governance rules, local inference is the only way to run AI on sensitive data.
The DGX Spark as the anchor
This setup is where a longer story ended up. It started with a Minecraft server on a single box and arrived at four AI agents sharing a 35-billion parameter model on a desktop supercomputer. The hardware changed and the idea stayed the same. Make whatever you have work as one system, and get the infrastructure out of the way so the people doing the work can focus on the work.
The Spark is the right anchor for a research AI cluster. Its unified memory handles models that would otherwise need server-class hardware. Orion turns it from one person’s workstation into a resource the whole institution shares, and the researchers using it never have to feel like they are sharing.
If you are thinking about what a local AI inference setup looks like for your institution, we are happy to work through the specifics. Book a demo
Alex Hatfield is the CEO and co-founder of Juno Innovations. Juno builds Orion, the customer-hosted unified compute plane. Orion runs inside your own environment, air-gapped by design, and gives your people one place to use the compute you already have, from GPUs and CPUs to VMs and bare metal.
WRITTEN BY
Alex Hatfield
Alex co-founded Juno to fix how enterprise compute gets done. He leads product vision and customer strategy, and spends most of his time working directly with infrastructure and research teams pushing the limits of what their hardware can do.
