Private inference · Applied AI
Inference on hardware you can point at
Open-weight models served from a cluster we own and operate — plus the engineering to put them to work. No third-party hop, no training on your prompts, no per-token meter that spikes the month a feature lands.
Models we serve
GLM-5.2
Zhipu AI
DeepSeek
DeepSeek AI
MiniMax-H3
MiniMax
20+
more models
Open weights throughout, plus any model you already run — fine-tunes included.
Models
Open weights, served fast
Every model on the cluster is open weights. If you leave, the model comes with you — what you are buying is the operation, not access to something you cannot take elsewhere.
GLM-5.2
Zhipu AI
General reasoning, long context, tool use
DeepSeek
DeepSeek AI
Code generation, maths, structured output
MiniMax-H3
MiniMax
Long-form generation and multilingual work
Bring your own
Any open weights
We host and tune the model you already trust
Running something else? We host open-weight models you already depend on, fine-tunes included.
Inference
A base URL change, not a migration
The API is OpenAI-compatible. If your code already talks to a chat-completions endpoint, it talks to ours — the integration is a configuration change, and the benchmark tells you whether it was worth making.
OpenAI-compatible API
Point your existing client at a new base URL. If your code already talks to a chat-completions endpoint, it talks to ours.
Dedicated capacity
Reserve whole units rather than sharing a queue. Throughput you can plan around instead of a rate limit you discover in production.
Your data does not train anything
Prompts and completions are not retained for training and are not passed to a third-party provider. The hardware is ours.
Open weights, no lock-in
Every model we serve is open weights. If you leave, the model comes with you — the weights are not the product, the operation is.
Predictable cost
Reserved capacity is billed as capacity, not as a per-token meter that spikes the month a feature goes viral.
Run it in your own racks
The same stack can be installed on your hardware, in your building, with us operating it or handing it over.
Consulting
The hard part is rarely the model
Most teams do not need a better model. They need retrieval they can inspect, agents that fail safely, and a way to tell whether last week’s change helped or hurt.
Local inference
Specify, install and tune a cluster in your own environment — model selection, serving stack, quantisation, throughput.
- Hardware sizing
- Serving stack
- Quantisation
- Benchmarks
RAG and retrieval
Get answers grounded in your own documents, with retrieval you can inspect and evaluation that tells you when it regresses.
- Chunking and indexing
- Hybrid retrieval
- Grounding
- Eval harness
Agent orchestration
Multi-step agents that call your systems: tool design, guardrails, retries, and the boundary between deterministic code and the model.
- Tool design
- MCP servers
- Guardrails
- Cost control
Evaluation and monitoring
Know whether a change helped. Task-specific evals, regression suites, and monitoring that surfaces drift before a user does.
- Golden sets
- Regression suites
- Drift alerts
- A/B harness
How it works
Benchmark first, quote second
We price against what your workload actually did on the cluster, not against what we hope it will do.
- 01
Scope
A call to establish what you are running, what it costs today, and whether we are the right answer. Often we say no.
- 02
Benchmark
We run your workload — your prompts, your documents — on the cluster and show you throughput, latency and quality against your current setup.
- 03
Deploy
Reserved capacity on our cluster, or the same stack installed in your racks. Either way you leave with something running.
FAQ
Questions we get asked
- How is this different from a hosted API?
- The hardware is ours and the models are open weights, so nothing is passed to a third-party provider and nothing is retained for training. You reserve capacity rather than sharing a queue, which makes throughput something you plan rather than discover.
- Can we run this on our own hardware?
- Yes. The same serving stack can be specified, installed and tuned in your environment — we either operate it for you or hand it over with documentation and a benchmark suite.
- Which models can you serve?
- The models listed above, plus any open-weights model you already depend on. If you have fine-tuned something, we host it as-is.
- How is it priced?
- Reserved capacity is priced per unit, per month, against the throughput your benchmark showed. Consulting is scoped per engagement. Both are quoted after the benchmark rather than before it.
- What does an engagement look like?
- Scope, benchmark, deploy. The benchmark uses your workload, not ours, and it is where most of the useful conversation happens — including the cases where the answer is that your current provider is already fine.
Bring us a workload and we’ll benchmark it
Your prompts, your documents, our cluster — measured against whatever you run today. If your current provider is already the right answer, we will tell you that.