Skip to content

Private inference · Applied AI

Inference on hardware you can point at

Open-weight models served from a cluster we own and operate — plus the engineering to put them to work. No third-party hop, no training on your prompts, no per-token meter that spikes the month a feature lands.

Models we serve

  • GLM-5.2

    Zhipu AI

  • DeepSeek

    DeepSeek AI

  • MiniMax-H3

    MiniMax

  • 20+

    more models

Open weights throughout, plus any model you already run — fine-tunes included.

Models

Open weights, served fast

Every model on the cluster is open weights. If you leave, the model comes with you — what you are buying is the operation, not access to something you cannot take elsewhere.

Flagship

GLM-5.2

Zhipu AI

General reasoning, long context, tool use

Code

DeepSeek

DeepSeek AI

Code generation, maths, structured output

Long context

MiniMax-H3

MiniMax

Long-form generation and multilingual work

Custom

Bring your own

Any open weights

We host and tune the model you already trust

Running something else? We host open-weight models you already depend on, fine-tunes included.

Inference

A base URL change, not a migration

The API is OpenAI-compatible. If your code already talks to a chat-completions endpoint, it talks to ours — the integration is a configuration change, and the benchmark tells you whether it was worth making.

OpenAI-compatible API

Point your existing client at a new base URL. If your code already talks to a chat-completions endpoint, it talks to ours.

Dedicated capacity

Reserve whole units rather than sharing a queue. Throughput you can plan around instead of a rate limit you discover in production.

Your data does not train anything

Prompts and completions are not retained for training and are not passed to a third-party provider. The hardware is ours.

Open weights, no lock-in

Every model we serve is open weights. If you leave, the model comes with you — the weights are not the product, the operation is.

Predictable cost

Reserved capacity is billed as capacity, not as a per-token meter that spikes the month a feature goes viral.

Run it in your own racks

The same stack can be installed on your hardware, in your building, with us operating it or handing it over.

Consulting

The hard part is rarely the model

Most teams do not need a better model. They need retrieval they can inspect, agents that fail safely, and a way to tell whether last week’s change helped or hurt.

Local inference

Specify, install and tune a cluster in your own environment — model selection, serving stack, quantisation, throughput.

  • Hardware sizing
  • Serving stack
  • Quantisation
  • Benchmarks

RAG and retrieval

Get answers grounded in your own documents, with retrieval you can inspect and evaluation that tells you when it regresses.

  • Chunking and indexing
  • Hybrid retrieval
  • Grounding
  • Eval harness

Agent orchestration

Multi-step agents that call your systems: tool design, guardrails, retries, and the boundary between deterministic code and the model.

  • Tool design
  • MCP servers
  • Guardrails
  • Cost control

Evaluation and monitoring

Know whether a change helped. Task-specific evals, regression suites, and monitoring that surfaces drift before a user does.

  • Golden sets
  • Regression suites
  • Drift alerts
  • A/B harness

How it works

Benchmark first, quote second

We price against what your workload actually did on the cluster, not against what we hope it will do.

  1. 01

    Scope

    A call to establish what you are running, what it costs today, and whether we are the right answer. Often we say no.

  2. 02

    Benchmark

    We run your workload — your prompts, your documents — on the cluster and show you throughput, latency and quality against your current setup.

  3. 03

    Deploy

    Reserved capacity on our cluster, or the same stack installed in your racks. Either way you leave with something running.

FAQ

Questions we get asked

How is this different from a hosted API?
The hardware is ours and the models are open weights, so nothing is passed to a third-party provider and nothing is retained for training. You reserve capacity rather than sharing a queue, which makes throughput something you plan rather than discover.
Can we run this on our own hardware?
Yes. The same serving stack can be specified, installed and tuned in your environment — we either operate it for you or hand it over with documentation and a benchmark suite.
Which models can you serve?
The models listed above, plus any open-weights model you already depend on. If you have fine-tuned something, we host it as-is.
How is it priced?
Reserved capacity is priced per unit, per month, against the throughput your benchmark showed. Consulting is scoped per engagement. Both are quoted after the benchmark rather than before it.
What does an engagement look like?
Scope, benchmark, deploy. The benchmark uses your workload, not ours, and it is where most of the useful conversation happens — including the cases where the answer is that your current provider is already fine.

Bring us a workload and we’ll benchmark it

Your prompts, your documents, our cluster — measured against whatever you run today. If your current provider is already the right answer, we will tell you that.