Skip to content

One platform. Two complementary features.

Analyse predicts what a workload needs before it runs.

Diagnose explains what happened after it runs. Together they close the loop between prediction and outcome: each execution makes the next prediction more accurate. On serving fleets, the same loop runs behind the scenes, with nothing for engineers to run.

Expanse Analyse

  • Memory, runtime and failure risk, before the queue
  • Predictions as distributions, with confidence
  • Same submit script, no workflow change
Expanse analyses a GROMACS workload and presents GPU, memory, runtime, completion probability, failure risk, and right-sizing evidence alongside the execution view.
Product demonstration
How it fits together

From submission to certainty.

Expanse sits in the path a job already takes. It does not replace a single step in your pipeline.

  1. 1

    Developer

    Submits a job as usual

  2. 2

    Analyse predicts

    Runtime, memory, GPUs, risk

  3. 3

    Scheduler executes

    SLURM, Kubernetes, or Nomad

  4. 4

    Diagnose explains

    Logs, metrics, exit state

  5. 5

    Cluster improves

    Feeds the next prediction

Serving

Always on for
serving fleets.

Expanse learns from everything that runs, batch or serving. The difference on a serving fleet is delivery: nobody runs a command. The same predictions work behind the scenes, sizing each model to what it actually needs under real traffic: more models multiplexed per GPU, and workloads reshuffled across hardware types without the guesswork.

  • Nothing to run

    No commands, no instrumentation. Expanse watches the fleet and characterises every model automatically.

  • Multiplex with confidence

    Each model’s slice is sized to what it actually uses, so models pack tighter without interfering.

  • Build on top of it

    The full prediction distribution is exposed through the API, so your own schedulers and simulators can consume it.

Before

Overprovisioned GPUs, wasted capacity, fewer workloads per node.

GPU 0

chat + chat

SM
BW
Breach

GPU 1

chat + chat

SM
BW
Breach

GPU 2

embed + rerank

SM
BW
Breach

GPU 3

embed + rerank

SM
BW
Breach

With Expanse

Right-sized models, higher GPU utilisation, more workloads on the same hardware.

GPU 0

chat + embed

SM
BW
OK

GPU 1

chat + embed

SM
BW
OK

GPU 2

chat + rerank

SM
BW
OK

GPU 3

chat + rerank

SM
BW
OK
Deployment

Installs into the infrastructure you already run.

  • One command to install

    An Ansible playbook sets up the cluster side; a Helm chart deploys the models to your GPUs.

  • Same submit script

    Jobs are submitted exactly as they are today. No new tooling for engineers or researchers to learn.

  • Cloud, on-prem or hybrid

    On-prem clusters, your own cloud accounts (AWS and others), or both. Your code and telemetry never leave infrastructure you control.

  • Your scheduler keeps scheduling

    SLURM, Kubernetes, and Nomad stay exactly where they are. Expanse sizes the requests; the scheduler executes them.

Integrations

Works with the scheduler youalready have.

Kubernetes logo

Kubernetes

One command installs Expanse on your existing SLURM, Kubernetes, or Nomad cluster. It automatically collects detailed telemetry alongside your workloads, bringing complete infrastructure visibility without changing your existing workflows.

SLURM logo

SLURM

Expanse arrives pretrained on real HPC and AI workloads, then continuously adapts to your environment. Every workload improves future predictions, making the platform more accurate over time.

Nomad logo

Nomad

Run Analyse before execution to predict GPU allocation, memory usage, runtime, and completion probability. Every prediction is presented as a confidence distribution, helping teams make infrastructure decisions with greater certainty.

Kubernetes®, Nomad® and other marks are trademarks of their respective owners. Use of these marks does not imply endorsement or affiliation.

Security & privacy

Private by design.

  • Self-hosted

    Expanse deploys inside your cluster: the models on your GPUs, the console and management layer on your servers. There is no hosted third-party environment.

  • You can see everything it sees

    Every signal Expanse collects is visible in your console: the telemetry, the job records, the predictions. Nothing the model uses is hidden from your team.

  • No workload data leaves the cluster

    Code, logs, telemetry, and job data remain inside your infrastructure. Nothing is uploaded or processed externally.

  • We’ve run clusters like yours

    Built by engineers with experience running large-scale HPC and AI infrastructure, Expanse is designed for environments where data never leaves the cluster.

Common questions

FAQ

01 Does Expanse replace my scheduler?

No. SLURM, Kubernetes, and Nomad stay in control. Expanse supplies evidence-backed resource requests; your scheduler continues to execute them.

02 How do we know the predictions are accurate?

Every prediction is presented as a distribution with confidence, then compared with the observed execution. The evidence and outcome stay visible to your team.

03 What happens when a prediction is wrong?

The observed outcome feeds the next prediction. Expanse keeps the original recommendation and the execution evidence together, so the model can adapt without hiding uncertainty.

04 How long until it’s accurate on our cluster?

Expanse starts producing confidence-bounded recommendations immediately and becomes specific to your environment as completed workloads feed back into the model.

05 Does my code leave the cluster?

No. Code, logs, telemetry, and job data remain inside infrastructure you control. The models and management layer are deployed in your environment.

06 We already monitor our cluster. What’s different?

Monitoring shows what is happening now. Expanse predicts resource fit before execution and connects the observed outcome to a cited explanation afterwards.

07 We built something like this internally. Why Expanse?

Expanse maintains the prediction, evidence, and feedback loop across schedulers and workload types, while your engineers keep the submit paths they already use.

08 Can I use Analyse without Diagnose?

Yes. Analyse and Diagnose can be used independently. Together, the observed diagnosis improves the evidence available to future Analyse predictions.

09 Which schedulers and clouds are supported?

Expanse integrates with SLURM, Kubernetes, and Nomad across on-premises clusters, your own cloud accounts, or hybrid environments.

Ready to make
compute certain?

Talk to our team about how Expanse fits your cluster, your scheduler, and your workloads.