Early access open — join the waitlist

GPU orchestration — anywhere

The world is your cluster.

PanoFabric aggregates GPUs anywhere — your clusters, your cloud accounts, spot markets — into one fabric, and its performance model plans the fastest, cheapest way to run your training or inference job across it.

The problem

Compute is everywhere.
Yours is stuck in one place.

0×

Price spread, same GPU

The same H100 rents on-demand from $1.99 to $10.00 per GPU-hour depending on provider. You pay whichever one you happen to be locked into.

0%

Of GPU time is actually used

Teams believe their clusters are about 60% busy. When real clusters were measured, GPUs were doing work only 14% of the time — the rest sat idle while jobs queued elsewhere.

0days

Erased by one dead node

One node failure kills an unprotected three-week pretraining run. Restart, re-queue, re-burn.

The platform

One control plane for training and inference on every GPU you can reach.

Slurm clusters, Kubernetes, SSH boxes, cloud accounts and spot markets — enrolled once, planned together, run as one machine.

Performance-modeled launches

Don't guess your parallelism config.

  • The performance model searches the whole space — sharding, replicas, batch size, placement — before a single GPU-hour burns.
  • Set a cost target or a deadline; get the plan that hits it, with receipts.
$ panofabric run llama3-8b.yaml --dry-runBALANCED
est. cost$6,840
eta44 h
throughput21.4 k tok/s
spot · 4090
0
your Slurm · A100
0
cloud · H100
0
4 islands · quorum-protected −29% vs single-cloud

One fabric, every GPU

Enroll any GPU. Run everywhere.

  • Your own machines — SSH boxes, Slurm clusters, Kubernetes — alongside cloud and spot capacity.
  • Nodes enroll: no inbound firewall holes, no VPN.
  • One control plane, one queue, one bill view.
ssh fleet slurm cluster kubernetes aws · gcp · nebius spot markets FABRIC pretrain sft · rl inference eval

Fault-tolerant decentralized training

Training that survives the real world.

  • Pretraining, SFT and RL across regions and providers with fault tolerance.
  • Nodes join and leave mid-run; training continues.
  • A dead spot instance is a blip, not a restart.
ISLAND A · us-east ISLAND B · spot eu-west training loss — continues through failure
quorum 7/7 · stepping step 41,208

Decentralized inference

Serve from wherever is cheap and close.

  • Serve large models across cheap, scattered capacity instead of one premium region.
  • Placement-aware routing sends every request to the nearest healthy shard within latency budget.
  • The same performance model keeps the cost side down while you scale out.

Built for teams

A control plane your platform team will sign off on.

  • Multi-tenant with real auth, per-tenant credentials and workspace isolation.
  • A live dashboard for every run — steps, loss, utilization, spend.
  • Thin pip install panofabric CLI/SDK. Self-host it, or use our hosted plane.
app.panofabric.ai/runs
llama3-8b-pretrain running 4 islands · 64 GPUs · $118/h · step 41,208
qwen-7b-sft-dpo queued plan: 16× A100 (your Slurm) · est $312 total
loss · llama3-8b-pretrain
fabric utilization 92% · 118/128 GPUs busy

How it works

Three commands from scattered to woven.

Enroll panofabric enroll

Point it at your SSH boxes, Slurm login node, K8s context or cloud creds. Nodes dial out, egress-only, and appear in your fabric.

Describe job.yaml

Model, data, and a target — a budget or a deadline. Your training loop stays yours.

Run panofabric run

The perf model picks placement, parallelism and islands; the fabric executes, heals around failures, and streams you the receipts.

Where one provider stops being enough.

Six situations where the constraint isn't your model or your team — it's that your compute lives behind a single account, in a single region, under a single set of rules.

Blocked on quota

Multi-cloud capacity, because scarcity is structural

The accelerator you need is always sold out somewhere. Quota is rationed per provider and per region, and the queue is measured in weeks. Pool capacity across clouds, GPU providers and spot markets so a run lands on whatever is genuinely free — and let the performance model choose the mix instead of you refreshing a quota page.

Regulated & sovereign workloads

Your jurisdiction, your hardware, your control plane

Regulation, procurement rules and customer contracts increasingly dictate where a model may be trained and served. Pin workloads to named countries, providers or your own racks, keep data on infrastructure you control, and self-host the entire control plane. Nodes enroll egress-only — nothing dials in, and nothing leaves that you didn't send.

Research & public sector

Federate clusters without merging them

Every institute has a cluster, and every cluster is idle at a different hour. Join Slurm systems across departments, universities or national centres into one queue while each site keeps its own admins, allocations and hardware. No central procurement, no migration: enrollment is a login node and an outbound connection.

Fleet owners

Make the GPUs you already bought count

Most fleets are busy a fraction of the time while their own teams wait in a queue somewhere else. Enroll what you own, let the scheduler backfill the gaps across sites, and burst to rented capacity only for the overflow you genuinely can't cover.

Cost-sensitive training

Spot-first runs that survive preemption

Preemptible capacity is the cheapest compute on the market and the least trusted, because one eviction used to mean restarting the run. With fault-tolerant islands, an evicted node costs a step rather than a week — which turns spot from a gamble into a default.

Serving open models

Inference without the premium-region bill

Serving costs scale with traffic and with whichever region you happened to start in. Place replicas across cheap, scattered capacity and route every request to the nearest healthy shard that fits your latency budget.

Early access

We're onboarding a small number of design partners.

Tell us where your GPUs live and we'll tell you when it's your turn — or skip the line and talk to us this week.

Want to skip the line? Book a call →

You're on the list.

Check your inbox to confirm your spot. Know someone else wrangling GPUs? Send them this page — every referral bumps you up.

Skip the line — book a call

The questions that block a call.

Is my code and data safe on BYO nodes?

Your nodes stay yours: machines dial out to the control plane, nothing dials in, no inbound firewall holes, no VPN. Code and data stay on your infrastructure and your storage; the control plane sees scheduling metadata and the logs you choose to ship. Self-hosting the whole plane is also supported.

Which clouds and schedulers do you support?

BYO: plain SSH fleets, Slurm clusters, and Kubernetes. Cloud: the major providers and GPU clouds — AWS, GCP, Azure, Nebius, and any K8s cluster — plus spot capacity across all of them, in one fabric.

What happens when a node dies mid-run?

The islands reform without it and training keeps stepping; a replacement joins and syncs when capacity allows. A dead spot instance costs you seconds of progress, not the run.

Self-hosted or hosted?

Both. Run the control plane on your own infra, or use our hosted plane with multi-tenant auth, per-tenant credentials and workspace isolation. Same CLI and dashboard either way.

Which models and frameworks?

PyTorch first. Pretraining, SFT and RL for open models, LoRA and full fine-tunes, plus decentralized serving for OSS multimodal models.

What does it cost?

Early access — let's talk. Design partners get hands-on help planning their first runs and pricing that reflects being early. Book the call and bring a real workload.

Loose threads, everywhere.One fabric.

20 minutes. Bring a real workload — we'll plan it across the fabric live, and you keep the plan.

Book a call

PanoFabric early-access demo · 20 min

Bring a real workload. We'll enroll a node, plan your job across the fabric, and kill a worker live so you can watch it heal.

Request sent.

We'll reply within one business day with a couple of times that fit.

Prefer email? Write to hello@panocular.ai.