Bare Metal · Dedicated GPU infrastructure

Dedicated GPUs, nodes, and clusters.

Deployed with our base stack — OS, NVIDIA drivers, orchestration, and monitoring — and operated by Nestor. No spot pool, no shared capacity, no ticket queue. You run the workload; we run the infrastructure underneath it.

Dedicated single-tenant nodesNVLink & multi-node fabricSLURM or KubernetesRoot access, operated by Nestor
View GPU capacity ↓
nestor / managed environment
env-7af · live
Workload path
01
Workload
02
Managed env
03
GPU cluster
04
API · jobs · rollouts
Status
  • Capacity reserved
    8× H200 · region us-east
  • Drivers configured
    CUDA 12.6 · NCCL 2.21
  • Cluster online
    slurm · 8 nodes · healthy
  • Monitoring active
    gpu · node · job
  • Support channel open
    shared · 1 business day
operated by nestoruptime · 14d
Why Nestor

For workloads a console and a credit card can't solve.

Self-serve GPU clouds work until they don't: spot interruptions, shared pools, generic environments, and a support ticket queue when infrastructure blocks model work. Nestor is the other model — a dedicated environment scoped to your workload, operated by a team you can actually reach.
Dedicated hardware.

GPUs reserved for you for the full term — no spot pool, no shared capacity, no noisy neighbors.

Configured around your workload.

Drivers, runtime, orchestration, storage, and access patterns tuned to your model and framework — not a generic image.

A human operator.

Direct support channel for onboarding, environment changes, and incident triage.

What you get

The machine, not a slice of one.

Dedicated hardware, wired and tuned for GPUs that saturate — from a single node to a multi-node cluster, scoped to your workload and operated by Nestor.
Single-tenant nodes
Whole GPUs reserved for your term — root access, your kernel modules, your containers. No hypervisor slice, no shared tenant, no noisy neighbors.
Multi-node fabric
NVLink inside the node and a low-latency fabric across nodes, sized to the cluster — so distributed training scales instead of stalling on communication.
High-throughput storage
Local NVMe plus attached shared storage, provisioned to keep the GPUs fed — checkpoints, datasets, and caches without an I/O bottleneck.
Managed orchestration
SLURM or Kubernetes, configured and operated by Nestor — job scheduling, multi-node launches, and health-aware rescheduling.
Base stack, preconfigured
OS, NVIDIA drivers, CUDA, NCCL, containers, and monitoring installed and tuned before you log in. You get a cluster that's ready, not a bare rack.
Built for

The jobs a shared endpoint can't run.

Bare Metal is for the workloads you operate yourself — compute-heavy, long-running, and multi-node. You run the job; Nestor runs the infrastructure underneath it.
Distributed training

Multi-node runs across NVLink and fabric, NCCL-tuned and checkpointed, with health-aware rescheduling when a node drops.

Fine-tuning & post-training

Full-parameter or adapter runs on dedicated capacity — your framework, your data pipeline, your schedule.

RL / continual training

Generation, reward, eval, and training loops on capacity planned for bursty, multi-stage post-training.

Large-scale batch

Embeddings, evaluations, and data processing at fleet scale — long jobs on GPUs that are yours for the term.

Self-managed inference

Run your own serving stack when you want the GPUs, not a managed endpoint. Want us to run the model instead? That's Inference Endpoints.

Managed scope

What Nestor manages.

Six operational layers between your workload and the hardware — scoped during onboarding and run by Nestor for the duration of the commit.
Layer 01
Capacity planning

Workload mapped to GPU type, cluster size, term, and region.

Layer 02
Provisioning

Dedicated environments, not a shared pool.

Layer 03
Orchestration

Managed SLURM or Kubernetes.

Layer 04
Runtime setup

CUDA, drivers, containers, frameworks, inference engines.

Layer 05
Networking & storage

Private access, attached storage, data movement.

Layer 06
Monitoring & support

GPU health, utilization, node status, and a direct channel.

Infrastructure

GPU capacity.

Capacity Nestor deploys through its partner network — scoped, contracted, and operated per deployment.

GPUMemoryInterconnectTypical use
NVIDIA B300
Blackwell Ultra. Built for trillion-parameter training.
288 GB HBM3eNVLinkFrontier training, large-scale post-training, high-memory inference.
NVIDIA B200
Blackwell. Native FP8 / FP4 throughput.
192 GB HBM3eNVLinkTraining, FP8 inference, large model serving.
NVIDIA H200 SXM
Hopper. 1.4× HBM vs H100, ideal for long-context workloads.
141 GB HBM3eNVLinkLong-context inference, fine-tuning, memory-heavy workloads.
NVIDIA H100 SXM
Hopper. NVLink, distributed training proven.
80 GB HBM3NVLinkDistributed training, fine-tuning, high-throughput inference.
NVIDIA A100 SXM
Ampere. Proven workhorse for training and inference.
80 GB HBM2eNVLinkFine-tuning, research, batch inference.
NVIDIA RTX Pro 6000 Blackwell
Blackwell. 96 GB GDDR7. Fits 70B FP8 on a single card.
96 GB GDDR7PCIeCost-efficient inference, 70B-class serving, development.
NVIDIA L40S
Ada. 48 GB GDDR6. Steady inference and lightweight fine-tunes.
48 GB GDDR6PCIeInference, embeddings, lightweight fine-tuning.
Onboarding

From workload to running infrastructure.

A direct path from architecture review to a stable, supported environment — scoped with your team before any capacity is committed.
  1. Step 01
    Architecture review
    You share workload, model, framework, and timing; we scope GPU config, orchestration, and term.
  2. Step 02
    Environment setup & acceptance
    We provision, configure, and hand off; your team validates before billing starts.
  3. Step 03
    Ongoing operation
    Nestor stays on as the infrastructure operator: monitoring, changes, support, and future capacity.
Security & access

Private environments with controlled access.

Environments are scoped to one customer. Access patterns, network boundaries, and container provenance are agreed during onboarding and operated by Nestor.
  • Environments scoped to a single customer
  • SSH-key access; VPN or private connectivity where available
  • Private containers and custom images supported
  • Export control and acceptable-use restrictions apply
Note

Nestor is in private beta. We do not currently claim SOC 2, ISO, HIPAA, or other compliance certifications. Security and compliance roadmap is available on request.

Specific networking, storage, or access requirements are scoped with your team during onboarding rather than assumed by default.

Operated by

Who operates Nestor

Nestor is operated from New York by an infrastructure team with backgrounds building GPU cloud and edge infrastructure products at AWS and leading GPU cloud providers. We work directly with a vetted network of data center and capacity partners across the US.
Contact
sales@nestor.software·We reply within one business day.
Inference Endpoints

Want the model served instead of the machines delivered?

Bring a model or a container and we serve it from a private, OpenAI-compatible endpoint on dedicated GPUs — and operate everything underneath.

Inference Endpoints →
Onboarding selected teams

Running something a self-serve cloud can't hold?

Nestor is onboarding selected AI teams in private beta. Tell us what you're trying to run — we'll follow up within one business day.

nestor/compute

Scope a deployment

Tell us what you're trying to run.

We'll follow up within one business day.

By submitting, you agree to be contacted about your request.