Multi-node runs across NVLink and fabric, NCCL-tuned and checkpointed, with health-aware rescheduling when a node drops.
Dedicated GPUs, nodes, and clusters.
Deployed with our base stack — OS, NVIDIA drivers, orchestration, and monitoring — and operated by Nestor. No spot pool, no shared capacity, no ticket queue. You run the workload; we run the infrastructure underneath it.
- Capacity reserved8× H200 · region us-east
- Drivers configuredCUDA 12.6 · NCCL 2.21
- Cluster onlineslurm · 8 nodes · healthy
- Monitoring activegpu · node · job
- Support channel openshared · 1 business day
For workloads a console and a credit card can't solve.
GPUs reserved for you for the full term — no spot pool, no shared capacity, no noisy neighbors.
Drivers, runtime, orchestration, storage, and access patterns tuned to your model and framework — not a generic image.
Direct support channel for onboarding, environment changes, and incident triage.
The machine, not a slice of one.
- Single-tenant nodes
- Whole GPUs reserved for your term — root access, your kernel modules, your containers. No hypervisor slice, no shared tenant, no noisy neighbors.
- Multi-node fabric
- NVLink inside the node and a low-latency fabric across nodes, sized to the cluster — so distributed training scales instead of stalling on communication.
- High-throughput storage
- Local NVMe plus attached shared storage, provisioned to keep the GPUs fed — checkpoints, datasets, and caches without an I/O bottleneck.
- Managed orchestration
- SLURM or Kubernetes, configured and operated by Nestor — job scheduling, multi-node launches, and health-aware rescheduling.
- Base stack, preconfigured
- OS, NVIDIA drivers, CUDA, NCCL, containers, and monitoring installed and tuned before you log in. You get a cluster that's ready, not a bare rack.
The jobs a shared endpoint can't run.
Full-parameter or adapter runs on dedicated capacity — your framework, your data pipeline, your schedule.
Generation, reward, eval, and training loops on capacity planned for bursty, multi-stage post-training.
Embeddings, evaluations, and data processing at fleet scale — long jobs on GPUs that are yours for the term.
Run your own serving stack when you want the GPUs, not a managed endpoint. Want us to run the model instead? That's Inference Endpoints.
What Nestor manages.
Workload mapped to GPU type, cluster size, term, and region.
Dedicated environments, not a shared pool.
Managed SLURM or Kubernetes.
CUDA, drivers, containers, frameworks, inference engines.
Private access, attached storage, data movement.
GPU health, utilization, node status, and a direct channel.
GPU capacity.
Capacity Nestor deploys through its partner network — scoped, contracted, and operated per deployment.
| GPU | Memory | Interconnect | Typical use |
|---|---|---|---|
NVIDIA B300 Blackwell Ultra. Built for trillion-parameter training. | 288 GB HBM3e | NVLink | Frontier training, large-scale post-training, high-memory inference. |
NVIDIA B200 Blackwell. Native FP8 / FP4 throughput. | 192 GB HBM3e | NVLink | Training, FP8 inference, large model serving. |
NVIDIA H200 SXM Hopper. 1.4× HBM vs H100, ideal for long-context workloads. | 141 GB HBM3e | NVLink | Long-context inference, fine-tuning, memory-heavy workloads. |
NVIDIA H100 SXM Hopper. NVLink, distributed training proven. | 80 GB HBM3 | NVLink | Distributed training, fine-tuning, high-throughput inference. |
NVIDIA A100 SXM Ampere. Proven workhorse for training and inference. | 80 GB HBM2e | NVLink | Fine-tuning, research, batch inference. |
NVIDIA RTX Pro 6000 Blackwell Blackwell. 96 GB GDDR7. Fits 70B FP8 on a single card. | 96 GB GDDR7 | PCIe | Cost-efficient inference, 70B-class serving, development. |
NVIDIA L40S Ada. 48 GB GDDR6. Steady inference and lightweight fine-tunes. | 48 GB GDDR6 | PCIe | Inference, embeddings, lightweight fine-tuning. |
From workload to running infrastructure.
- Step 01Architecture reviewYou share workload, model, framework, and timing; we scope GPU config, orchestration, and term.
- Step 02Environment setup & acceptanceWe provision, configure, and hand off; your team validates before billing starts.
- Step 03Ongoing operationNestor stays on as the infrastructure operator: monitoring, changes, support, and future capacity.
Private environments with controlled access.
- Environments scoped to a single customer
- SSH-key access; VPN or private connectivity where available
- Private containers and custom images supported
- Export control and acceptable-use restrictions apply
Nestor is in private beta. We do not currently claim SOC 2, ISO, HIPAA, or other compliance certifications. Security and compliance roadmap is available on request.
Specific networking, storage, or access requirements are scoped with your team during onboarding rather than assumed by default.
Who operates Nestor
Want the model served instead of the machines delivered?
Bring a model or a container and we serve it from a private, OpenAI-compatible endpoint on dedicated GPUs — and operate everything underneath.
Running something a self-serve cloud can't hold?
Nestor is onboarding selected AI teams in private beta. Tell us what you're trying to run — we'll follow up within one business day.