Private beta · Dedicated inference

Your model, served. Everything underneath, operated.

Bring an open model, private weights, a fine-tune, or a complete Docker image. Nestor deploys it to a private, OpenAI-compatible endpoint on dedicated GPUs — and operates the runtime, networking, monitoring, and incident response for the life of the deployment.

Dedicated GPUsOpenAI-compatible APIBenchmarked and quoted per workloadOperated by Nestor
What's included ↓
What you get

A private endpoint that behaves like a product, on capacity that's yours.

  • Private HTTPS endpoint — your-model.endpoints.nestor.software
  • OpenAI-compatible API — point your existing client at a new base URL
  • API-key authentication
  • Dedicated GPUs — no shared pool, no rate limits, no noisy neighbors
  • Persistent model storage for weights, adapters, and caches
  • Health checks, logs, and endpoint metrics
  • A direct operator channel — the people who deployed it are the people who answer
What you can bring

A model, a runtime, or a whole container.

Models

Hugging Face models, private weights, fine-tuned checkpoints. Text, vision, audio, embedding, and reranking workloads.

Runtimes

vLLM by default. SGLang, TensorRT-LLM, TGI, or a runtime you specify.

Containers

Complete Docker images from public or private registries — custom CUDA dependencies, environment variables and encrypted secrets, your entrypoint and port, persistent volumes.

If it serves HTTP and runs on a GPU, we can likely operate it. Unusual stacks are scoped during workload review rather than turned away.

What Nestor operates

Six layers between your model and the hardware.

Run by Nestor for the duration of the commitment — so your team stays on the model, not the machine underneath it.
01
Capacity & hardware selection

Your model and traffic shape mapped to the most cost-efficient suitable GPU. You buy an outcome; hardware choice is our problem.

02
Runtime

Engine configuration, quantization, batching, context length, parallelism — tuned to your latency and throughput targets.

03
Serving

TLS, authentication, routing, and the endpoint itself.

04
Storage

Model weights, adapters, and caches on persistent volumes.

05
Monitoring

Endpoint health, throughput, latency, and GPU telemetry — watched by us, visible to you.

06
Operations

Restarts, redeploys, runtime upgrades, and incident response through a channel with a human on the other end.

How it works

From workload to a private endpoint, priced on measurement.

  1. Step 01
    Workload review

    You share the model or container, expected traffic shape, and latency target.

  2. Step 02
    Benchmark & quote

    We run your workload on candidate hardware and quote a price per million tokens at a stated throughput, with a monthly minimum. The number comes from measurement on the GPUs you'll actually run on — not a rate card.

  3. Step 03
    Deploy & accept

    Endpoint goes live on dedicated capacity. Your team validates before billing starts.

  4. Step 04
    Operate

    Nestor runs the deployment. Scaling, model updates, and changes are scoped through your channel.

When dedicated beats serverless

For traffic that's outgrown the per-token meter.

Sustained production volume.

At real traffic, per-token serverless list pricing is the expensive way to buy inference. A dedicated endpoint typically lands materially below it — with capacity that's guaranteed.

Latency you can plan around.

No cold starts, no rate limits, no sharing the batch with strangers.

Models serverless won't host.

Fine-tunes, private weights, custom architectures, and full custom containers.

Isolation.

Your prompts and weights never touch shared infrastructure. Single-tenant by construction.

Pricing

Quoted per workload. You'll know the number before anything is deployed.

A monthly GPU commitment plus a management fee, expressed as price per million tokens at your benchmarked throughput. Custom container deployments carry a higher operations fee than managed models — the surface area is larger and we staff it.

No self-serve meter, no surprise bill.

Bare Metal

Need the machines themselves?

Some teams want the endpoint. Some want the GPUs underneath it. The same workload can move between the two — start with a managed endpoint and take over the capacity later, or the reverse.

Bare Metal →
Dedicated inference · Private beta

Running a model that's outgrown serverless?

Tell us the model and the traffic. We'll benchmark it and come back with a quote within a few business days.

nestor/compute

Scope a deployment

Tell us what you're trying to run.

We'll follow up within one business day.

By submitting, you agree to be contacted about your request.