Hugging Face models, private weights, fine-tuned checkpoints. Text, vision, audio, embedding, and reranking workloads.
Your model, served. Everything underneath, operated.
Bring an open model, private weights, a fine-tune, or a complete Docker image. Nestor deploys it to a private, OpenAI-compatible endpoint on dedicated GPUs — and operates the runtime, networking, monitoring, and incident response for the life of the deployment.
A private endpoint that behaves like a product, on capacity that's yours.
- Private HTTPS endpoint — your-model.endpoints.nestor.software
- OpenAI-compatible API — point your existing client at a new base URL
- API-key authentication
- Dedicated GPUs — no shared pool, no rate limits, no noisy neighbors
- Persistent model storage for weights, adapters, and caches
- Health checks, logs, and endpoint metrics
- A direct operator channel — the people who deployed it are the people who answer
A model, a runtime, or a whole container.
vLLM by default. SGLang, TensorRT-LLM, TGI, or a runtime you specify.
Complete Docker images from public or private registries — custom CUDA dependencies, environment variables and encrypted secrets, your entrypoint and port, persistent volumes.
If it serves HTTP and runs on a GPU, we can likely operate it. Unusual stacks are scoped during workload review rather than turned away.
Six layers between your model and the hardware.
Your model and traffic shape mapped to the most cost-efficient suitable GPU. You buy an outcome; hardware choice is our problem.
Engine configuration, quantization, batching, context length, parallelism — tuned to your latency and throughput targets.
TLS, authentication, routing, and the endpoint itself.
Model weights, adapters, and caches on persistent volumes.
Endpoint health, throughput, latency, and GPU telemetry — watched by us, visible to you.
Restarts, redeploys, runtime upgrades, and incident response through a channel with a human on the other end.
From workload to a private endpoint, priced on measurement.
- Step 01Workload review
You share the model or container, expected traffic shape, and latency target.
- Step 02Benchmark & quote
We run your workload on candidate hardware and quote a price per million tokens at a stated throughput, with a monthly minimum. The number comes from measurement on the GPUs you'll actually run on — not a rate card.
- Step 03Deploy & accept
Endpoint goes live on dedicated capacity. Your team validates before billing starts.
- Step 04Operate
Nestor runs the deployment. Scaling, model updates, and changes are scoped through your channel.
For traffic that's outgrown the per-token meter.
At real traffic, per-token serverless list pricing is the expensive way to buy inference. A dedicated endpoint typically lands materially below it — with capacity that's guaranteed.
No cold starts, no rate limits, no sharing the batch with strangers.
Fine-tunes, private weights, custom architectures, and full custom containers.
Your prompts and weights never touch shared infrastructure. Single-tenant by construction.
Quoted per workload. You'll know the number before anything is deployed.
No self-serve meter, no surprise bill.
Need the machines themselves?
Some teams want the endpoint. Some want the GPUs underneath it. The same workload can move between the two — start with a managed endpoint and take over the capacity later, or the reverse.
Running a model that's outgrown serverless?
Tell us the model and the traffic. We'll benchmark it and come back with a quote within a few business days.