Our GPU colocation plans
Fixed-price plans sized for the power draw of a GPU server. Request a quote in one click and we get back to you within 24 business hours.
GPU Starter
A 1 to 2 GPU machine for inference, fine-tuning or image generation.
- Reserved power800 W
- Rack space2U
- Guaranteed bandwidth100 Mbps
- IPv4 addresses2
- 10 Gbps uplink portFree
GPU Pro
A 4U multi-GPU server, the most common format for AI in production.
- Reserved power1.5 kVA
- Rack space4U
- Guaranteed bandwidth250 Mbps
- IPv4 addresses4
- 10 Gbps uplink portFree
GPU Cluster
Several GPU servers side by side, connected in the same rack space.
- Reserved power3 kVA
- Rack space8U
- Guaranteed bandwidth500 Mbps
- IPv4 addresses8
- 10 Gbps uplink portFree
Custom
An 8-GPU server, more than 3 kVA or half a rack: build your own space.
- Rack space1 to 24U
- Power0.5 to 6 kVA
- Guaranteed bandwidth50 to 1,000 Mbps
- 10 Gbps uplink portFree
Build your custom space
10 Gbps uplink port, IPv6 /64 and redundant A/B power included. We send you a firm quote within 24 business hours.
- Tier III+ infrastructure (N+1 redundancy)
- IPMI / KVM monitoring on request
- Multi-carrier transit
- IPv6 /64 included on every plan
- Setup fee: €49.99 (waived with a 12-month commitment)
Why host your GPU server in a data center
A multi-GPU machine runs hot, makes noise and draws as much power as a heater. It belongs in a rack, not under a desk.
Reserved electrical power
Each plan reserves a power envelope, from 800 W to 3 kVA, on redundant A/B power: if one feed fails, the other takes over without interruption for your running training jobs.
Network to serve your models
Free 10 Gbps uplink port, multi-carrier transit, guaranteed bandwidth of 100 to 500 Mbps depending on the plan, dedicated IPv4 and IPv6 /64. Enough to expose an inference API or pull in your datasets.
Your data stays yours
The server belongs to you and stays in Aubervilliers, at the Equinix PA6 data center. Training data, model weights, prompts: nothing goes through a third-party cloud platform.
Remote control of your machine
IPMI/KVM access on request to reboot, reinstall an NVIDIA driver or switch kernels without travelling. Our team receives, racks and cables your hardware.
How much power does your GPU server need?
In GPU colocation, power determines the plan, well before rack space. To estimate the power to reserve: add up the maximum draw of each GPU and of the processor (CPU), then about 150 W for the motherboard, RAM, NVMe SSDs and fans. Then add a 20% margin to absorb load peaks.
| Graphics card | Video memory | Max. power draw | Typical uses |
|---|---|---|---|
| GeForce RTX 4090 | 24 GB GDDR6X | 450 W | LLM inference up to ~30B quantised parameters, Stable Diffusion, LoRA fine-tuning |
| GeForce RTX 5090 | 32 GB GDDR7 | 575 W | Same use as the 4090 with more video memory |
| RTX 6000 Ada | 48 GB GDDR6 | 300 W | AI and 3D rendering, dual-slot blower card suited to rack chassis |
| L40S | 48 GB GDDR6 | 350 W | Inference and image generation in production, multi-GPU servers |
| A100 PCIe | 80 GB HBM2e | 300 W | Training and fine-tuning of mid-sized models |
| H100 PCIe | 80 GB HBM2e | 350 W | Training and inference of large language models |
Example: a server with 2 RTX 4090s (2 × 450 W) and a 170 W processor draws at most 900 + 170 + 150 = 1,220 W. With a 20% margin, that comes to about 1.45 kVA: the GPU Pro plan fits. A single card of this type fits in GPU Starter. NVIDIA reference values; check the datasheet of your exact model, as some custom cards draw more.
GPU colocation, cloud, dedicated server or office?
Four ways to run your AI models, with very different costs, levels of control and flexibility.
| GPU colocation | Cloud GPU rental | Rented dedicated server or VPS | Server in the office | |
|---|---|---|---|---|
| Hardware | Yours, paid off over time | Rented, never owned | Rented from the host, rarely fitted with GPUs | Yours |
| Virtualisation | None: physical machine, direct GPU access | Often a virtual machine | VPS: shared virtual server; dedicated: none | None |
| Cost | Fixed monthly fee | Billed by the hour, adds up with continuous use | Monthly server rent | Electricity and air-conditioning bill |
| Administration | Full root access, you manage OS, firewall and backups | Through the provider's console | Root access, managed service depending on the plan | Everything is on you, on site |
| Power | Redundant A/B in the data center | Handled by the provider | Handled by the host | One socket, no backup |
| Network | 10 Gbps uplink, fixed IP addresses | Depends on the plan | Bandwidth depends on the plan | Business internet connection |
| Data | On your machine, in France | At the provider | At the host | On your premises |
| Best for | 24/7 use: inference, APIs, long training runs | One-off needs of a few hours | Workloads without GPUs, or no hardware to buy | Prototyping, a single card |
No hardware yet? Dedicated hosting remains an option: a dedicated server saves you the purchase. For a standard server without GPUs, our colocation plans start at 1U and 100 W. Still unsure? Our guide colocation or managed dedicated server compares both approaches, while data center colocation: how it works explains rack units, racks and connection.
Managing your colocated GPU server
ElypseCloud is your physical host, not your managed service provider: the machine stays fully under your control.
System and drivers
You have full root access. Install the distribution of your choice, for example Ubuntu Server or Debian, then the NVIDIA drivers, CUDA and, if you work with containers, Docker with the NVIDIA Container Toolkit.
SSH access and firewall
Key-based SSH login, password disabled, and a firewall (UFW or nftables) that only leaves SSH and your inference API port open. Our first security settings apply as they are.
Storage and backups
Model weights and datasets are best kept on NVMe SSDs. Backups remain your responsibility in colocation: keep a backup copy off the machine for anything you cannot download again.
What the data center provides
High availability of the infrastructure (A/B power, N+1 redundancy), guaranteed bandwidth, fixed IP addresses and IPMI/KVM access on request to take back control remotely.
What you can run in GPU colocation
LLM inference
Serve an open-source model (Llama, Mistral, Qwen…) with vLLM, Ollama or llama.cpp behind your own API, with no per-token billing.
Training and fine-tuning
Adapt a model to your data (LoRA, QLoRA) for days with PyTorch and CUDA, without watching an hourly meter.
Image and video generation
Stable Diffusion, Flux or ComfyUI in production for an application or a creative studio.
3D rendering and compute
Blender or Unreal render farm, simulation, scientific computing: anything that benefits from thousands of CUDA cores.
FAQ - GPU Colocation
Everything about colocating GPU and AI servers at the Equinix PA6 data center.