Two GPU cards of NVIDIA’s newest Blackwell generation are now available in the OneCloudPlanet cloud: RTX PRO 4500 Blackwell (32 GB GDDR7) and RTX PRO 6000 Blackwell (96 GB GDDR7). Both come with hourly billing and preinstalled NVIDIA drivers and CUDA: an instance is created from the console in minutes and is ready to work right away.

The optimal entry point into the newest generation — $0.69/hr (in a working-day 8×5 mode that is ≈$121/mo).

What it fits: inference of open 7–8B LLMs (Llama 3.1 8B, Qwen 8B) and up to ~14B quantized — corporate chatbots and RAG over internal documents; image generation (SDXL, Flux); rendering, video processing and ML experiments. 32 GB of fast GDDR7 is comfortable headroom where previous card generations were already tight.

The headline feature — 96 GB of VRAM in one card: quantized 30–70B-class models (Llama 3.1 70B, Qwen 72B) run without multi-GPU setups and the sharding complexity they bring. Also: multimodal models like Qwen3-VL-32B, long-context RAG, LoRA/QLoRA fine-tuning of mid-size models.

Where the 70B class used to require several cards — now it is a single instance. Current prices — on the Cloud GPU page.

The fastest path — ready images from the Solution Marketplace: Ollama, DeepSeek, Llama or the Dify AI platform — a working stack in minutes, with no driver battles. Billing is hourly: an evening experiment costs a couple of dollars, not a monthly budget.

Not sure which card fits your model — the “which GPU for which task” cheat sheet is on the AI / LLM hosting page.