← Blog
July 31, 2025

RTX 5090s are now live on Compute

The next rung on the GPU ladder

In July 2025, Hivenet added NVIDIA RTX 5090 capacity to Compute. This page preserves that launch record and points to the corrected benchmark and current access information.

Why the RTX 5090 joined the lineup

When Hivenet launched with RTX 4090 GPUs, they gave teams a consumer-GPU option for inference and other workloads that fit within the card's memory and feature limits. The RTX 5090 launch added 32 GB of GDDR7 memory and a newer GPU generation. The first launch batch also included capacity in the UAE-2 region.

This is historical launch context. Regions, presets, capacity, and prices can change, so current availability should be checked in the Compute console.

Benchmark highlights at a glance

The launch benchmark compared one RTX 4090, one RTX 5090, and one A100 80 GB with Llama 3.1 8B Instruct in BF16. In the recorded test:

  • At 1 request/s, the RTX 5090 recorded 45.41 ms average TTFT and 6,058.57 ms average end-to-end latency. The A100 recorded 296.44 ms and 7,080.9 ms, respectively.
  • At the high-load setting, the RTX 5090 recorded 3,802.09 output tokens/s and the A100 recorded 3,748.16 output tokens/s, a difference of about 1.4% in this run.
  • The RTX 4090 recorded 737.65 output tokens/s in the same high-load table.

No two-GPU RTX 5090 run was recorded. The earlier 7,604 tokens/s figure was a linear extrapolation and is not presented as a measured result. The previous launch chart has also been removed because it showed a different tensor/pipeline-parallel test without enough method detail to reconcile it with the cited benchmark PDF.

Read the full RTX 4090, RTX 5090, and A100 benchmark setup, results, and limits before applying these numbers to another workload.

How we ran the tests

  • Model: meta-llama/Meta-Llama-3.1-8B-Instruct
  • Precision: BF16
  • Dataset: ShareGPT
  • Context and output: 8,192-token context and 512-token output
  • Engine: vLLM 0.8.3 benchmark_serving.py
  • Scenarios: 1 request/s with 100 prompts; 1,100 requests/s with 1,500 prompts
  • Launch regions recorded in the article: France and UAE-2

The historical record does not include the exact A100 form factor, host CPU and system RAM, NVIDIA driver and CUDA versions, full command and flags, test date, run count, variance, or error behavior. The current Hivenet benchmark methodology explains the reporting standard used for newer tests.

What it means for your workload

This benchmark supports a narrow conclusion: for this Llama 3.1 8B BF16 setup, the RTX 5090 recorded much lower TTFT than the A100 and similar high-load throughput. It does not prove a universal performance, energy-efficiency, cost, or multi-GPU scaling advantage.

The RTX 4090 has since been retired from Hivenet's launchable fleet. For current RTX 5090 pricing and availability, read the current pricing and availability article, then check the active preset in the Compute console.

How to launch an RTX 5090 now

The current Compute quickstart covers account setup, credits, instance creation, location and GPU selection, connectivity, and cleanup. Use that guide rather than the historical console steps or screenshots on this page.

Historical July 2025 screenshot of the Compute console

For the active price and launchable configurations, use the current Hivenet Compute pricing page and the console.

Your next workload belongs on Hivenet.

Pick one AI, compute, or storage workload and see the difference for yourself. Spin it up in minutes, or let our team map your fastest path to production.

Shader gradient background