
This page reports one Llama 3.1 8B Instruct BF16 serving benchmark. The results apply to the recorded setup and load profiles, not to every model or AI workload.
We compared one RTX 4090, one RTX 5090, and one A100 80 GB using vLLM 0.8.3, ShareGPT prompts, an 8,192-token context, and a 512-token output. The test covered a moderate-load latency scenario and a high-load throughput scenario.
The measured single-GPU results are useful for this setup. They do not establish a general winner across other models, precisions, context lengths, training jobs, or multi-GPU systems.
Without repeated runs and variance data, the small throughput difference between the RTX 5090 and A100 should be read as a result from this run, not proof of a consistent advantage.
Hivenet's current benchmark methodology and reporting standard explains the fields expected in newer benchmark reports.
| Parameter | Value |
|---|---|
| Context length | 8,192 tokens |
| Output length | 512 tokens |
| Model | meta-llama/Meta-Llama-3.1-8B-Instruct |
| Precision | BF16 |
| Batch size | Automatic, based on GPU memory |
| Dataset | ShareGPT |
| Benchmark tool | vLLM 0.8.3 benchmark_serving.py |
| Attribute | Value |
|---|---|
| Request rate | 1 request/s |
| Number of prompts | 100 |
| Goal | Measure average TTFT and end-to-end latency |
| Attribute | Value |
|---|---|
| Request rate | 1,100 requests/s |
| Number of prompts | 1,500 |
| Goal | Measure output-token throughput and sustained requests/s |
| GPU | Avg ITL (ms) | Avg TPOT (ms) | Avg TTFT (ms) | Avg E2E latency (ms) |
|---|---|---|---|---|
| RTX 4090 | 19 | 19 | 349.9 | 9,759.07 |
| RTX 5090 | 12.14 | 12.14 | 45.41 | 6,058.57 |
| A100 80 GB | 13.25 | 13.25 | 296.44 | 7,080.9 |

| GPU | Avg output-token throughput (tokens/s) | Sustained requests/s |
|---|---|---|
| RTX 4090 | 737.65 | 1.47 |
| RTX 5090 | 3,802.09 | 7.58 |
| A100 80 GB | 3,748.16 | 7.58 |

This is a historical benchmark, and the surviving record does not contain every field required by Hivenet's current methodology. The exact A100 form factor, host CPU and system RAM, NVIDIA driver and CUDA versions, complete command and flags, test date, run count, variance, and error or saturation behavior were not recorded in the article or source PDF. We have not filled those gaps with assumptions.
The test covers one 8B model in BF16 with fixed context and output lengths. It does not establish results for quantized models, other model sizes, training, fine-tuning, longer contexts, different batching or concurrency, or multi-GPU scaling.
For this Llama 3.1 8B BF16 serving setup, the RTX 5090 recorded much lower TTFT than the A100 and similar high-load throughput. The RTX 4090 was slower in both recorded scenarios. Hardware selection still depends on memory capacity, reliability features, software support, concurrency, and the workload you actually plan to run.
For broader RTX 4090 workload fit, see the RTX 4090 AI guide. For current Hivenet availability and price context, use the current RTX 4090 retirement and RTX 5090 pricing article and confirm the active preset in the Compute console.
Pick one AI, compute, or storage workload and see the difference for yourself. Spin it up in minutes, or let our team map your fastest path to production.