Compute - Inference API

Run open-weight models through dedicated, managed endpoints.

Deploy production inference without operating the serving stack yourself. Hivenet Inference API gives you an OpenAI-compatible endpoint, dedicated replica capacity, and regional deployment on Hivenet infrastructure.

OpenAI-compatible API

Dedicated endpoints

1-4 replicas

Per-replica pricing

Billed by the second

Regional deployment

Qwen · Llama · Mistral · Falcon · GPT OSS · Gemma

Managed inference with dedicated capacity and a familiar API.

Running inference yourself gives you control, but it also leaves your team operating GPUs, serving engines, routing, authentication, monitoring, and capacity. Managed model APIs remove most of that work, but often give you less control over infrastructure placement and how you pay for sustained workloads.

Hivenet Inference API sits between those two approaches.

Hivenet operates the serving layer.

Choose a model and the capacity you need. Hivenet runs the infrastructure, router, serving environment, endpoint, and operational layer behind it.

The API experience stays familiar.

Use an OpenAI-compatible endpoint with common client patterns. Point your application at the Hivenet base URL, authenticate with an API key, and call the model.

The cost follows dedicated capacity.

Pay for the replica capacity behind the endpoint rather than a separate meter for every input and output token. Runtime is billed by the second.

For teams running production inference at real volume.

Hivenet Inference API is for teams that want to use open-weight models in production without becoming responsible for the infrastructure that serves them.

Production teams with steady inference traffic

Use dedicated capacity for workloads where request volume is sustained enough that reserving the endpoint makes sense.

SMBs with growing AI bills

Move suitable production workloads from per-token APIs to infrastructure-backed pricing you can see before you deploy.

Teams with residency needs

Choose from available deployment locations and keep the endpoint tied to the selected infrastructure location.

Developers using OpenAI-compatible tooling

Keep familiar request patterns with the OpenAI SDK and other compatible clients instead of rebuilding the application around another proprietary API.

Managed endpoint or compute instance?

Use Hivenet Inference API when you want the endpoint. Use Compute with Hivenet when you want the machine and full control over the runtime. Use Hivenet Router when you want to operate the routing layer yourself.

Need

Use

Why

I want an OpenAI-compatible endpoint

Hivenet Inference API

Hivenet operates the serving layer and endpoint

I want predictable dedicated inference capacity

Hivenet Inference API

Choose replica capacity and pay for its runtime

I want to run vLLM, SGLang, llama.cpp, PyTorch, custom models, or fine-tuning myself

GPU/CPU rental with Hivenet

You control the instance, runtime, and software stack

I want to route across inference engines I operate myself

Hivenet Router

You run the open-source routing layer and control its policies

I need a custom AI system around sensitive data or unusual deployment requirements

Private AI

Hivenet can help scope the model, data, infrastructure, and support path

Swipe left to see more

Learn more about our products

Built for the AI workloads teams actually run.

Open-weight models can handle a wide range of production tasks when the model is matched to the workload and tested against real data.

RAG

Serve retrieval-augmented generation workflows for internal knowledge, customer support, documentation, and business data.

Structured extraction

Extract dates, entities, categories, and structured fields from documents, messages, tickets, invoices, or records.

Summarization

Summarize documents, conversations, support threads, research, and operational content.

Classification

Classify messages, records, tickets, documents, and workflow inputs through a dedicated endpoint.

Code assistance

Run code-related workflows where the available model meets your quality, context, latency, and cost requirements.

Internal tools

Build assistants, automation, and internal AI features without operating the inference stack behind them.

From model to endpoint in a few steps.

Inference API is available from the Hivenet Compute console. Choose the deployment you need; Hivenet handles the serving environment behind it.

1

Choose a location

Select from the deployment locations currently available for your workload.

2

Choose a model and variant

Pick a model from the managed catalog, then choose the available variant that fits your context, quality, performance, and cost requirements.

The variant determines the underlying serving configuration. You do not need to configure the GPU, precision, or serving engine yourself.

3

Set replica capacity

Choose between 1 and 4 replicas, subject to available capacity.

The console shows the expected throughput and price as you adjust the deployment.

4

Review and deploy

Name the endpoint and review the model, location, replicas, hourly rate, and projected cost before starting it.

5

Create a key and connect

Generate an API key for your organization and use the deployment URL with an OpenAI-compatible client.

Dedicated endpoint controls you can see.

The endpoint exposes the important operational choices instead of hiding them behind a generic API plan.

Capacity

1–4 replicas

Set the replica count when you deploy and change it later without rebuilding the endpoint.

Billing

Per second

Replica runtime is metered by the second rather than rounded into a monthly commitment.

Integration

OpenAI-compatible

Use familiar OpenAI-compatible clients and request patterns with the Hivenet endpoint.

Routing

One router per customer

Your deployments share a customer-specific router rather than a single router shared with unrelated customers.

Lifecycle

Start · stop · terminate

Stop capacity when you do not need it, start it again later, or terminate the deployment when the workload is finished.

A curated model catalog for production workloads.

Hivenet Inference API starts from a managed catalog rather than an empty GPU machine. Exact models, variants, and locations depend on current availability and benchmark readiness.

Model family

Good starting point for

What to evaluate

Qwen

Extraction, RAG, structured output, multilingual workloads

Context, quality, throughput, structured output

Llama

RAG, summarization, assistants, internal tools

Context, latency, tool support, workload fit

Mistral

Instruction workloads, summarization, tools

Context, throughput, tool support

Falcon

Efficient general-purpose inference

Model size, throughput, workload fit

GPT OSS

General instruction and reasoning workloads

Quality, model size, latency

Gemma

Compact model workloads and experimentation

Model size, quality, throughput

Swipe left to see more

Qwen for production extraction and RAG

Qwen is a strong starting point for teams testing structured extraction, RAG, and production workflow automation.

Llama and Mistral workloads

Run widely adopted model families for RAG, summarization, internal tools, and model-serving experiments.

Need your own weights?

The managed catalog is the starting point for Inference API.

If you need custom weights, fine-tuned models, or direct control over the serving environment, use Compute with Hivenet or talk to Hivenet about a custom deployment.

Predictable pricing for dedicated endpoints.

Hivenet Inference API uses per-replica pricing billed by the second. You pay for provisioned inference capacity while it is running, rather than paying separately for each token generated.

Per-replica pricing

The model variant determines the infrastructure configuration and rate behind each replica.

Billed by the second

Pay for actual replica runtime. Stop the deployment when you no longer need the capacity and billing pauses once it has stopped.

Tokens included

Input and output tokens are included in the dedicated endpoint rate for the current product.

API calls and egress included

There is no separate charge for API calls or egress in the current dedicated endpoint pricing model.

Tier

Example use

Price

1 × RTX 5090

Smaller models and variants that fit a single GPU

from €1.10/hr

2 × RTX 5090

Production variants needing more memory or throughput

from €2.10/hr

4 × RTX 5090

Larger or higher-capacity model variants

from €3.80/hr

Swipe left to see more

Talk to sales about pricing

Regional infrastructure for dedicated AI endpoints.

Hivenet Inference API runs on Hivenet-operated, Policloud-backed infrastructure. Deployment location remains part of the product rather than disappearing behind an abstract API endpoint.

One location in the current release

Choose an available location for your first Inference deployment. Subsequent deployments currently inherit that location.

Single-tenant routing

One router per customer keeps the routing path isolated from unrelated customer traffic and easier to understand operationally.

Full-stack operation

Hivenet operates the router, serving runtime, endpoint infrastructure, authentication layer, monitoring, and billing behind the managed service.

Policloud logotype

We own the hardware

Run inference on a Policloud-backed path instead of routing production AI entirely through default hyperscaler APIs.

Test quality and throughput on your real workload.

Inference performance depends on the model, variant, context, request shape, output length, replica count, and traffic pattern. A benchmark is useful only when those conditions are visible.

Benchmarks with context

See what was actually tested

Hivenet benchmark results should show the model, variant, hardware configuration, request rate, and workload conditions behind the result.

Model evaluation

Your prompts decide the fit

Test the model against your real inputs, expected outputs, quality requirements, latency target, and traffic before moving production volume.

Endpoint metrics

Watch the signals that matter

Track requests, tokens, latency, errors, cost, and time to first token where available.

Keep your client code familiar.

Hivenet Inference API uses an OpenAI-compatible API so existing integrations can usually keep the same client pattern and change the endpoint configuration.

Point the client at your deployment

Each deployment has its own endpoint URL. Use it as the base_url for the client.

Use an organization API key

Create a Hivenet Inference API key and use it to authenticate calls to your organization's endpoints.

Send the model actually deployed

Specify the served model name in the request rather than relying on a generic default.

# Use the OpenAI client, pointed at Hivenet
from openai import OpenAI

client = OpenAI(
   api_key="HIVENET_API_KEY",
   base_url="https://api.hivenet.example/v1"
)

response = client.chat.completions.create(
   model="qwen-example",
   messages=[{"role": "user", "content": "Summarize this document."}]
)

Built around real production needs.

Hivenet Inference API is designed for production workloads where dedicated capacity, model choice, cost visibility, and infrastructure placement matter.

Production workload

High-volume document automation

A business automation customer uses a dedicated Qwen endpoint for part of a production extraction workflow.

The useful question is not whether an open model can replace every API call. It is whether a tested model can handle a specific production workload well enough to move that traffic onto dedicated capacity.

Best-fit buyer

Teams with meaningful, steady API spend

Dedicated endpoints make the most sense when a team already has production traffic and can identify workloads with sufficiently predictable demand.

Variable experiments and small sporadic workloads may be better suited to a different infrastructure path.

Talk through your use case

Need a different AI infrastructure path?

Hivenet Inference API is the managed endpoint path. If you want to operate more of the stack yourself, or need a different data and infrastructure setup, use the Hivenet product that matches the job.

GPU/CPU rental

Rent RTX 5090 or vCPU instances when your team wants full control over the instance, framework, and serving stack.

Private AI

Work with Hivenet on guided AI projects involving sensitive data, model choice, deployment support, or custom requirements.

RAG

Build retrieval systems on your own data using Hivenet's AI and storage paths.

S3-compatible storage

Store datasets, documents, and AI pipeline artifacts with S3-compatible tools and free egress.

FAQ

Common questions

Move one production AI workload to Hivenet.

Bring your current API usage, model needs, latency target, region requirements, and quality bar. We'll help you decide whether a managed foundational model endpoint is the right path.

Shader gradient background

PoliCloud + Hivenet

30% Off Hivenet Plans!

PoliCloud, powered by Hivenet’s technology, is redefining sovereign cloud storage. To celebrate our partnership, we’re offering 30% off all Hivenet plans—for a limited time!

*Offer ends March 31, 2025. Don't miss out!

Read our Terms & Conditions