hostfleet Find my setup
AI infrastructure, made visible

Host AI.
Without guessing.

See what runs where, what hardware it needs, and what it can cost — before you deploy anything.

Sourced33 GPU types · 21 providers · verified 2026-09-24
33GPU types tracked
21providers compared
3hosting paths explained
0fake benchmarks
Start with the workload

What are you trying to run?

Pick the closest shape. We will separate the app server, the agent runtime, and the model so you do not buy a GPU for the wrong job.

Your first decision

Do you need to own model serving?

Use an API for the fastest start. Rent serverless GPU for custom weights or bursty jobs. Keep a GPU warm only after the workload proves it needs one.

Open the model hosting guide
Your app
→
Model endpoint
GPU belongs here — not under the web app.
The core choice

Three ways to use a model

There is no universal winner. Traffic shape, control, and engineering time change the answer.

API↗
01

Use a model API

Best first move for most apps. No GPU setup, scaling, or idle hardware.

  • Best for prototypes and variable usage
  • You manage prompts and product code
  • Tradeoff less infrastructure control
Start here unless you know why not
GPUϟ
02

Rent serverless GPU

Bring custom weights and pay for active compute. Cold starts and minimums still matter.

  • Best for bursty jobs and custom models
  • You manage image and inference code
  • Tradeoff startup delay and limits
Good bridge to self-hosting
24/7▦
03

Keep a GPU running

Predictable capacity and maximum control — with a bill that continues while idle.

  • Best for steady, measured workloads
  • You manage the complete serving stack
  • Tradeoff cost and operations
Earn this complexity with data
EstimatedPut your traffic into the calculatorCompare usage-based API, serverless GPU, and always-on GPU monthly cost.
Calculate my cost →
VRAM, not vibes

What GPU fits the model?

These are conservative first-test targets for 4-bit weights, short context, and one active sequence — capacity estimates, not performance benchmarks.

Smaller modelMore VRAM →
Qwen3 8B8B parameters
16 GB
T4from $0.59/hr
Mistral Small 24B24B parameters
24 GB
L4from $0.44/hr
Qwen3 32B32B parameters
32 GB
RTX 5090from $0.25/hr
Llama 3.3 70B70B parameters
48 GB
A40 / A6000from $0.35/hr
Sourced + estimated

Model sizes come from official model cards. Raw weight math and first GPU targets are estimates with headroom. Context, concurrency, runtime, and quantization can require more.

Read the sizing method →
Useful before you deploy

Tools, not another wall of text

Every tool exposes its inputs, claim type, source date, and limitations.

Deep research

Go deeper only when you need to

The short visual answer comes first. These longer pages keep the source trail, assumptions, and implementation detail available for search, GEO, and serious buyers.

ai hosting Serverless GPU pricing 2026: H100 rates and what scale-to-zero actually bills Source-checked H100 rates and idle-tail rules, now with Cloud Run GPU instance billing and an explicit unpriced scale-to-zero boundary. 2026-10-04 · Read the evidence → ai hosting What GPU do you need to run Llama 70B? VRAM, context, and KV-cache guide Llama 70B GPU VRAM and KV-cache math, selected cloud costs, and why a listed Koyeb B200 price is not deployable-capacity evidence. 2026-10-06 · Read the evidence → ai hosting Best hosting for AI agents on a budget (June 2026): choose by workload, not by AI branding A June 26, 2026 HostFleet refresh on budget AI agent hosting, split by scheduled jobs, always-on workers, and small self-hosted stacks. 2026-07-10 · Read the evidence → deploy ai apps Where to deploy your Lovable, Bolt, or v0 app: a decision guide (April 2026) Lovable, Bolt.new, and v0 all generate working apps — but the hosting story is different for each one. A feature-by-feature guide to what each platform supports and where your code actually runs. 2026-04-21 · Read the evidence → deploy ai apps What breaks when AI-generated apps hit production: documented footguns A field guide to the failure modes you'll hit when you ship an app written by Lovable, Bolt, v0, or Cursor. Synthesized from public GitHub issues, Reddit threads, and vendor docs — every claim linked. 2026-04-21 · Read the evidence → ai hosting What it costs to run an AI side project on a VPS for 30 days (June 24, 2026): honest budget ranges A practical June 24, 2026 guide to what an AI side project really costs on a VPS for 30 days, using current Hostinger, DigitalOcean, and Hetzner pricing plus explicit workload assumptions. 2026-06-24 · Read the evidence →