Why I did this
I had $10 of Hugging Face compute credits and a question I couldn’t find a straight answer to: everyone recommends “scale to zero to save money,” but almost nobody measures what that actually costs when the traffic comes back.
So I turned it into a weekend experiment. I deployed a dedicated Inference Endpoint, scaled it to zero three times, timed every recovery, and paid attention to everything the hourly bill doesn’t show. This post is the full write-up — method, raw numbers, and the cost math, including one claim from my first draft that turned out to be wrong.
The short version
- I deployed Qwen/Qwen3.8-27B (revision
1d4bf0f2) on 1x Nvidia A100 (80 GB) via Hugging Face’s vLLM catalog recipe ($2.50/hr, AWS us-east-1), scaled it to zero three times, and timed recovery. - In my tests, waking this deployment took ~6.6–10 minutes before a successful answer completed (median 6.6 min). Warm answers completed in about 2 seconds (TTFT median 1.49 s).
- Every wake-up began with 36–50 consecutive HTTP 503s — no request queue, hard failures until the replica was healthy. (The docs describe 502s; this deployment served 503s.)
- Boot time is billed. Estimated startup compute cost was ~
0.28–0.42 per wake. This recipe’s automatic 15-minute idle timeout adds a tail of billed idle time to every isolated request (worked example below). - Total experiment cost: ~$1.60 (my estimate from billed minutes; the dashboard Usage panel is the authoritative number).
How Hugging Face Inference Endpoints actually work
An endpoint is a managed container, not a magic API. HF provisions a VM at your chosen vendor/region, pulls a serving image (here vllm/vllm-openai:v0.27.1), downloads the weights from the Hub, health-checks it, and exposes a stable URL. You never touch Kubernetes, CUDA, or weight storage.
The lifecycle is pending → initializing → running, plus two “off” states: scaledToZero and paused.
Billing applies per replica-minute while initializing or running. The two off states cost $0 — but they are not the same:
scaledToZeroauto-wakes on the next request, with a cold start.pausedstays down until you explicitly resume it.
Scale-to-zero triggers on idleness, not utilization: this recipe fires after 15 minutes with no requests (configurable per deployment). And critically: there is no queueing during initialization — requests arrive and fail immediately. Client-side retry is your job.
The setup
| Item | Value |
|---|---|
| Model | Qwen/Qwen3.8-27B, revision 1d4bf0f2 (Hub catalog recipe) |
| Engine | vLLM 0.27.1 (managed container, port 8000, OpenAI-compatible /v1) |
| Hardware | AWS us-east-1, 1x Nvidia A100, 80 GB VRAM, 11 vCPUs, 145 GB RAM |
Engine flags (endpoint config via hf endpoints describe) | --gpu-memory-utilization 0.95, --max-model-len 262144, MTP speculative decoding (num_speculative_tokens: 3), tool parser qwen3_coder |
| Rate | $2.50/hr per replica, billed by the minute |
| Auth | Private (personal HF token, Authorization: Bearer) |
| Client | A small Python script — streaming chat completions against /v1/chat/completions |
Procedure: deploy → 10 warm streaming requests → hf endpoints scale-to-zero (verifying the scaledToZero state before each run) → poll until the first successful response, counting every 503 → repeat ×3 → pause.
How “boot time” is defined here — and its limits
- Boot time = from the first request sent after scale-to-zero until the first successful response fully completed (streaming, short prompt, 64 max tokens). This includes polling delay (10 s interval) and the final response, so it slightly overstates the time when the endpoint became ready. It is not “time to first HTTP 200.”
- The script stores its own run timestamps in local time; run 2’s boot time was reconstructed from separate UTC shell timestamps (first 503 → first 200) after my polling tool hit a 9-minute timeout while the endpoint was still booting. Runs 1 and 3 were measured by the script directly. The two methods agree within ~1 s where they overlap (run 1: script 398.4 s vs timestamps ~399 s).
- The script only speaks OpenAI-compatible streaming chat — it will not measure other endpoint types.
The results
Warm steady state (n=10)
| metric | value |
|---|---|
| TTFT median | 1.49 s (range 1.37–2.35 s) |
Cold boots (n=3)
| run | boot time | failed requests | TTFT after boot | boot cost |
|---|---|---|---|---|
| 1 | 398 s | 36 × 503 | 2.25 s | $0.28 |
| 2 | 601 s * | ~50 × 503 | 2.29 s | $0.42 |
| 3 | 399 s | 36 × 503 | 2.37 s | $0.28 |
* reconstructed from UTC timestamps (see above). All runs: identical config, same endpoint, same hour.
What n=3 can and can’t tell you
- Startup time varied by 50%. Two boots came in at ~399 s, one at 601 s. With three runs I cannot establish whether later boots benefit from caching, nor explain the outlier (plausible suspects: VM allocation, image pull, weight-download contention — the logs don’t say). The practical takeaway is defensive: don’t assume reboots get faster; size your retries for ~10 minutes.
- The failures are hard, not slow. Every request during initialization failed immediately with 503 (the docs describe 502 — on this deployment, this day, it was 503). There is no accept-and-queue behavior; availability is entirely the client’s problem.
- First request after boot sat near the upper end of the warm range. TTFT 2.25–2.37 s vs the 1.49 s warm median (warm range 1.37–2.35 s) — roughly 0.8 s above median. Whatever the residual warm-up is, it is marginal next to the 6.6–10 minute wall. I did not isolate its cause.
- Multimodal status is unclear. The catalog lists
image-text-to-text, and one startup logged “no registered multimodal processor; running in text-only mode.” I tested only text, so I can’t say whether images work — that warning alone doesn’t establish that they never do. - Engine choice decided whether the model ran at all. HF’s default engine (transformers toolkit) crashed on boot — its pinned
transformersversion didn’t recognize the newqwen3_5architecture (ValueError: ... model type 'qwen3_5' but Transformers does not recognize this architecture, from the failed deployment’s boot log). The vLLM catalog recipe ran it. On bleeding-edge models the “verified recipe” isn’t a convenience; it’s the difference between running and not running.
The cost math
Billing applies while initializing or running, by the minute. There are two distinct regimes, and they have different math.
Manual scale-to-zero (you trigger it right after use, no idle tail):
cost per wake = boot_time × rate # 399–601 s × $2.50/hr ≈ $0.28–$0.42
break-even gap = boot_time # ~6.6–10 min — below this, staying warm is cheaper AND fasterAutomatic scale-to-zero with a 15-minute idle timeout (this recipe’s default): after every request the endpoint stays warm for 15 more billed minutes ($0.625) before shutting down. An isolated request then costs roughly:
boot ($0.28–$0.42) + idle tail ($0.625) ≈ $0.90–$1.05 vs $2.50/hr for staying warmAs a simplified estimate, savings vs staying warm only materialize when traffic gaps exceed roughly 22–25 minutes — the 15-minute idle tail plus the 6.6–10-minute boot — where a “gap” runs from the previous completed response to the next request, the same clock the idle timer uses. Below that, the idle tail makes auto-scale-to-zero equal to or more expensive than never scaling down. The naive “break-even = boot time” claim — which my first draft of this post got wrong — only holds for the manual regime.
And the bill never shows the other cost: every wake is 6.6–10 minutes during which the endpoint simply does not exist for your users.
The engineering takeaway: the 503 problem
Since the server queues nothing during initialization, the client owns availability:
for attempt in range(max_retries):
r = requests.post(f"{url}/v1/chat/completions", headers=auth, json=payload, timeout=420)
if r.status_code == 200:
return r
time.sleep(10) # 503 while a new replica initializes (docs say 502 — retry both)When I’d use each option
| traffic pattern | best option |
|---|---|
| steady traffic | min replicas ≥ 1, never scale to zero |
| bursty, gaps well over the ~22–25 min threshold | automatic scale-to-zero + client retry with generous timeouts |
| dev/demo box you forget about | pause (scaledToZero still wakes — and bills — on traffic) |
| occasional single calls, no infra | serverless Inference Providers (pay per token) instead |
Caveats
n=3 boots, one model, one region (AWS us-east-1), one vendor, one day, one endpoint. Boot times depend on weight size, image state, VM allocation, and region; run 2 proves the variance is real for an identical config. The 503-vs-502 observation is scoped to this deployment. The total cost (~$1.60) is my reconstruction from billed minutes — the dashboard Usage panel has the exact figure.