What does this workload cost?
Drop a vLLM result. Get the cost, latency and a report you can hand to a customer. Files stay in this browser.
Count only results that meet your requirements
By default, “completed” means a successful response, not a correct task. Optional gates below change the cost denominator; absent request samples hold that calculation.
Gates are applied to the same successful requests, not multiplied percentages. Evaluation sidecars bind per-request pass/fail to an exact source hash. Their author remains responsible for the correctness criterion.
Your report
Import a result on either side. A comparison file is optional.
Price-only break-even calculator
At the same allocation size, Hot Aisle needs the following share of a comparator’s throughput to tie GPU-rental cost. These are price ratios, not measured speed.
Prices are per GPU-hour. Confirm the purchasable instance size. This is AI Cloud infrastructure, not Token Factory managed endpoints. Custom rates and whole-bill charges can be entered above.
Get a compatible result from your existing server
Add these flags to your existing, approved vllm bench serve command. Keep your actual model, request set and offered load.
--save-result --save-detailed \
--percentile-metrics ttft,tpot,itl,e2el --metric-percentiles 50,95,99 \
--result-filename workload-result.jsonThis page does not start benchmarks, install runtimes, download models or rent GPUs. Detailed results may contain generated text and errors. Those fields are excluded from reports. Pin your runtime and retain the original file privately.
Measurement plan · Supported formats and evaluation sidecars
Verify a source file against a report
Select a saved Evidence JSON and any original result file. Hash matching proves byte identity with the report’s commitment, not the producer’s honesty.
Assumptions, sources and what stays private
Cost starts with the allocation you enter and the sum of the benchmark windows. Setup, restarts, idle time, storage and taxes outside those windows are excluded unless you enter them. A whole-bill override divides your supplied charge by the outcomes in the imported set.
Repeated runs are weighted by measured duration. Percentiles are pooled only when every trial has successful-request samples; otherwise per-trial ranges stay separate. Samples from failed requests never improve latency statistics. Summary goodput is recorded but never relabeled as correctness.
Model, revision, precision, tokenizer, workload, cache and offered-load identities must match before a savings claim appears. Entered identities are declarations; imported benchmarks are producer evidence. There is no automatic independent-verification badge.
Files remain in memory and clear on reload. Reports omit prompts, responses, error bodies, original filenames and unknown metadata. They retain selected model/runtime identifiers, normalized timings and file hashes; review these identifiers before sharing. Provider names confer no endorsement.