Hot Aisle
Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8
$0.6908per 1,000 evaluator-passed, latency-qualified requests
- Result rate
- 1.202 / second
- Completed / attempted
- 8,622 / 8,622
- Trials / total measured seconds
- 1 / 3,606.58
- Allocation
- 1 GPUs × $2.99/hour + $0.0000/hour extras
- Cost basis
- modeled charge for summed benchmark windows; $3.00
- Price source / period
- campaign listed price; benchmark window model only / 2026-09-24
- Quote / cost note
- None
Successful-request latency (ms)
| p50 | p95 | p99 | |
|---|---|---|---|
| TTFT | 52.898 | 155.038 | 568.93 |
| pooled successful-request samples | |||
| E2E | 2,280.241 | 12,353.639 | 15,339.371 |
| pooled successful-request samples | |||
Conditions and input commitments
{
"identity": {
"model": "Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8",
"revision": "dcaee4d4dfc5ee71ad501f01f530e5652438fde0",
"precision": "FP8",
"tokenizer": "dcaee4d4dfc5ee71ad501f01f530e5652438fde0",
"workload": "sha256:6c508237082261ac927b96825e2f097782389b1cebc3f78360f0665a8db3501f",
"cache": "no-prefix-cache",
"load": "azure-code:54e9a6d2a4bd06ba1e060304b900abbc74cbea53de96506e60fe5bb4f2277fb6:2023-11-16T18:17:04Z:factor=0.95:duration=3600:seed=700000:workers=256:max_tokens=1024:temperature=0.2",
"runtime": "vllm/vllm-openai-rocm:v0.30.0@sha256:2e7da1ad1c66836802072588adea75f9f4991da5f9545b4318e91d422c22ce6a"
},
"gates": {
"ttft_ms": 1000,
"e2e_ms": 60000,
"include_client_queue": true,
"quality": true
},
"criterion": "evalplus-0.3.1-human-v0.1.10-mbpp-v0.2.0-base-and-plus-raw-completion-v1",
"holds": [],
"warnings": []
}Input commitments
SHA-256 ef907e4cbe45832890109c4ca0d37851ad8e4f087beb70421506af897e6a67ab
Record 0 · 605631 bytes