Verified benchmarks

Verified benchmarks

Hardware-class numbers you can trust — weekly auto schedule on a core L4 model set, Angestrom Verified on the Decision Card. Expand with BENCHMARK_TARGET_SET=expanded when ready.

Angestrom Verified

Measured on one hardware class

TTFT and throughput for a core high-demand model set on NVIDIA L4 · 24GB VRAM. Weekly auto schedule refreshes stale runs; A10/T4 activate when Modal/E2B endpoints are configured.

0verified models
0%pipeline coverage
Pipeline warming0/5 targets covered

Weekly schedule: 03:00 UTC on Sun · runner auto · set core · runner offline

Set BENCHMARK_VLLM_BASE_URL and/or BENCHMARK_MODAL_ENDPOINTS / BENCHMARK_E2B_ENDPOINTS; use BENCHMARK_RUNNER=auto|local|cloud

NVIDIA L4 · 24GB VRAM · VerifiedNVIDIA A10 · 24GB VRAM (planned)NVIDIA T4 · 16GB VRAM (planned)
No verified runs yet. From infra/ run ./run-benchmarks.sh (auto → local vLLM or cloud), or seed with BENCHMARK_RUNNER=fixture ALLOW_BENCHMARK_FIXTURES=1.

L4 Verified: BENCHMARK_RUNNER=local + BENCHMARK_VLLM_BASE_URL. A10/T4 Verified: set BENCHMARK_MODAL_ENDPOINTS / BENCHMARK_E2B_ENDPOINTS to OpenAI-compatible /v1 URLs, then BENCHMARK_RUNNER=cloud. Multi-probe averages keep the badge precise.