EVIDENCE / BENCHMARK PROTOCOL

Performance claims
need a reproduction path.

SPARK does not publish universal latency, throughput, GPU-efficiency, or cost-reduction numbers without the model, hardware, workload, and measurement method needed to interpret them.

Review the protocol
PUBLIC RESULT STATUS
NO INDEPENDENTLY VERIFIED PRODUCTION BENCHMARK PUBLISHED

Workload-specific reports may be produced during evaluation. A result moves to this public page only after disclosure and publication rights are approved.

MINIMUM DISCLOSURE

Six parts of a
reviewable result.

Any future benchmark table will link to a test record containing these fields. A chart without them is treated as illustrative UI, not proof.

01

Model artifact

Exact model name, revision, parameterization, quantization or precision, tokenizer, and serving configuration.

02

Hardware

Accelerator model and count, CPU and memory profile, interconnect, region, and tenancy.

03

Traffic

Input/output token distributions, context length, concurrency, request rate, warm-up, cache state, and duration.

04

Metrics

Time to first token, inter-token latency, end-to-end latency by percentile, throughput, errors, and cost basis.

05

Quality

Task-specific dataset, scoring method, evaluator version, human-review process, and confidence interval.

06

Reproduction

Date, code or configuration digest, exclusions, raw-output retention, and reviewer approval.

REPORT TEMPLATE

Compare against
the customer baseline.

FIELDBASELINESPARK CANDIDATESTATUS
Quality scoreRequiredRequiredNot published
P50 / P95 latencyRequiredRequiredNot published
Output tokens / secondRequiredRequiredNot published
Error and timeout rateRequiredRequiredNot published
Cost basisRequiredRequiredNot published