# AMD-Pro-Bench: LLM benchmark methodology

Research snapshot: **2026-10-11**. Workload scope: **LLM inference, serving, and language-model training only**, including model-linked latency, throughput, power, memory and correctness measurements. Gaming, rendering, image/video generation and standalone compute kernels are excluded from published data and exports.

Hardware scope: discrete **AMD Radeon PRO / Radeon AI PRO workstation products introduced since 2021**, including desktop, mobile-workstation and Apple MPX variants. The catalog records evidence and qualifies announcement versus availability dates. Newer listed products with uncertain shipment dates are explicitly qualified. Integrated APUs, server Instinct/V-series accelerators and pre-2021 products are excluded from the catalog. Mixed systems containing an eligible card preserve the full reported inventory without attributing their total speed to that card alone.

## Coverage and limits

This is a broad research collection, not proof that every internet report has been found. The search included GitHub and Codeberg repositories, raw benchmark files, vendor documentation, independent reviews, Phoronix charts, LLM benchmark registries, Reddit, forums, and video descriptions. Known unparsed, inaccessible, duplicate, and nonnumeric sources are explicitly retained in `coverage`. Failed experiments are not represented as zero performance. Source numbers were checked against available evidence; hardware tests were not rerun.

## Data model

The primary filters are Card, GPU count, Developer, Model family, Version and Quantization. `model_series` identifies a release line such as Qwen3.8; `model_version` identifies its checkpoint or variant, such as 27B or Flash-Next. Developer choices are ranked by distinct family/version combinations matching the other active filters, rather than measurement count. `model_developer` and `model_name` are browsing labels derived from the existing reviewed repository mappings. Upstream developers are distinguished from quantization publishers, and documented fine-tunes retain their own developer and model names. Precision/packaging suffixes are separated from model names. Family-only links and conflicting identities retain explicit uncertainty; browsing labels do not upgrade their evidence status. The source-reported `model` and `quantization` fields remain unchanged in details, the advanced Reported model / file filter, and exports. Different three-bit formats are separate values. Model family alone is insufficient. Unresolved aliases remain explicitly unresolved; they are never silently promoted to a known checkpoint.

One benchmark row represents one numeric LLM metric. A single run may have several rows: prefill, decode, TTFT, power, etc. A source campaign can contain many runs and can share a repository or author with other campaigns. Row counts are not counts of independent replications.

`sources` describes documentation strength, independence, author, dates, access, URLs and methodology. `benchmarks` preserves exact reported workload/model names, quantization, participating GPU count, parallelization, OS, engine/build, backend, CPU/RAM, context capacity, actual prompt/output lengths where established, concurrency, statistic, measurement scope, units, extraction evidence and source locator. Extra configurations remain in `settings`. SQLite also provides `cards`, `flags`, `coverage` and `metadata`; full source and record payloads are retained as JSON.

Missing values remain null. A default setting of an engine, a current repository README, a physical device inventory, or an archive directory name is not automatically evidence for an individual run. Configured context, actual prompt length, maximum active sequences and observed concurrency are separate concepts. `settings.past_context_tokens` is pre-existing KV depth in llama-bench and is not context capacity.

## Source runs, repository links and multiple selections

Every row includes a `source_url`, a `run_id`, a portable `source_run_url` and a `source_artifact_url` when explicit raw evidence is available. Safe absolute raw/attachment URLs are retained; relative repository paths become permalinks only with a full commit. A reviewed `settings.report_url` can correct the per-record report destination without changing campaign membership or existing run IDs. A run uses that URL only when all member report URLs agree. Clicking **View run** opens all metrics in that group and clears other filters; the action is omitted when already at that complete scope. **Open report** opens the external evidence. Broader report navigation appears only when it would change the scope. Custom filters support combining reports and runs.

The `runs` table groups rows within a source by an explicit reported run ID, then a manifest/result file and revision, then a reported campaign label. Initial posts and explicit follow-up campaign labels remain separate. When those boundaries are unreported, a visibly labeled whole-source fallback is used. A result file or run can include a sweep across models, GPU allocations, prompts and concurrency settings. This grouping is a navigation aid, not proof of identical experimental conditions or independent replication. As of this snapshot there are 1,272 groups: 14 reported run IDs, 1,163 result-file groups, 17 campaign-label groups and 78 whole-source fallbacks.

`model_references` stores repository evidence, verification dates, availability, link type and remaining uncertainty. Rows reference it through `model_link_id`, and include `model_repository_url` and `model_link_status` directly. **Tested weights** means the source connects a repository, tag, protocol or release-pinned alias to the reported configuration; the detailed reason distinguishes this from a hash-verified historical artifact. **Upstream model reference** and **Model family reference** do not claim to identify the tested quantized files. Gated repositories are marked; a current repository or mutable tag does not prove historical bytes. A GGUF filename, a UD prefix or another commenter's download link alone does not establish a quantizer repository.

The repository review covers all 9,917 rows: 1,193 tested-weight references, 8,629 upstream references, 75 family references and 20 unresolved rows. The unresolved rows include contradictory Qwen identification, unnamed 8B tests, and unnamed suite selections/aggregates. No repository is fabricated for them. Source-specific mappings and evidence are in `research/model-links-*.json`, with the second verification pass in matching quality reports. The compiler rejects conflicting mappings and unreviewed rows; an explicit unresolved mapping is allowed when evidence is insufficient.

All categorical selectors allow multiple choices: **OR within a filter, AND across filters**. Facet choices and counts honor every other active filter, including search, card, source, ranges and dates. A facet ignores its own choices so multiple selections remain possible. Unselected zero-match options are hidden; selected zero-match values remain visible so they can be removed. Clearing constraints restores the available choices. Repeated URL parameters preserve punctuation in model names and round-trip multiple choices; legacy single-selection links still work. Individual chips remove one choice, and browser back/forward restores prior selections. Filtered CSV/JSON exports include repository evidence and absolute source-run navigation links; full database exports remain portable.

## Evaluation scores and result views

The single explorer shows all measurements by default, including evaluation scores, subtests and token speeds. Initial ordering interleaves evaluations and other metrics so both remain discoverable; it is not a performance ranking. Benchmark family, Version, Test set and Harness are separate categorical filters. The other desktop rows cover hardware/search, model identity and runtime. Custom filters expose less-used fields, including metric, unit and displayed-result ranges. Legacy result-type and result-level URL restrictions are retired. Source-run navigation retains every metric in the run.

`result_kind` separates answer-quality/agent evaluations and diagnostics from performance measurements. Procyon performance points remain in the performance category. `benchmark_*` fields describe the reported benchmark, test set, harness/agent, scoring method, level, sample count or maximum, direction and comparison protocol. Unknown details stay null. Original numeric values, units and exact models are preserved. Fractions and explicit task counts can be shown as percentages; rounded fractions never fabricate integer pass counts. Rubric points retain their original maximum. Perplexity and loss identify lower as better. Result ranges operate on the displayed numeric scale (including percentages), while exports preserve original values. Named benchmark suites are grouped independently of metric type; generic families such as Token speed or Latency are used for reports without a named suite. Benchmark versions remain unknown when only an engine build, repository commit or receipt schema is available.

Evaluation charts require matching benchmarks, subsets, harnesses, scoring methods and scales; model/quantization comparisons still require checking settings and evidence. Numeric sorting groups different protocols rather than treating them as a single leaderboard. Overall aggregates and their task/category/depth subtests are not independent benchmark replications.

The new pi-bench campaigns keep override-capable LLM-judge approval separate from recorded FAIL_TO_PASS exit-zero counts. Test collection errors stay in the denominator and those test aggregates are marked partial. These are community subset results, not official full SWE-bench resolved rates. Terminal-Bench-Local Core19 is distinct from the complete Terminal-Bench suite; its two persistent timeout failures remain in the 19-task denominator and the reported retry policy is retained. Custom MMLU/HumanEval/LAB-Bench subsets and synthetic tool-use probes retain their actual sample definitions and scoring rules.

Contaminated SWE-bench runs, rollout-completion counts without solved-task scores, budget-invalid cells and DeepSWE claims without eligible-GPU attribution are recorded as gaps/exclusions rather than imported as clean evaluations. Full source reviews are in `research/evaluation-*-quality.md`.

## GPU topology and measurement scope

Participating GPU chip count, physical board count, card identity/variant and split type are independent fields. A W6800X Duo board has two GPU chips; a source that only reports enumerated devices does not establish physical board count. Tensor parallel, pipeline parallel, data parallel, combined TP/PP, llama.cpp layer split, row split, and independent replicas are distinct. Unknown or conflicting participation is flagged. TP=1 on one established device may be represented as single GPU, with exact command/settings preserved.

Per-request decode is not aggregate serving throughput. Prefill and decode are not combined. Wall-time output throughput, aggregate decode throughput, combined input-plus-output throughput, and per-request distributions retain their metric identities and source definitions. No per-GPU division or aggregate-to-user conversion is inferred. Mean, median, percentiles, ranges, cold starts and warm runs stay separate. Relative performance and performance per dollar are not absolute tok/s.

## Evidence strength

- High: original raw data or detailed original methods with traceable measured figures.
- Medium: original measured reports with useful details but material omissions.
- Low: incomplete, ambiguous, indirectly accessible or weakly supported reports.

These are documentation ratings, not probabilities of truth or hardware endorsements. Vendor reports and sponsorship are disclosed independently. An open repository can be well documented and still contain a correctness problem. Source-specific flags should always accompany interpretation.

## Quality and history

Compilation validates required fields, unique IDs, source foreign keys, finite numeric values, integer counts, ISO date format, topology consistency, common metric units, exact duplicates and conflicting values at identical detailed locators. SQLite integrity and foreign keys are checked. Automatic completeness measures metadata presence, not correctness.

A separate review revisits source tables, footnotes, follow-up comments and correction notes. Rechecked examples and unresolved conflicts are saved under the current `research/llm-*-quality.md` and `research/expansion-*-quality.md` reports and included in `audit.json`. Old campaigns are retained. An explicitly withdrawn/corrected result may be marked superseded; a different engine version or newer model does not itself invalidate an older measurement. Immutable repository commits and raw file identifiers are retained when available. The record notes describe mismatched model labels, quantization changes, unproven GPU assignment, corrected tokenizers, and invalidated optimizations.

The UI never infers a speedup from heterogeneous results. Record details preserve conditions and missing fields; legacy numeric sorts remain grouped by protocol or metric units. Charts show individual measurements, not pooled averages across unrelated tests.

## Reproduce the compilation

From the `site` directory:

```text
python scripts/build_database.py
node --test tests/*.test.mjs
python -m unittest discover -s tests -p "test_*.py"
npm run dev
```

The build uses only the Python standard library and no network. Add reviewed input files following `research/SCHEMA.md`; inspect and resolve build errors; then rebuild. The build writes `database/database.json`, `database/benchmarks.sqlite`, the compressed JSON/SQLite downloads, `public/data/index.json` with small verified data segments, `audit.json`, and this methodology. SHA-256 fingerprints of research inputs and review notes are embedded in the outputs. The SQLite `payload_json` columns retain fields beyond the indexed columns.

Schema 3 adds `runs` and `model_references` tables, source-run and repository foreign keys on benchmarks, and indexes for source-run/repository filtering. See `research/LINKS_SCHEMA.md` for mapping selectors and precedence. Schema 3.1 adds unified benchmark-family/version metadata and retains explicit raw-artifact and reviewed report URLs.

The benchmark explorer uses plain HTML/CSS/JavaScript served by a Sites Worker. It presents a fixed, versioned research snapshot, not a live benchmark feed. Anonymous visitor submissions are stored separately in private D1 storage and parsed into provisional suggestions for owner review. On-demand source fetching is limited to submitted URLs and supported public sources; there is no scheduled crawling. Hourly abuse-limit hashes expire after 24 hours. See SUBMISSIONS.md for the intake and curation workflow.
