Frontier monitor

Benchmarks matter. They just aren’t the whole buying decision.

Compare frontier models on intelligence and professional work, then keep agent-product benchmarks separate from the raw model underneath.

Loading current data…
Benchmark

Broad model capability signal. It is not a business-product ranking.

Model / variantScoreCost / taskSpeedContextSource

Coding benchmark = harness + tools + model.

These results are intentionally separate from raw-model scores. They measure the coding-agent product configuration that actually does the work.

Agent configurationAgent IndexDeepSWETerminalSWE-AtlasCost / taskTime
Do not collapse these

A high intelligence score does not prove a product has the connectors, computer-use surface or admin controls your business needs.

Variant matters

Reasoning effort and harness configuration materially change benchmark results. PriceBolt displays the measured variant instead of hiding it.

New releases

New products/models can appear as benchmark pending until an independent measurement exists.