How Krater tests AI models

Every comparison, ranking and "best model for X" article on Krater comes from the same procedure, run inside Krater's Model Arena on the models we actually offer. This page explains what we run, how we score it and how often we redo it, so you can judge our conclusions and reproduce them.

Every comparison, ranking and "best model for X" article on Krater comes from the same procedure, run inside Krater's Model Arena on the models we actually offer.

This page explains what we run, how we score it and how often we redo it, so you can judge our conclusions and reproduce them.

What we run

  • A fixed prompt set per task type (writing, coding, research and fact retrieval, data analysis, image and video prompts). Each set is 10 prompts and does not change between models within a test round.
  • The same prompt, same attachments and same settings (temperature default, no system prompt) go to every model side by side in Model Arena, so differences come from the model, not the setup.
  • We record the model name and version, the date, the response time and the credits consumed, because models change behind the same name.

How we score

  • Fact check: every factual claim in an answer is checked against a primary source. We count hallucinations (claims that are wrong or invented) per answer.
  • Task completion: did the output do what the prompt asked, without follow-up prompting? Scored pass, partial, fail.
  • Quality: two reviewers (Malte Prüser and Vitus Prüser) read the outputs blind, without seeing which model wrote them, and rank them. Disagreements are resolved by a third run.
  • Cost and speed are reported as measured, not from vendor pages.

What we do not do

  • We do not publish vendor benchmark numbers as our own results.
  • We do not test models we do not offer in Krater.
  • We do not adjust prompts per model to make one look better.

How often

  • Flagship comparisons are rerun monthly and when a provider ships a new version. Each article shows "Tested: Month Year" and the model versions used.

Where the results appear

  • Comparison articles, model pages and the best-of lists on the blog link back here.