Model Pit

A benchmark for your own prompt.

Run any prompt on several models and see how long each took, what its answer cost, and how many claims it backed or disputed.

Claims checked by Jev, TypeSafe's decision model.

Model presetsDefault

Benchmark for this prompt

Backs and Disputes count the claims below that each answer supports or contradicts, as checked by Jev. Stood alone counts the claims where one answer was the only one on its side against two or more.

ModelTimeOutput tokensTokens/sCostBacksDisputesStood alone
OpenAIGPT-6 Sol9.2s53858$0.0054300
ClaudeClaude Sonnet 59.3s71877$0.0073500
GeminiGemini 3.8 Flash9.6s1,196125$0.0045500
DeepSeekDeepSeek V4.1 Flash17.7s1,20068$0.0004000

Is it true that about 70% of online shopping carts are abandoned? Where does the figure come from, and how was it measured?

Run

Start from an example

Each one opens the studio with the text, settings and models filled in.

How it works

  1. 1

    Pick a test

    Start from an example (made-up sources, reasoning, instructions, speed, recent knowledge, counting) or write your own prompt.

  2. 2

    Run it

    Five models answer the example prompts. The synthesis runs by itself, and Jev checks which claims each answer backs or disputes.

  3. 3

    Read the benchmark

    Time, output tokens, tokens per second, cost, claims backed and disputed, and claims where a model stood alone.

Published scores sit next to yours.

Model Pit also keeps a table of published scores on GPQA Diamond, SWE-bench Verified, AIME 2025, MMLU-Pro, Humanity's Last Exam and LiveCodeBench for 45 models, each score with its source and date.

What it costs

A run on the 5 models above costs 11 credits, about $0.44 on the $20 bundle. The synthesis runs with every benchmark and is 3 credits. Credits never expire, and your first 20 are free.

You pay for questions you ask. That is the whole model.

No plan, no seat, no monthly minimum, nothing to cancel. You buy credits once and they never expire, every model shows its cost before you run it, and we refund any run that does not finish.

Free

$0

Try Keimodel with no commitment.

  • 20 credits on sign-up
  • The same models as the paid tiers
  • No credit card required
Get started

Starter

$5one-time

$0.05 per credit

Top up when your free credits run low.

  • 100 credits
  • Access to all models
  • Credits never expire
Buy credits
Most popular

Pro

$20one-time

$0.04 per credit

Best value for regular users.

  • 500 credits
  • Access to all models
  • Credits never expire
Buy credits

Max

$35one-time

$0.035 per credit

For power users running many comparisons.

  • 1,000 credits
  • Access to all models
  • Credits never expire
Buy credits

Most model responses cost 1 to 5 credits · A very long prompt or answer costs 1 credit per $0.035 of model cost · The synthesis is priced separately

Priced before you run

Every model carries its credit cost in the picker and its real input and output price per million tokens on the panel, and the total for the lineup you have built sits under the Run button. We meter nothing after the fact.

Or bring your own key

Add an OpenRouter key in Settings and every model response runs on your own account and costs you no credits at all. Only the synthesis, which reads every answer and lists the claims, stays on ours.

Questions

What does “stood alone” mean?

A claim where a model was the only one on its side while two or more others took the other side. Standing alone is not the same as being wrong: when every model shares the same mistake, nobody stands alone.

Why does the synthesis run by itself here?

The backed, disputed and stood-alone columns come from the synthesis and Jev's check, so Model Pit runs it after every prompt. It costs 3 credits.

Where do the published scores come from?

From the model developers' own announcements and from independent leaderboards. Each score lists its source and the date it was published.

What goes into Model stats?

For each model: its median time and cost per answer, and how many claims it backed or stood alone on across Keimodel runs. The stats keep counts only, not prompts or answers.

Benchmark your own prompt.

Pick the models, run the prompt, and read the table.

Start free20 free credits to start, no card required