Run any prompt on several models and see how long each took, what its answer cost, and how many claims it backed or disputed.
Claims checked by Jev, TypeSafe's decision model.
Backs and Disputes count the claims below that each answer supports or contradicts, as checked by Jev. Stood alone counts the claims where one answer was the only one on its side against two or more.
| Model | Time | Output tokens | Tokens/s | Cost | Backs | Disputes | Stood alone |
|---|---|---|---|---|---|---|---|
| GPT-6 Sol | 9.2s | 538 | 58 | $0.0054 | 3 | 0 | 0 |
| Claude Sonnet 5 | 9.3s | 718 | 77 | $0.0073 | 5 | 0 | 0 |
| Gemini 3.8 Flash | 9.6s | 1,196 | 125 | $0.0045 | 5 | 0 | 0 |
| DeepSeek V4.1 Flash | 17.7s | 1,200 | 68 | $0.0004 | 0 | 0 | 0 |
Is it true that about 70% of online shopping carts are abandoned? Where does the figure come from, and how was it measured?
RunEach one opens the studio with the text, settings and models filled in.
Start from an example (made-up sources, reasoning, instructions, speed, recent knowledge, counting) or write your own prompt.
Five models answer the example prompts. The synthesis runs by itself, and Jev checks which claims each answer backs or disputes.
Time, output tokens, tokens per second, cost, claims backed and disputed, and claims where a model stood alone.
Model Pit also keeps a table of published scores on GPQA Diamond, SWE-bench Verified, AIME 2025, MMLU-Pro, Humanity's Last Exam and LiveCodeBench for 45 models, each score with its source and date.
A run on the 5 models above costs 11 credits, about $0.44 on the $20 bundle. The synthesis runs with every benchmark and is 3 credits. Credits never expire, and your first 20 are free.
No plan, no seat, no monthly minimum, nothing to cancel. You buy credits once and they never expire, every model shows its cost before you run it, and we refund any run that does not finish.
Free
Try Keimodel with no commitment.
Starter
$0.05 per credit
Top up when your free credits run low.
Pro
$0.04 per credit
Best value for regular users.
Max
$0.035 per credit
For power users running many comparisons.
Most model responses cost 1 to 5 credits · A very long prompt or answer costs 1 credit per $0.035 of model cost · The synthesis is priced separately
Every model carries its credit cost in the picker and its real input and output price per million tokens on the panel, and the total for the lineup you have built sits under the Run button. We meter nothing after the fact.
Add an OpenRouter key in Settings and every model response runs on your own account and costs you no credits at all. Only the synthesis, which reads every answer and lists the claims, stays on ours.
A claim where a model was the only one on its side while two or more others took the other side. Standing alone is not the same as being wrong: when every model shares the same mistake, nobody stands alone.
The backed, disputed and stood-alone columns come from the synthesis and Jev's check, so Model Pit runs it after every prompt. It costs 3 credits.
From the model developers' own announcements and from independent leaderboards. Each score lists its source and the date it was published.
For each model: its median time and cost per answer, and how many claims it backed or stood alone on across Keimodel runs. The stats keep counts only, not prompts or answers.
Pick the models, run the prompt, and read the table.