Bench Log

Run the same fixed set of coding and reasoning tasks on every new model, score it, and see how it stacks up against the last one.

0 models logged
How to use this
  1. Open Task library below, tap a task to expand it, and tap Copy prompt. Paste that into the model you want to test.
  2. Tap + Log a new model above and enter the model's name.
  3. Paste the model's answer into the box under that same task, then tap Judge (or Auto-judge all pasted responses once you've pasted a few) for an automatic Fail / Partial / Pass with a reason — or just tap Fail / Partial / Pass yourself.
  4. Repeat for as many of the 12 tasks as you want. Skipping some is fine — only the ones you score count toward the average.
  5. Tap Save model. It shows up in the Leaderboard, ranked by overall score with coding and reasoning split out. Log the same name again later to add a new run.
  6. Optional: open Judge settings to point auto-judging at your own API instead of the built-in one.
Judge settings
Judge

A custom endpoint only works if that provider allows direct browser requests and this sandbox can reach that domain — use "Test connection" before relying on it. Your key is saved in this app's private storage so you don't have to retype it each time; treat it like anywhere else you'd paste a key.

Model name

Paste each response under its task, then tap Judge — it's graded automatically against that task's rubric. You can always override the score by tapping Fail / Partial / Pass yourself.

How to score

Pass — gets the core answer right and the reasoning holds up. Minor style nitpicks don't count against it.
Partial — right direction but misses something material: an edge case, a wrong number, a claim that doesn't check out, or a correct answer reached by hand-wavy or lucky reasoning.
Fail — wrong answer, a right-sounding answer built on flawed reasoning, or it dodges the task (over-asks instead of attempting it, refuses without cause).

Unsure between Partial and Pass? Default to Partial — it keeps the leaderboard meaningful instead of everything clustering near 100%. Each task below names the correct answer or the specific thing to check for, so you don't have to work it out yourself.

Leaderboard

Loading…

Task library

Coding
Reasoning