How We Compare Models in One Workspace (A Method, Not a Leaderboard)
A repeatable bake-off: same prompt, shared notes, score the task, then decide whether to add Grok or DeepSeek. Rankings rot. A method stays useful.
By the flAim team
We do not publish a “best model” leaderboard. Rankings rot the week a lab ships. A method you can rerun in a workspace does not.
This is the protocol we use internally and the one the Model arena template is built for: same prompt, shared notes, score the task.
The protocol
- 1
Write the brief in shared notes. Task, constraints, what “done” looks like. Agents read the same source of truth. Do not paste a slightly different brief into each tab.
- 2
Pin the models as agents. Typical first pass: one OpenAI, one Claude, one Gemini. Add DeepSeek or Grok when the task cares about cost, speed, or a different style — not because a blog said so.
- 3
Send once in Chat to compare. Same prompt to every agent. You are measuring disagreement, not collecting vibes from three private chats. For everyday work, @ one agent, then another in that same thread, then fan out when you want everyone on the next question.
- 4
Score the task, not the prose. Rubric of three: correctness against the brief, missing constraints, and whether you would ship it. Style is a tie-breaker.
- 5
Write the winner into the notes. Next week’s bake-off starts from that. The workspace remembers. Your sidebar does not.
What we do not do
| Anti-pattern | Why it lies | Replace with |
|---|---|---|
| Global ranking posts | The catalog moved. Your job did not match their prompts. | A rubric for this brief, this week. |
| Different prompts per model | You compared prompting skill, not models. | One prompt when you bake off. @ one then another when you are conversing. |
| Winner-take-all forever | Claude can win research and lose a tight rewrite. | Re-run when the job changes. |
| Ignoring cheaper models | Frontier is not always the task. | Add DeepSeek or a Flash-class model when cost or speed is the constraint. |
When to add Grok or DeepSeek
After the three-lab baseline, not instead of it. Add a model when you have a hypothesis: cheaper everyday chat, a different tone, or a second opinion on a claim the first three already split on. Adding six agents on day one is noise.
How to attach providers is in Documentation. You can start on managed credits with no API key.
Run a bake-off in the Model arena template
Three agents, one shared task brief, same prompt. Score the job — not the vibe.
Related: Chat vs Build · Add agents · What is a multi-agent workspace
We're building this in the open and we read everything. If you try it and something's clumsy, tell us. We'd rather hear it from you than not hear it.