← All posts

Two models on one problem

7 min readVigneshMarkdown

Models are trained differently, so they miss different things. A second model is a second pair of eyes that did not grow up with the first. Pasting one model's answer into the other is already a workflow, run by hand. The work is in arranging the two so they help each other, not repeat each other.

Loopgate ships fourteen patterns. Here are three that put two models on one problem. I ran each on this machine for this post.

One honest note first. Claude Code was not signed in here on the day, so every step that a pattern puts on Claude ran on pi with GPT-6 Sol instead. The other side ran on Codex with GPT-6 Astra. Two harnesses, two models, one lab. They still disagreed, as you will see.

Pair programming

A driver writes, a navigator reads. The navigator says continue with what is left, or ready. Up to six rounds.

I ran it on a small sample package: a slugify function and six failing tests. The driver was on pi, the navigator on Codex.

The pair-programming run: Round 1 of 6, driver → navigator, ready. The driver's reply: added slugify, npm test passed all 6 tests, the new file is untracked. The navigator's reply on Codex GPT-6 Astra: reviewed against the README and tests, independently ran npm test, and "the handoff's untracked note is outdated"
One round, fifty seconds, fifty cents. The navigator ran the tests itself.

One round, 50 seconds, $0.50. The driver cost two cents. The navigator cost 48.

The task was too small for a real disagreement. But look at what the navigator did. It ran the tests itself instead of trusting the report. It also caught that the driver's "the new file is untracked" was no longer true by the time it looked. Small. Exactly the kind of small that ends up in a pull request description and misleads a reviewer.

Parallel plans, then build

Two planners plan the same task at the same time. Neither sees the other. A lead reads both and writes one plan: what both agreed, where they disagreed, which it chose and why. You lock it. Then build, review and fix in a bounded loop.

The run's card: task, then plan_claude and plan_codex side by side, synthesize, lock, build, review, fix, decide, ship. "The $10 cost ceiling was reached. $10.23 spent."
The track as the run drew it. The two plans ran side by side.

The task was a real one on this repository: add --status <state> to loopgate runs, refuse an unknown state in one sentence, and cover it with a test.

The step names come from the pattern. plan_claude ran on pi with GPT-6 Sol, plan_codex on Codex with GPT-6 Astra.

Both agreed on the shape: filter in the CLI, validate before starting the daemon, no API change. They disagreed on the edges. The first plan left the error and the empty result unworded. The second wrote both sentences out. It also noticed that the client returns fifty runs a page, so a filter has to walk the pages, and it asked for a test where two runs start at the same moment.

The lead took the second plan's contract. In its own words:

synthesize on pi GPT-6 Sol high: "Plan — loopgate runs --status <state>", Agreement and decisions: both plans place the filter in the CLI… then "Claude's plan leaves error and empty-result wording open; Codex specifies them…", saved as PLAN.md
The lead's plan says where the two disagreed and which reading it took.

Choose Codex's explicit output contract because it is predictable to people and scripts, and its tied-timestamp test protects the existing cursor semantics.

I locked it. The plans and the lead cost $2.54 together.

Then the build ran into my ceiling. I had set ten dollars for every run in this post. The builder on Codex with GPT-6 Astra spent the rest, and the run stopped on budget at $10.23, mid-build.

That is the ceiling doing its job. The card says what it spent, and Raise it is one press. I did not press it. A post about patterns does not need a ten-line flag merged. The plan is the part I wanted you to see.

Root cause by two investigators

Two investigators look at the same bug, each on its own harness, neither seeing the other. A lead writes one root cause from both, with the evidence, where they disagree and which reading it takes. You choose: fix it now, or finish with the report.

I gave it the bug from the last post: the composer drops a step's engine choice when setup words open a draft.

One investigator, on pi, found it in 44 seconds for 39 cents. The composer keeps the choice in local state. When setup words open a draft, it creates the chat without those choices, and the draft fits the engines again from scratch.

The other investigator, on Codex with Read only access, found nothing. Loopgate refused its very first command, pwd && rg --files -g '…', as more than a plain read. It tried two other ways to read files. Both were refused too, and it stopped and said so.

Two replies. look_codex: "Investigation is blocked; I cannot establish the cause from evidence. Loopgate rejected all repository-access attempts…". look_claude on pi: "Cause: the composer keeps the plan engine choice only in its local overrides state…"
One investigator blind, one with the answer. Both said exactly which they were.

That is Loopgate's rule, not Codex's. A Read only step gets plain reads and nothing else, and its list of plain reads knows rg --files but not the -g filter, and a quoted glob is not a path it can check. So the whole line failed the test. The step was not even told that: what Codex got back was Loopgate denied this tool. Use a workflow node for delegation and declared checks., which is not the reason. A read-only investigator should not go blind on a file search, and it should hear the rule it hit. Both are mine to fix. I found them because the run put them in front of me.

What I like is what the lead did with it. It did not average a blind answer with a seeing one:

The source confirms look_claude's diagnosis… look_codex had no repository access, so its omission-versus-overwrite question was unresolved rather than contradictory.

I chose Fix it now. The fixer on Codex wrote a regression test and started on the change. Eight minutes and $8.88 later, the run stopped on its ten dollar ceiling.

The lead's reply, "You chose Fix it now at choose", then the round band: Round 1 of 2, fix, stopped on budget, 8m, $8.88
Stopped on budget. The worktree keeps what the fixer wrote.

The root cause cost $1.28. The fix would have cost more than I gave it.

Or say it

You do not have to find the pattern. Say it in the composer, with Who on Auto.

Two planners, one on Claude Code and one on Codex, then a lead merges the plans, then build and review

The copilot picked parallel-plans-then-build off the shelf.

The draft: "Two planners will work in parallel on Claude Code and Codex…", parallel-plans-then-build, Loopgate's pick, the Task row waiting, and the steps with both planners on Codex
The reply says Claude Code. The card says Codex. Trust the card.

Its reply says one planner runs on Claude Code. The card shows both on Codex, because Claude Code was not signed in, and Loopgate fitted every step to an engine this machine can run. The words were wrong. The card was right. The card is what Start uses.

Then I asked for something the shelf does not have.

Two planners on two harnesses, then a lead merges the plans, then build, then a security reviewer and a test reviewer in parallel, then I approve, and email me the summary.

The copilot wrote a workflow for it, and said what it could not build.

The draft: plan-build-parallel-reviews-approve, written for this task, with plan_codex and plan_pi, synthesize, build, security_review and test_review side by side. Above it, Left out: "Emailing you the summary is not supported by the available workflow blocks"
Left out names what it did not build. Loopgate has no email block.

Left out is the line I care about. It did not quietly drop the email. It did not invent a block. It said so, and the run would show the summary instead.

One built for one project

A pattern is a starting point. second-brain-bootstrap is one of my own, for one project.

A lead plans, I lock the plan, two builders work in parallel worktrees, the lead verifies the merged result for up to four rounds, and I accept it.

The builder canvas: task, plan, approve_plan, team, then runtime and interface in parallel, lead with "up to 4 rounds", then review_result and done, with exhausted under it after 4 rounds
second-brain-bootstrap in the builder. Every route it can take is drawn, including the one after four rounds.

What these cost

Three runs and two drafts for this post. Pair programming $0.50. Parallel plans $10.23, stopped on budget. Root cause $10.16, stopped on budget. The copilot's drafts, under a cent each. Plus the first post's run at $1.02.

Two of the four runs across both posts hit my ceiling. Both times the step spending was a builder or fixer on GPT-6 Astra. Next time I give that step a cheaper model and keep the expensive one for the planners. The engine is per step, so that is one change on the card.

The patterns guide lists all fourteen, and what each one gives you.