Claude Code shows speed lead as agent framework costs diverge

Composio tested DeepSeek V4 Flash across Claude Code, Codex, OpenCode, and Oh My Pi on 30 tasks. Claude Code was fastest, OpenCode was cheapest, and Oh My Pi had the highest success rate, showing that the agent framework can change both cost and performance.

WTF Index NEUTRAL
◄ Terminator 0 Idiocracy 0 ►

This is a routine performance and cost benchmark of AI agent frameworks without clear evidence of danger or social degradation.

Claude Code shows speed lead as agent framework costs diverge

Choosing an AI model is only part of the cost and performance equation. Composio’s test of DeepSeek V4 Flash across four agent frameworks shows that the surrounding software layer can strongly affect how fast tasks finish, how often they succeed, and how much each successful result costs.

The comparison covered Claude Code, Codex, OpenCode, and Oh My Pi across 30 tasks using real-world tools like Gmail, GitHub, Slack, and Notion. The result was not a simple ranking. Each framework had a different strength, and none led across every category.

What Composio Tested

Composio ran DeepSeek V4 Flash through four agent frameworks: Claude Code, Codex, OpenCode, and Oh My Pi. The tasks were tied to practical tool use rather than abstract prompts, with Gmail, GitHub, Slack, and Notion included among the tools.

That matters because agent frameworks do more than pass text to a model. They shape how the model plans, calls tools, produces outputs, and completes work. In this test, those differences were large enough to change the outcome even when the underlying model stayed the same.

The clearest takeaway is that the framework can be a major part of the total system behavior. The model was DeepSeek V4 Flash throughout, but cost, speed, and success rate still shifted depending on whether the task ran through Claude Code, Codex, OpenCode, or Oh My Pi.

The Winners Changed By Category

Oh My Pi posted the highest success rate, completing 17/30 tasks successfully. That made it the strongest performer on task success in this comparison, but it also came with the slowest average completion time at 272 seconds per task.

Claude Code went in the opposite direction. It was the fastest framework in the test, averaging 122 seconds. At the same time, it was the most expensive option per successful task at $0.195.

OpenCode was the lowest-cost framework, reaching $0.073 per successful task. Its success rate was slightly behind the others at 14/30, making it the only framework described as trailing on overall success.

Codex was part of the four-framework comparison, though the source does not single it out as the top performer on success rate, cost, or speed. The broader finding is still important: the field was close enough on success that the biggest practical differences appeared in cost and completion time.

Cost And Speed Varied Sharply

The test found nearly a 3x price difference depending on the framework. For anyone comparing agent systems, that is a meaningful spread because it appears at the level of successful tasks, not just raw model usage.

Speed also changed substantially. The source reports a 2.2x speed difference across frameworks. Claude Code finished fastest at 122 seconds, while Oh My Pi took 272 seconds per task.

One notable detail is that Claude Code was still the most expensive even though it used the fewest tool calls and generated the least output tokens. That shows why token count and tool-call count do not always tell the full story of cost in an agent framework.

For teams evaluating AI tools, the practical lesson is straightforward: measuring only model output is not enough. The wrapper around the model can influence how much work gets done, how quickly it gets done, and what each successful result costs.

Success Was Close, But Not Identical

The overall success rates were described as close, with OpenCode slightly behind at 14/30 and Oh My Pi ahead at 17/30. That narrow spread makes the cost and speed differences stand out more clearly.

Still, the framework choice did affect individual outcomes. Seven tasks passed or failed based solely on which framework ran them. In other words, the same model could succeed or fail depending on the agent environment around it.

That point is important for interpreting the results. A framework that looks strong on average may still miss specific tasks. A cheaper framework may be attractive if cost matters most, while a faster one may be more useful when latency is the main concern.

What The Results Mean

Composio’s findings point to a more nuanced way to compare agent systems. Claude Code, Codex, OpenCode, and Oh My Pi should not be judged only by whether they can connect to tools or run a model. The operational tradeoffs are part of the product.

Based on the source, the comparison breaks down into three clear signals:

  • Best success rate: Oh My Pi, with 17/30 successful tasks.
  • Fastest completion: Claude Code, at 122 seconds.
  • Lowest cost per successful task: OpenCode, at $0.073.

No single framework dominated every dimension. That makes the right choice depend on what matters most in a given workflow: reliability, speed, or cost.

The larger message is that agent framework selection can materially change results even when the AI model remains constant. In Composio’s test, the model was DeepSeek V4 Flash, but the framework determined whether some tasks passed, how quickly they completed, and how expensive success became.