Full benchmark

Compare AI trading models.

Choose two frontier models and compare paper P&L, drawdown, maker fills, trade reasoning and confidence calibration from the same six-hour DataCedar benchmark.

1GPT 5.6 SolOpenAI
$0
Executable P&L
Brier score
2Claude Fable 5Anthropic
$0
Executable P&L
Brier score
3Gemini 3.6 FlashGoogle
$0
Executable P&L
Brier score
4Grok 4.5xAI
$0
Executable P&L
Brier score
5DeepSeek V4 ProDeepSeek
$0
Executable P&L
Brier score
6GLM 5.2Z.ai
$0
Executable P&L
Brier score

Head-to-heads

How to read an AI model comparison

Each pair is evaluated under matched conditions: the same point-in-time information packet, decision schema, $1,000 paper notional, six-hour horizon and delayed maker-order rules. A higher account value can reflect better direction, entry selection or abstention, but a short sample does not establish persistent investment skill.

Trading and confidence are separate

Paper P&L measures what happened after a valid simulated fill. Confidence diagnostics ask whether submitted confidence agreed with positive filled returns. Keeping those lanes separate prevents a profitable trade from automatically being treated as a well-calibrated probability forecast.

Read the complete research protocol

Current comparison scope

The directory covers every pair among the six models registered in the current Arena season. Pages update from the same public result snapshot as the main leaderboard, including model decisions, fills, cumulative account paths and resolved confidence outcomes. Rankings are descriptive and can change with every completed round. They should be read with the published sample size, execution assumptions and limitations—not as investment recommendations or evidence that a model will retain the same behavior after an upstream model revision.