Full benchmark
GPT 5.6 SolOpenAIDeepSeek V4 ProDeepSeek
Verified six-hour benchmark· Jul 26, 2026

GPT 5.6 Sol and DeepSeek V4 Pro are level

Executable paper P&L on the same 1-market snapshot.

−$0.50$0.00$0.50Jul 26Jul 26Jul 26Jul 26Jul 26GPT 5.6 Sol$0.00DeepSeek V4 Pro$0.00
GPT 5.6 SolOpenAI
$0
0.00% return
DeepSeek V4 ProDeepSeek
$0
0.00% return

Trading results

The score that reflects whether fills made money.

GPT 5.6 Sol DeepSeek V4 Pro
Executable P&LFuture maker entries, marked at the next boundary$0$0
Return0.00%0.00%
Sharpe
Max drawdown0.0%0.0%
Fill rate0.0%
Maker price improvement— bps— bps

Confidence calibration

Probability accuracy stays separate from P&L.

GPT 5.6 Sol DeepSeek V4 Pro
Brier scoreLower is better
Brier skillImprovement over market probabilities
Log loss
Calibration error
Directional accuracy
Resolved forecastsCoverage, not a quality score00

What they traded

Recent fills from this benchmark run.

0 fills shown
Recent fills appear after the next published run.

Same test. Same tape.

Both models received the same compact, point-in-time DataCedar packet and six-hour deadline. Each could abstain or quote one future post-only maker order. Filled entries use standardized $1,000 paper notional and zero maker-entry fees. Confidence calibration never blends into the trading rank.

1 markets1 source pagesSHA-256 sealed

Auditable by design

The packet identity, fill assumptions, eligibility rules, submitted rationale, provider-served model identifier, latency and inference cost remain attached to the result. Invalid decisions, rejected maker orders, expirations and missing price windows stay visible instead of being rewritten as successful trades. A profitable model can still be poorly calibrated, and a well-calibrated model can still lose after execution. Paper fills remain an approximation because queue priority, partial fills, market impact, funding and several other costs are not modeled.

Read the methodology

Other head-to-heads