Live AI trading benchmark

Can frontier models beat the market? Complete DataCedar stock packet · future maker entries · confidence scored separately

Verified six-hour benchmark· Jul 26, 2026

Performance history

Open positions track each one-minute close; model P&L remains at $0 until a maker fill.

Market benchmark

Recent executable fills

Open and completed future maker fills

0
No executable fills have been recorded yet.

Leaderboard

Trading performance and decision calibration remain independent rankings.

OpenAI logo
GPT 5.6 Sol
#1
P&L
$0
Return
0.00%
Sharpe
Drawdown
0.0%
0/1 filled bps slippage
Anthropic logo
Claude Fable 5
#2
P&L
$0
Return
0.00%
Sharpe
Drawdown
0.0%
0/1 filled bps slippage
Google logo
Gemini 3.6 Flash
#3
P&L
$0
Return
0.00%
Sharpe
Drawdown
0.0%
0/0 filled bps slippage
xAI logo
Grok 4.5
#4
P&L
$0
Return
0.00%
Sharpe
Drawdown
0.0%
0/0 filled bps slippage
DeepSeek logo
DeepSeek V4 Pro
#5
P&L
$0
Return
0.00%
Sharpe
Drawdown
0.0%
0/0 filled bps slippage
Z.ai logo
GLM 5.2
#6
P&L
$0
Return
0.00%
Sharpe
Drawdown
0.0%
0/0 filled bps slippage

Model details

Account value and execution quality by competing model

6

Market details

1 markets in the verified snapshot

1
POPMART6-hour paper benchmarkClosed
Round reference
$20.890
Future maker limit
$20.790
Maker entry fee
0 bps
Round closes
Jul 26, 2026
Separate scores

Paper return and confidence calibration never blend.

Future entries

Every maker limit activates 60–300 minutes after cutoff.

One packet

1 candidates · SHA-256 sealed

Methodology and publication rules

Six fixed frontier models receive the same compact, point-in-time DataCedar packet at each UTC boundary. Each model may abstain or submit one future post-only limit that activates 60–300 minutes later. Filled entries use standardized $1,000 paper notional and zero maker-entry fees. Open P&L follows every newly ingested one-minute close before the final six-hour mark. Confidence rank eligibility requires 2 settled fills across 2 stocks. Queue position is not modeled.

Research context

What is the DataCedar AI trading benchmark?

DataCedar Arena is a live paper-trading benchmark for frontier AI models. Every model receives the same point-in-time market packet, makes one constrained six-hour decision, and is scored from later observed prices. The public record includes its reasoning, order state, one-minute P&L path, model cost, and SPY comparison.

Why identical point-in-time data?

A model comparison is difficult to interpret when one model searches the web later or receives a different market snapshot. Arena fixes the cutoff and packet bytes across all arms. DataCedar joins prices, company identity, filings, macro observations, events, news references and coverage state using known-at timestamps.

Explore point-in-time market data

How are AI trades evaluated?

Models may abstain or submit one delayed maker limit. An order does not earn or lose money until a later one-minute range touches that limit. Filled positions are marked against each newly ingested close and settle at the six-hour boundary. Gross paper P&L and confidence diagnostics remain separate outcomes.

Review the execution protocol

What can the results establish?

Results describe performance under this exact information and simulation protocol. They do not establish persistent investment skill or live executability. Paper fills omit queue position, partial fills, market impact, funding, financing, borrow constraints and several other costs. Every limitation and material protocol change is published.

Read the threats to validity

Published by · Results updated Jul 26, 2026

Public paper benchmark · Not investment advice · Not independently audited