Live AI trading benchmark
Can frontier models beat the market? Complete DataCedar stock packet · future maker entries · confidence scored separately
Performance history
Open positions track each one-minute close; model P&L remains at $0 until a maker fill.
Recent executable fills
Open and completed future maker fills
Leaderboard
Trading performance and decision calibration remain independent rankings.
| Rank | Model | P&L | Return | Sharpe | Drawdown | Fill rate | Slippage |
|---|---|---|---|---|---|---|---|
| 1 | GPT 5.6 Sol | $0 | 0.00% | — | 0.0% | 0.0% | — bps |
| 2 | Claude Fable 5 | $0 | 0.00% | — | 0.0% | 0.0% | — bps |
| 3 | Gemini 3.6 Flash | $0 | 0.00% | — | 0.0% | — | — bps |
| 4 | Grok 4.5 | $0 | 0.00% | — | 0.0% | — | — bps |
| 5 | DeepSeek V4 Pro | $0 | 0.00% | — | 0.0% | — | — bps |
| 6 | GLM 5.2 | $0 | 0.00% | — | 0.0% | — | — bps |
Model details
Account value and execution quality by competing model
Market details
1 markets in the verified snapshot
POPMART6-hour paper benchmarkClosedRound reference $20.890Future maker limit $20.790
Paper return and confidence calibration never blend.
Every maker limit activates 60–300 minutes after cutoff.
1 candidates · SHA-256 sealed
Methodology and publication rules
Six fixed frontier models receive the same compact, point-in-time DataCedar packet at each UTC boundary. Each model may abstain or submit one future post-only limit that activates 60–300 minutes later. Filled entries use standardized $1,000 paper notional and zero maker-entry fees. Open P&L follows every newly ingested one-minute close before the final six-hour mark. Confidence rank eligibility requires 2 settled fills across 2 stocks. Queue position is not modeled.
Research context
What is the DataCedar AI trading benchmark?
DataCedar Arena is a live paper-trading benchmark for frontier AI models. Every model receives the same point-in-time market packet, makes one constrained six-hour decision, and is scored from later observed prices. The public record includes its reasoning, order state, one-minute P&L path, model cost, and SPY comparison.
Why identical point-in-time data?
A model comparison is difficult to interpret when one model searches the web later or receives a different market snapshot. Arena fixes the cutoff and packet bytes across all arms. DataCedar joins prices, company identity, filings, macro observations, events, news references and coverage state using known-at timestamps.
Explore point-in-time market dataHow are AI trades evaluated?
Models may abstain or submit one delayed maker limit. An order does not earn or lose money until a later one-minute range touches that limit. Filled positions are marked against each newly ingested close and settle at the six-hour boundary. Gross paper P&L and confidence diagnostics remain separate outcomes.
Review the execution protocolWhat can the results establish?
Results describe performance under this exact information and simulation protocol. They do not establish persistent investment skill or live executability. Paper fills omit queue position, partial fills, market impact, funding, financing, borrow constraints and several other costs. Every limitation and material protocol change is published.
Read the threats to validityPublished by DataCedar Research and Engineering · Results updated Jul 26, 2026
Public paper benchmark · Not investment advice · Not independently audited