Live AI trading benchmark

Can frontier models beat the market? Bounded active-market sample · future maker entries · confidence scored separately

Awaiting first eligible round

Performance history

Open positions track each one-minute close; model P&L remains at $0 until a maker fill.

Waiting for the first real fill
The future maker orders are still pending. Live P&L begins only after an order is touched, then updates with each newly ingested one-minute close.

Recent executable fills

Open and completed future maker fills

0
No executable fills have been recorded yet.

Leaderboard

Trading performance and decision calibration remain independent rankings.

No certified execution ranking yet.

Model details

Account value and execution quality by competing model

0
Model details appear after the first scored cycle.

Market details

— markets in the verified snapshot

0
Market details appear with the next verified tape.
Separate scores

Paper return and confidence calibration never blend.

Future entries

Every maker limit activates 60–300 minutes after cutoff.

One packet

The same point-in-time packet for every model.

Methodology and publication rules

Six fixed frontier models receive the same compact, point-in-time DataCedar packet at each UTC boundary. Each model may abstain or submit one future post-only limit that activates 60–300 minutes later. Filled entries use standardized $1,000 paper notional and zero maker-entry fees. Open P&L follows every newly ingested one-minute close before the final six-hour mark. Confidence rank eligibility requires 0 settled fills across 2 stocks. Queue position is not modeled.

Research context

What is the DataCedar AI trading benchmark?

DataCedar Arena is a live paper-trading benchmark for frontier AI models. Every model receives the same point-in-time market packet, makes one constrained six-hour decision, and is scored from later observed prices. The public record includes its reasoning, order state, one-minute P&L path, model cost, and SPY comparison.

Why identical point-in-time data?

A model comparison is difficult to interpret when one model searches the web later or receives a different market snapshot. Arena fixes the cutoff and packet bytes across all arms. DataCedar joins prices, company identity, filings, macro observations, events, news references and coverage state using known-at timestamps.

Explore point-in-time market data

How are AI trades evaluated?

Models may abstain or submit one delayed maker limit. An order does not earn or lose money until a later one-minute range touches that limit. Filled positions are marked against each newly ingested close and settle at the six-hour boundary. Gross paper P&L and confidence diagnostics remain separate outcomes.

Review the execution protocol

What can the results establish?

Results describe performance under this exact information and simulation protocol. They do not establish persistent investment skill or live executability. Paper fills omit queue position, partial fills, market impact, funding, financing, borrow constraints and several other costs. Every limitation and material protocol change is published.

Read the threats to validity

Published by

Public paper benchmark · Not investment advice · Not independently audited