OpenAI's Latest In House Chip verus Rubin NVL72New

Compare Jalapeño (Teacup) with Vera Rubin (July) NVL72 on DeepSeek R1 at 8K / 1K.

AgentX / live results

Compare Realistic Agentic Inference Perf

Long Context Multi Turn Inference Performance. Compare Across MI355X, GB300 NVL72, GB200 NVL72, B200, H200, H100, RTX Pro, etc.

Explore InferenceX

Start with a concise cost overview across active models and key platforms, or open the full dashboard for every model, chip, framework, and metric. AgentX is our long-context, multi-turn coding scenario.

Compare NVIDIA GB300 NVL72, GB200 NVL72, B300, B200, H200, H100, AMD MI355X, MI325X, MI300X and soon VR200 NVL72, AMD MI455X UALoE72, TPUv7 Ironwood, etc across DeepSeekv4 Pro, Qwen, Kimi, GLM, MiniMax, gpt-oss, Llama and other models.

Or jump straight to the benchmarks for one model:

Every Result Is Transparently done through Public GitHub Actions Automation

Every data point on the dashboard is produced by a public GitHub Actions workflow run. The recipe lives in the repo, the run executes on the actual target hardware, and the full logs and artifacts are publicly viewable. Click any point on a chart to jump straight to the run that produced it. All reproducible, auditable, and open source.

1,000+ new benchmark datapoints added per week on average. Browse every new model, chip, framework, and configuration as it lands.

Public Actions runs
Every benchmark executes on GitHub Actions with full logs visible while the run is in progress.
Open recipes
Every model, framework, precision, and parallelism setting is committed to the public repo as a shell script.
Weekly DB snapshots
The full benchmark database is published as a public GitHub Release every week so the historical dataset stays auditable.