AI Organization & Research Lab

OpenAI

Pioneering AI research and deployment company behind GPT-4o, o1, o3-mini, and ChatGPT.

Foundation Models by OpenAI

Active and preview models in the OpenAI intelligence series.

4 Models
OpenAI Commercial API

OpenAI o3-mini

Next-generation cost-efficient reasoning model optimized for STEM, competitive math, and coding.

Context Window

200k tokens

Architecture / Params

Not Disclosed

text Reasoning Tool Calling
Verified Benchmark Scores
AIME (2024/2025)87.3%
SWE-bench Verified49.3%
LiveCodeBench68.2%
Input / Output

$1.10 / $4.40 / MTok

Full Dossier
OpenAI Commercial API

OpenAI o1

Flagship deep reasoning model trained with reinforcement learning for frontier science, math, and coding.

Context Window

200k tokens

Architecture / Params

Not Disclosed

textimage Reasoning Tool Calling
Verified Benchmark Scores
SWE-bench Verified48.9%
AIME (2024/2025)83.3%
GPQA Diamond75.7%
Input / Output

$15.00 / $60.00 / MTok

Full Dossier
OpenAI Commercial API

GPT-4o mini

High-speed, ultra-affordable multimodal small model designed to replace GPT-3.5 Turbo at 60% lower cost.

Context Window

128k tokens

Architecture / Params

Not Disclosed

textimage Tool Calling
Verified Benchmark Scores
SWE-bench Verified20.2%
LiveCodeBench32.4%
GPQA Diamond40.2%
Input / Output

$0.15 / $0.60 / MTok

Full Dossier
OpenAI Commercial API

GPT-4o

OpenAI's omni-modal flagship model natively processing text, audio, images, and vision in real time.

Context Window

128k tokens

Architecture / Params

Not Disclosed

textimageaudio Tool Calling
Verified Benchmark Scores
SWE-bench Verified38.8%
LiveCodeBench45.3%
GPQA Diamond53.6%
Input / Output

$2.50 / $10.00 / MTok

Full Dossier

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.