OpenAI Commercial APIStatus: active

GPT-4o

OpenAI's omni-modal flagship model natively processing text, audio, images, and vision in real time.

Released: May 13, 2024
Last Verified: Aug 10, 2026
Context Window

128k

Tokens

Max Output

16.384k

Output limit

Architecture

Multimodal Transformer

Model family

Parameters

Not Disclosed

Total / Active

Input Token Cost

$2.50

Per 1M tokens

Output Token Cost

$10.00

Per 1M tokens

Model Overview

GPT-4o (Omni) is OpenAI's versatile multimodal model featuring native omni architecture, ultra-fast response latency, 128k context window, structured JSON schema outputs with 100% adherence, and fine-tuning support.

Developer Implementation Notes

Supports Realtime Audio API with low latency WebRTC connection. Strict structured outputs ensure deterministic JSON schema matching.

Key Strengths

  • Native omni-modal processing (audio, vision, text)
  • Strict structured outputs (100% JSON schema validation)
  • Fine-tuning available on text and vision datasets
  • 50% prompt caching discount ($1.25/MTok)

Limitations & Boundaries

  • Lacks deep chain-of-thought reasoning tokens (handled by o1/o3-mini)
  • 128k context window is smaller than 2M on Gemini or 200k on Claude

Capabilities & Modalities

Modalities:text, image, audio
Reasoning Chains:No
Tool / Function Calling:Supported
Structured Outputs:Supported
Fine-Tuning:Supported
Local Deployment:Cloud Only

Best Production Use Cases

  • Real-time voice and audio agents
  • Enterprise structured JSON extraction pipelines
  • Multimodal document and visual understanding

Verified Benchmark Results

Standardized evaluations with methodology notes and authoritative citation links.

View all industry benchmarks →
Coding Official
SWE-bench Verified

38.8%

Pass@1 standard harness

Coding Official
LiveCodeBench

45.3%

Pass@1, 0-shot code generation

Reasoning & Math Official
GPQA Diamond

53.6%

Zero-shot Chain-of-Thought

General Knowledge & Frontier Official
MMLU-Pro

72.5%

5-shot standard benchmark run

Tool & Agentic Official
BFCL (Berkeley Function Calling)

88.5%

AST accuracy on single & multi-turn tool calling

Multimodal Official
MMMU

69.1%

Multimodal visual benchmark

Pricing & Inference Cost Calculator

Token Cost Estimator – GPT-4o

Calculate projected inference spend with prompt caching

Quick Workload Presets

Input Tokens / Req2,000
100100k200k
Output Tokens / Req500
5016k32k
Number of Requests1,000
125k50k
Cache Hit Rate0%
0% (No cache)50%90% (Max)
Total Input Spend

$5.0000

2.00M input tokens

Total Output Spend

$5.0000

0.50M output tokens

Estimated Total Cost

$10.00

Avg: $0.01000 / req

Rates: $2.50 in / $10.00 out per million tokens (OpenAI API)Last verified: Aug 10, 2026

Compare with Similar Models

Anthropic

Claude 3.7 Sonnet

Anthropic's first hybrid reasoning frontier model with dynamic thinking budget control and state-of-the-art coding capabilities.

Compare GPT-4o vs Claude 3.7 Sonnet
Google DeepMind

Gemini 2.0 Pro (Experimental)

Google's flagship intelligence model engineered for complex coding, mathematical proofs, and 2M token context.

Compare GPT-4o vs Gemini 2.0 Pro (Experimental)
OpenAI

OpenAI o3-mini

Next-generation cost-efficient reasoning model optimized for STEM, competitive math, and coding.

Compare GPT-4o vs OpenAI o3-mini

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.