Google DeepMind Commercial APIStatus: active

Gemini 2.0 Flash

Google's high-speed multimodal workhorse with 1M token context, native tool use, and real-time audio/video streaming.

Released: Dec 11, 2024
Last Verified: Aug 10, 2026
Context Window

1,000k

Tokens

Max Output

8.192k

Output limit

Architecture

Multimodal MoE

Model family

Parameters

Not Disclosed

Total / Active

Input Token Cost

$0.10

Per 1M tokens

Output Token Cost

$0.40

Per 1M tokens

Model Overview

Gemini 2.0 Flash is built from the ground up for speed, low latency, and multimodal capabilities across text, audio, image, and video inputs. Features 1M token context window, built-in Google Search grounding, code execution sandbox, and ultra-competitive pricing ($0.10/$0.40 per MTok).

Developer Implementation Notes

Multimodal Live API enables bidirectional voice and video streaming over WebSockets. Grounding with Google Search supported natively.

Key Strengths

  • Massive 1,000,000 token context window
  • Native video and audio ingestion at low cost
  • Extremely cheap pricing ($0.10 input / $0.40 output per MTok)
  • Built-in Google Search grounding and code execution

Limitations & Boundaries

  • Maximum output token limit of 8,192
  • Can hallucinate fine details in obscure edge cases without grounding

Capabilities & Modalities

Modalities:text, image, audio, video
Reasoning Chains:No
Tool / Function Calling:Supported
Structured Outputs:Supported
Fine-Tuning:Supported
Local Deployment:Cloud Only

Best Production Use Cases

  • Video and audio analysis & transcription
  • Massive document and codebase RAG
  • Real-time interactive voice agents

Verified Benchmark Results

Standardized evaluations with methodology notes and authoritative citation links.

View all industry benchmarks →
Coding Official
LiveCodeBench

48.0%

Pass@1, 0-shot code generation

Reasoning & Math Official
GPQA Diamond

62.1%

Zero-shot Chain-of-Thought prompt

General Knowledge & Frontier Official
MMLU-Pro

74.8%

5-shot Chain-of-Thought standard

Multimodal Official
MMMU

73.1%

Multimodal visual benchmark

Tool & Agentic Official
BFCL (Berkeley Function Calling)

87.4%

Function calling evaluation

Pricing & Inference Cost Calculator

Token Cost Estimator – Gemini 2.0 Flash

Calculate projected inference spend with prompt caching

Quick Workload Presets

Input Tokens / Req2,000
100100k200k
Output Tokens / Req500
5016k32k
Number of Requests1,000
125k50k
Cache Hit Rate0%
0% (No cache)50%90% (Max)
Total Input Spend

$0.2000

2.00M input tokens

Total Output Spend

$0.2000

0.50M output tokens

Estimated Total Cost

$0.40

Avg: $0.00040 / req

Rates: $0.10 in / $0.40 out per million tokens (Google AI Studio)Last verified: Aug 10, 2026

Compare with Similar Models

Anthropic

Claude 3.7 Sonnet

Anthropic's first hybrid reasoning frontier model with dynamic thinking budget control and state-of-the-art coding capabilities.

Compare Gemini 2.0 Flash vs Claude 3.7 Sonnet
Google DeepMind

Gemini 2.0 Pro (Experimental)

Google's flagship intelligence model engineered for complex coding, mathematical proofs, and 2M token context.

Compare Gemini 2.0 Flash vs Gemini 2.0 Pro (Experimental)
OpenAI

OpenAI o3-mini

Next-generation cost-efficient reasoning model optimized for STEM, competitive math, and coding.

Compare Gemini 2.0 Flash vs OpenAI o3-mini

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.