Anthropic Commercial APIStatus: active

Claude 3.5 Haiku

Ultra-fast and cost-efficient frontier intelligence model for high-throughput and sub-agent workflows.

Released: Nov 4, 2024
Last Verified: Aug 10, 2026
Context Window

200k

Tokens

Max Output

8.192k

Output limit

Architecture

Dense Transformer

Model family

Parameters

Not Disclosed

Total / Active

Input Token Cost

$0.80

Per 1M tokens

Output Token Cost

$4.00

Per 1M tokens

Model Overview

Claude 3.5 Haiku offers rapid execution speeds and strong reasoning at low per-token pricing, matching previous generation flagship models (Claude 3 Opus) across multiple benchmarks while operating at fraction of the latency.

Developer Implementation Notes

Ideal for sub-agent routing, real-time chatbots, and high-frequency API tasks. Input pricing is $0.80/MTok with $0.08/MTok cached.

Key Strengths

  • Very low latency and high tokens/second throughput
  • Matches Claude 3 Opus on coding benchmarks at 1/15th the price
  • 200k context window supported

Limitations & Boundaries

  • Text-only; image input is not supported on Haiku 3.5
  • Less nuanced on multi-turn complex mathematical derivations

Capabilities & Modalities

Modalities:text
Reasoning Chains:No
Tool / Function Calling:Supported
Structured Outputs:Supported
Fine-Tuning:No
Local Deployment:Cloud Only

Best Production Use Cases

  • Customer support routing & triage
  • High-volume synthetic data generation
  • Fast code autocomplete and linter fixes

Verified Benchmark Results

Standardized evaluations with methodology notes and authoritative citation links.

View all industry benchmarks →
Coding Official
SWE-bench Verified

40.6%

Pass@1 with standard developer scaffold

Coding Official
LiveCodeBench

43.1%

Pass@1, 0-shot code generation

Reasoning & Math Official
GPQA Diamond

41.6%

Zero-shot Chain-of-Thought

General Knowledge & Frontier Official
MMLU-Pro

65.1%

5-shot Chain-of-Thought standard

Pricing & Inference Cost Calculator

Token Cost Estimator – Claude 3.5 Haiku

Calculate projected inference spend with prompt caching

Quick Workload Presets

Input Tokens / Req2,000
100100k200k
Output Tokens / Req500
5016k32k
Number of Requests1,000
125k50k
Cache Hit Rate0%
0% (No cache)50%90% (Max)
Total Input Spend

$1.6000

2.00M input tokens

Total Output Spend

$2.0000

0.50M output tokens

Estimated Total Cost

$3.60

Avg: $0.00360 / req

Rates: $0.80 in / $4.00 out per million tokens (Anthropic API)Last verified: Aug 10, 2026

Compare with Similar Models

Anthropic

Claude 3.7 Sonnet

Anthropic's first hybrid reasoning frontier model with dynamic thinking budget control and state-of-the-art coding capabilities.

Compare Claude 3.5 Haiku vs Claude 3.7 Sonnet
Google DeepMind

Gemini 2.0 Pro (Experimental)

Google's flagship intelligence model engineered for complex coding, mathematical proofs, and 2M token context.

Compare Claude 3.5 Haiku vs Gemini 2.0 Pro (Experimental)
OpenAI

OpenAI o3-mini

Next-generation cost-efficient reasoning model optimized for STEM, competitive math, and coding.

Compare Claude 3.5 Haiku vs OpenAI o3-mini

Data Accuracy & Verification Notice

AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.