← Agent Frameworks 🕐 22 min read
Agent Frameworks

LLM Vendor Pricing History, Cost Benchmarking, and Enterprise Contract Intelligence (2026)

Frontier LLM API prices have fallen 80–99% since March 2023 depending on the tier. GPT-4-equivalent capability now costs $2–3 per million input tokens; in March 2023 it cost $30.

Audience: FinOps leads, procurement teams, CTOs negotiating AI platform contracts
Last verified: June 17, 2026


Frontier LLM API prices have fallen 80–99% since March 2023 depending on the tier. GPT-4-equivalent capability now costs $2–3 per million input tokens; in March 2023 it cost $30. The pace of price cuts — averaging one significant reduction every 90 days across the major providers — has structural implications for procurement: any per-token rate locked beyond 12 months is virtually guaranteed to become above-market. The smart procurement posture in 2026 is on-demand-first with model-substitution rights, not committed rates.

This note covers the full arc: price history with timestamps, a comprehensive June 2026 pricing reference table across all major providers, performance-per-dollar benchmarking methods, EDP negotiation tiers, multi-cloud cost normalization under FOCUS 1.4, price-path scenario planning, and third-party benchmarking resources for ongoing vendor comparison.


1. Price Trajectory 2023–2026

The Core Fact

Andrew Ng summarized the trajectory in August 2024: GPT-4 cost $36 per million tokens (blended 80/20 input/output) at its March 2023 launch. By mid-2024 GPT-4o cost $4 per million on the same blended basis — a ~79% annual price decline. By mid-2026 the equivalent capability (GPT-4o, Claude Sonnet 4.6) costs $3–5 per million blended, and mid-tier open-weight models served on commodity inference infrastructure (DeepSeek V3, Llama 4 via Groq) cost $0.20–$0.90 per million. The total collapse from GPT-4’s launch to mid-2026 equivalent capability is approximately 90–95%.

Milestone Price Table (Input / Output, $/1M tokens)

Model Launch Date Input Output Notes
GPT-4 8K Mar 2023 $30.00 $60.00 First widely available frontier API
GPT-4 32K Mar 2023 $60.00 $120.00 Extended context premium
GPT-4 Turbo Nov 2023 $10.00 $30.00 67% input cut at DevDay
Claude 3 Haiku Mar 2024 $0.25 $1.25 Entry-tier price shock — set new floor
Claude 3 Sonnet Mar 2024 $3.00 $15.00 Mid-tier benchmark
Claude 3 Opus Mar 2024 $15.00 $75.00 Flagship, matched GPT-4 quality
GPT-4o (launch) May 2024 $5.00 $15.00 50% cut vs GPT-4 Turbo on input
Gemini 1.5 Flash May 2024 $0.075 $0.30 Google’s aggressive mid-tier entry
GPT-4o (Oct 2024) Oct 2024 $2.50 $10.00 Second 50% cut in 5 months
Claude 3.5 Sonnet Jun 2024 $3.00 $15.00 Surpassed Opus at Sonnet price — pivotal
Claude 3.5 Haiku Oct 2024 $0.80 $4.00 Revised down from $1.00 in Dec 2024
Gemini 2.0 Flash Dec 2024 $0.10 $0.40 Commodity-tier reference price
GPT-4.1 Nano Q1 2025 $0.10 $0.40 OpenAI’s sub-$1 flagship entry
Gemini 2.5 Flash Q1 2025 $0.30 $2.50 Current Google speed/cost champion
Gemini 2.5 Pro Q1 2025 $1.25 $10.00 Tiered: higher rate above 200K ctx
Claude Haiku 4.5 2026 $1.00 $5.00 Haiku 4.x — more capable than 3.5
Claude Sonnet 4.6 2026 $3.00 $15.00 Current mid-tier standard
Claude Opus 4.8 May 28, 2026 $5.00 $25.00 Tops Artificial Analysis index at 61.4
o3 2026 $2.00 $8.00 Reasoning model; billing includes hidden tokens
o4-mini 2026 $0.55 $2.20 Best reasoning per dollar current generation
GPT-4o (current) 2026 $2.50 $10.00 Stable since Oct 2024

Key milestones by category:

  • First 10x drop: GPT-4 ($30/$60) → GPT-4 Turbo ($10/$30) — 7 months, November 2023
  • Second 10x drop: GPT-4 Turbo → GPT-4o-class ($2.50/$10) — 12 months additional
  • Haiku-class emergence: Claude 3 Haiku at $0.25 input (March 2024) proved sub-$1 frontier-adjacent capability was commercially viable
  • Parity inversion: Claude 3.5 Sonnet (June 2024) matching or exceeding Opus quality at Sonnet price — permanently ended the “pay more for better” assumption at the frontier tier

What Drives Price Cuts

Three forces compound:

  1. GPU efficiency gains. H100 → H200 → B100 inference throughput per dollar improved 3–5x from 2023 to 2026. Training and inference cost curves both dropped sharply. Providers pass through some of this as price competition.

  2. Competition. Google’s Gemini 1.5 Flash at $0.075 input (May 2024) forced OpenAI’s October 2024 GPT-4o cut. Meta’s open-weight Llama series, served by Groq and Together AI at commodity margins, created a price ceiling for any task where Llama-class quality suffices. DeepSeek V3’s release in late 2024 triggered another round of cuts — showing that non-US competition accelerates the pressure.

  3. Economies of scale. Anthropic, OpenAI, and Google now run at token volumes orders of magnitude above 2023 levels. Fixed infrastructure costs spread over more tokens; providers sustain margins at lower per-unit prices.

Procurement Implication

The average gap between major price drops across the big three providers has been approximately 90 days. A 24-month per-token rate lock signed today will almost certainly be above-market within 6–9 months. Never accept per-token price locks exceeding 12 months. The correct contract structure is a spend commitment (which earns discounts) with pricing indexed to the provider’s published list rate, subject to downward revision at each contract anniversary.


2. Current Pricing Reference Table — June 2026

All prices in USD per 1 million tokens. “Batch” = async batch API discount. “Context” = maximum context window.

Anthropic Claude

Model Input Output Batch Input Batch Output Context Notes
Claude Opus 4.8 $5.00 $25.00 $2.50 $12.50 200K Flagship; tops Artificial Analysis at 61.4
Claude Sonnet 4.6 $3.00 $15.00 $1.50 $7.50 200K Production workhorse
Claude Haiku 4.5 $1.00 $5.00 $0.50 $2.50 200K Speed/cost optimized
Prompt cache (read) $0.30 90% discount on cached input
Prompt cache (write) $3.75 One-time write cost

Bedrock pricing: Identical per-token rates to direct API. Bedrock adds data transfer costs for large volumes and regional pricing differentials; the effective delta is small (<5%) for typical enterprise workloads but can reach 6–8% for high-volume streaming use cases.

OpenAI

Model Input Output Batch Input Batch Output Context Notes
GPT-4o $2.50 $10.00 $1.25 $5.00 128K Stable flagship
o3 $2.00 $8.00 $1.00 $4.00 200K Reasoning; output bill includes reasoning tokens
o4-mini $0.55 $2.20 $0.28 $1.10 200K Best reasoning value currently
GPT-4.1 Nano $0.10 $0.40 $0.05 $0.20 128K Commodity tier

Reasoning token billing warning. OpenAI’s o-series models generate “reasoning tokens” — internal chain-of-thought — that are billed as output tokens but not visible in the response. A query showing 500 output tokens in the response may consume 3,000+ billed tokens. Budget calculations must account for a 3–6x reasoning multiplier on complex tasks. Build this into any cost model before comparing o3/o4-mini to non-reasoning models.

Azure OpenAI: Enterprise Agreement pricing is available. Discounts via Azure EA are negotiated as part of the broader Microsoft enterprise commitment; AI workloads are increasingly included in EA conversations. Published Azure OpenAI rates match OpenAI API rates; EA discounts run 10–20% for customers with large Azure commitments.

Google Gemini (via AI Studio / Vertex AI)

Model Input (≤200K) Output (≤200K) Input (>200K) Output (>200K) Context Notes
Gemini 2.5 Pro $1.25 $10.00 $2.50 $15.00 1M Tiered pricing on context length
Gemini 2.5 Flash $0.30 $2.50 $0.30 $2.50 1M Flat rate; best long-context value
Gemini 2.0 Flash $0.10 $0.40 $0.10 $0.40 1M Commodity reference tier

Vertex AI vs AI Studio: Vertex adds enterprise data governance (VPC-SC, audit logs, DLP), no persistent training on customer data by default, regional deployment options. List rates are the same; Committed Use Discounts (CUDs) via Google Cloud apply to Vertex spend when included in a broader GCP CUD negotiation.

Context window cost trap. The Gemini 2.5 Pro pricing step at 200K context — from $1.25 to $2.50 input — means a document-processing workload that routinely exceeds 200K tokens pays 2x on input. Model your P90 context usage before assuming Gemini 2.5 Pro’s headline rate applies.

Meta Llama (Open-Weight, Hosted Inference)

Meta’s Llama models are open-weight (free download). Hosted inference costs vary significantly by provider:

Model Provider Input Output Notes
Llama 4 Maverick AWS Bedrock $0.50 $1.50 Enterprise SLA; Bedrock premium
Llama 3.3 70B Together AI $0.88 $0.88 Mid-tier hosted
Llama 3.3 70B Groq $0.59 $0.79 Speed-optimized
Llama 3.3 70B DeepInfra $0.23 $0.40 Cheapest verified host as of Apr 2026
Llama 3.1 8B Groq $0.05 $0.08 Sub-$0.10 commodity tier

Bedrock premium for open-weight models is substantial: Llama 3.3 70B on Bedrock costs ~3x vs DeepInfra for the same model. The Bedrock premium buys VPC isolation, CloudTrail, SOC 2 / HIPAA eligibility, and EDP credit accumulation.

Mistral AI

Model Input Output Context Notes
Mistral Large 2 $2.00 $6.00 128K Competitive mid-tier; strong European data residency story
Mistral Small 3 $0.10 $0.30 32K Sub-$0.50 blended; document classification use case

Cohere

Model Input Output Context Notes
Command R+ $2.50 $10.00 128K Strong RAG/structured-extraction benchmark
Command R $0.15 $0.60 128K Budget RAG tier

Cohere’s differentiator is native reranking (Rerank v3 at ~$2/1M documents) tightly integrated with Command R+. For RAG pipelines where retrieval quality matters, the combined retrieval+generation cost may be lower than frontier models with prompt engineering workarounds.


3. Performance-Per-Dollar Benchmarking

Why MMLU Is No Longer the Differentiator

As of 2026, frontier scores on MMLU sit at 90%+ across all major models. MMLU is no longer useful for provider comparison. The relevant benchmarks are:

  • GPQA Diamond — graduate-level scientific reasoning; real differentiation between tiers
  • SWE-Bench Verified / SWE-Bench Pro — code generation on real GitHub issues; Claude Opus 4.8 at 69.2%
  • LiveCodeBench — coding; updated continuously, resists memorization
  • HumanEval — code generation; still used, but saturating at the frontier
  • Humanity’s Last Exam (HLE) — hardest public benchmark; frontier discrimination
  • LMSYS Chatbot Arena ELO — human preference signal; strong proxy for real-world quality

Artificial Analysis Intelligence Index v4.0 is the most useful single composite: a weighted average of MMLU-Pro, GPQA Diamond, HumanEval, MATH, LiveCodeBench, and others, normalized to a 0–100 score. As of June 2026, Claude Opus 4.8 tops the index at 61.4.

Value Score: Quality Per Dollar

A practical value score = (Intelligence Index score) / (blended price per 1M tokens). “Blended” assumes 80% input, 20% output at typical enterprise prompt/completion ratios.

Model Intelligence Index (approx) Blended $/1M Value Score (index/dollar)
Claude Opus 4.8 61.4 $9.00 6.8
Claude Sonnet 4.6 ~52 $5.40 9.6
GPT-4o ~50 $4.00 12.5
o4-mini ~48 $0.88 54.5
Gemini 2.5 Pro ~50 $3.00 16.7
Gemini 2.5 Flash ~44 $0.74 59.5
Llama 3.3 70B (Groq) ~42 $0.63 66.7
Mistral Large 2 ~44 $2.80 15.7
Gemini 2.0 Flash ~38 $0.18 211

Note: Value scores compress at the frontier because intelligence gains are marginal while cost differences are large. Gemini 2.0 Flash’s extreme value score reflects its role as a commodity routing tier, not frontier capability.

The Pareto Frontier (June 2026): Three models sit on the efficient frontier of quality vs. cost:

  • o4-mini — best reasoning value; use for structured problem-solving at sub-$1 blended
  • Gemini 2.5 Flash — best long-context value; 1M context window at $0.30 input is unmatched
  • Claude Sonnet 4.6 — best all-round production value; $3/$15 with 200K context, strong code and instruction following

Opus 4.8 and GPT-4o are below the frontier on pure cost-efficiency; you pay the Opus premium for top-1 benchmark performance on the hardest tasks (agentic coding, scientific reasoning). For most enterprise document and extraction workloads, Sonnet 4.6 or Gemini 2.5 Flash is the correct model.

Workload-Specific Recommendations

Code generation (production-grade, agentic):
Claude Opus 4.8 (SWE-Bench Pro 69.2%) if quality is paramount. o4-mini for interactive coding assistants where cost matters. Gemini 2.5 Flash for high-volume code review pipelines.

Document classification / extraction:
Gemini 2.5 Flash or Mistral Small 3. For long documents (>50K tokens), Gemini 2.5 Flash’s flat pricing and 1M context window is the clear winner. Cohere Command R is competitive for RAG-integrated pipelines.

Long-form reasoning (multi-step, complex analysis):
o3 or Claude Opus 4.8. Note the o3 reasoning-token billing multiplier — budget 4x the visible output token count for total cost.

Structured extraction (JSON, entities, schemas):
Claude Sonnet 4.6 (strong instruction following) or Gemini 2.5 Flash (cost). Cohere Command R+ for RAG-fed extraction pipelines.

High-volume, latency-sensitive routing (classification, triage, intent detection):
Gemini 2.0 Flash ($0.10/$0.40) or Llama 3.1 8B via Groq ($0.05). These are 50–200x cheaper than frontier models with acceptable quality for deterministic classification tasks.

How to Build Your Own Cost-Adjusted Benchmark

A workload-specific benchmark is more valuable than any published leaderboard for procurement decisions:

  1. Sample your production traffic. Pull 200–500 representative real inputs across your P50, P90, P99 prompt lengths.
  2. Define quality metrics. For extraction: F1 on entity set. For classification: accuracy on labeled holdout. For generation: human rating or embedding similarity to gold standard.
  3. Run all candidate models in parallel against the same inputs. Use LiteLLM’s unified interface for provider-agnostic execution.
  4. Record simultaneously: quality score, latency (TTFT + total), input tokens, output tokens, estimated cost at list price.
  5. Plot each model on a 2D scatter: quality (y-axis) vs. blended cost per query (x-axis). Models on the Pareto frontier (upper-left) are candidates. Models below the frontier are dominated — avoid them.
  6. Add latency as a color dimension. Some Pareto-efficient models are only efficient at acceptable latency; filter out any model exceeding your P95 latency SLO before comparing costs.

Run this benchmark quarterly. As new model versions ship, the frontier shifts. A model that was Pareto-efficient in January may be dominated by March.


4. Enterprise Contract Negotiation

AWS Bedrock — EDP Structure

The AWS Enterprise Discount Program (EDP) is the primary lever for Bedrock cost optimization. Key mechanics:

How Bedrock integrates with EDP:
Bedrock consumption accumulates against your EDP commitment at full list price, and the EDP discount applies at billing. Bedrock spend added to your EDP commitment can help you reach the next discount tier faster — this is useful for organizations already at $3–4M AWS annual spend, where adding $1M of Bedrock pushes them into the next tier.

Observed EDP discount tiers (2026 benchmarks):

Annual Commitment Typical EDP Discount Range
$1M – $3M 8–12%
$3M – $5M 12–18%
$5M – $10M 18–25%
$10M – $20M 25–30%
$20M – $100M 30–43%

Source: VendorBenchmark, Redress Compliance, The Negotiation Experts (2026 surveys)

These are ranges, not guarantees. The actual discount depends on your growth trajectory, service mix, and whether you can credibly threaten migration. Pure AI workloads with no other AWS dependency have less leverage; multi-service AWS shops with SageMaker, EC2, and Bedrock combined have more.

AI-specific Bedrock negotiation (emerging in 2026):
High-throughput token workloads (>500M tokens/month) can negotiate provisioned throughput pricing. Provisioned throughput offers predictable capacity and can run 6–8% cheaper than on-demand at scale. The risk is over-commitment if volumes don’t materialize — unused provisioned throughput is not refunded.

Key EDP negotiation tactics:

  • Include Bedrock explicitly in your EDP scope document — it was sometimes excluded in 2024 negotiations and defaulted to standard pricing
  • Negotiate model substitution rights: if a newer Claude or Llama version launches at lower cost, you retain the right to switch without renegotiation
  • Include quarterly pricing review clauses indexed to provider list rates
  • Request audit rights on token metering — not standard, but achievable for $5M+ commitments

Anthropic Direct vs. AWS Bedrock — When to Go Direct

Go direct to Anthropic when:

  • Annual Claude API spend exceeds $5M (Anthropic’s threshold for meaningful committed-spend discounts)
  • You require data processing agreements specific to your jurisdiction (EU DPA, HIPAA BAA) that Bedrock’s standard terms don’t cover
  • Your workload is Claude-only with no dependency on other AWS services (EDP leverage is limited)
  • You need access to early/preview models before Bedrock GA rollout (Anthropic API gets new Claude versions days to weeks before Bedrock)

Stick with Bedrock when:

  • Your AWS EDP already covers substantial non-AI spend — the incremental discount from adding Bedrock to an existing EDP typically exceeds standalone Anthropic committed-spend discounts
  • You need Bedrock’s AWS-native integrations (VPC endpoints, CloudTrail, S3, Lambda, SageMaker Pipelines)
  • Your security/compliance posture requires AWS-native controls (HIPAA, FedRAMP, SOC 2 within AWS boundary)
  • You use multiple foundation models (Claude + Llama + Nova + Titan) — Bedrock’s unified billing and IAM is worth the complexity overhead

Bedrock vs. Direct rate parity: Per-token rates are identical. The delta comes from: (a) EDP discount on Bedrock-via-AWS vs. standalone Anthropic commitment discount, (b) data transfer costs on Bedrock for high-throughput streaming, © provisioned throughput pricing on Bedrock.

Azure EA for OpenAI

Azure OpenAI pricing matches OpenAI API list rates at standard tiers. Enterprise value is unlocked via:

  • Azure Committed Use: OpenAI spend can be included in Azure EA monetary commitment, burning down pre-purchased credits
  • EA discount application: Customers with existing Azure EA (typically 10–20% discount on compute) can negotiate to apply that discount to Azure OpenAI spend with Microsoft in renewals
  • Content filtering and compliance: Azure OpenAI includes Azure-native content filtering, logging via Azure Monitor, and RBAC — valuable for regulated industries

Practical note: Microsoft’s AI pricing negotiations in 2026 are more complex because Copilot seats (M365 Copilot at $30/user/month), Azure OpenAI API, and GitHub Copilot are increasingly bundled. Procurement teams should model all three in the same negotiation rather than treating them as separate line items.

Contract Red Flags

Any enterprise AI contract should be reviewed for these clauses before signature:

  1. Per-token price locks without downward adjustment. A 2-year contract at June 2026 rates will be above-market by early 2027. Acceptable: spend commitment with pricing indexed to provider’s published list rate. Unacceptable: fixed per-token rate for >12 months.

  2. No model substitution clause. If the contract names specific model versions (e.g., “claude-sonnet-4-6”), you lose the right to migrate to cheaper/better versions without renegotiation. Require language allowing substitution to any model in the same provider family at the same or lower cost.

  3. No audit rights on token metering. You should have the right to verify token counts against your own measurement. Without audit rights, overbilling is undetectable at scale. For $1M+ contracts, require monthly usage reports at the model, project, and team level.

  4. Automatic renewal at current rates. The standard renewal trap: contract auto-renews at current year’s rates without a renegotiation trigger. Require a 90-day pre-renewal window with price-reduction right if the provider has cut list rates since contract execution.

  5. Reasoning token opacity. For o-series or other chain-of-thought models, ensure the contract specifies whether reasoning tokens are billed and how they are reported. Without this, cost overruns on reasoning-heavy workloads are common and unauditable.

  6. Data training opt-out not default. Verify that the contract’s default is no training on your data, not an opt-out-required model. Both Anthropic and OpenAI enterprise agreements default to no training; some third-party hosted model providers do not.


5. Price Normalization for Multi-Cloud Comparison

FOCUS 1.4 — Current State

The FinOps Open Cost and Usage Specification v1.4 was ratified by the FinOps Foundation Steering Committee on June 4, 2026 — 13 days before this note was written. It is the first version to formally include AI token economics columns in the core specification (not extension-only status).

FOCUS 1.4 adds:

  • TokenInputQuantity, TokenOutputQuantity, TokenCacheReadQuantity, TokenCacheWriteQuantity columns for AI billing normalization
  • InferenceType attribute (on-demand, batch, provisioned)
  • ModelFamily and ModelVersion attributes for cross-provider model comparison

Current limitations: Microsoft has announced plans to support FOCUS 1.4 (Azure OpenAI will export FOCUS-compliant data), but the rollout timeline is H2 2026. AWS FOCUS export for Bedrock is available as of June 2026 for core columns; the AI-specific columns are in GA as part of the June 4 ratification but may take 60–90 days for all providers to implement fully. In practice, many enterprises are still hand-normalizing AI billing data.

Practical Normalization Approach (Provider-Agnostic)

Until FOCUS 1.4 exports are universally available, the standard normalization approach is:

Step 1: Standardize on $/1M input tokens and $/1M output tokens.
All major providers quote at this granularity. This is the lowest-common-denominator unit and should be your procurement-facing KPI.

Step 2: Model blended rate for workload comparison.
Define your workload’s typical input:output ratio. Enterprise document processing is often 80:20 (input-heavy). Agentic coding pipelines are often 60:40 or even 50:50 (output-heavy). Blended rate = (input_fraction × input_price) + (output_fraction × output_price). Always compare blended rates — comparing only input prices or only output prices produces misleading rankings.

Step 3: Normalize for context window utilization.
A 128K context model and a 1M context model are not equivalent for the same task if your prompts average 50K tokens. For tasks that fit in 8K, both are equivalent — use the cheaper provider’s per-token rate. For tasks that require 500K tokens, only Gemini 2.5 Flash and Gemini 2.5 Pro are viable — compare their prices, not their prices against shorter-context models.

Step 4: Add reasoning token multiplier for o-series.
For OpenAI o3/o4-mini, multiply expected output tokens by 3–6x to get total billed tokens. For Claude Opus 4.8’s extended thinking mode (when enabled), apply a similar multiplier. This step is often omitted and causes 3–5x cost overruns on reasoning-intensive workloads.

Step 5: Add batch discount where applicable.
Anthropic and OpenAI both offer 50% batch discounts for async workloads (non-latency-sensitive processing). For document classification, data enrichment, and offline extraction pipelines, batch pricing is almost always the right choice. The effective blended rate for a batch pipeline is roughly half the on-demand equivalent.

LiteLLM for Multi-Provider Cost Normalization

LiteLLM (open-source, MIT license) is the standard tool for cross-provider normalization in 2026. Key capabilities for FinOps:

  • Unified cost ledger: LiteLLM tracks input tokens, output tokens, and cost per call across all providers using a single normalized schema. Costs are recorded in $/1M tokens equivalents.
  • Provider routing by cost: You can configure routing rules such as “use Gemini 2.5 Flash for any request where Claude Sonnet cost would exceed $0.01” — automated cost-ceiling routing.
  • Budget guardrails: Per-team, per-project, and per-user daily/monthly spend budgets with hard cutoffs. Prevents runaway costs from agentic loops.
  • Model fallback on cost: If a premium model exceeds budget, LiteLLM can fall back to the next cheapest model in your approved list automatically.

LiteLLM’s hosted gateway: $149/month (Hobby, 10M tokens) to $1,999/month (Enterprise, 500M tokens). Self-hosted is free. For most enterprises with >$50K/month AI spend, the cost normalization and budget-control value justifies the gateway overhead.

Portkey and Helicone are alternative gateways with similar normalization capabilities; choice is largely infrastructure preference (Portkey skews toward LangChain integrations, Helicone toward observability-first workflows).


6. Cost Forecasting Under Price Uncertainty

Price-Path Scenarios

Three scenarios for 2026–2028 frontier model pricing:

Bull Case: 40% annual price cuts continue
Driven by: continued GPU efficiency gains (B200/GB200 becoming the inference standard), intensifying competition from open-weight models (Llama 5, DeepSeek V4), and Google’s infrastructure cost advantages compounding. Claude Sonnet-class capability drops from $3 input to $1.80 (2027) to $1.08 (2028).

Base Case: 20% annual cuts
Price compression moderates as the “easy” efficiency gains saturate. Competition stabilizes around 4–5 major providers. Claude Sonnet-class: $3 → $2.40 (2027) → $1.92 (2028).

Bear Case: Pricing stabilizes
Providers consolidate, open-weight competition plateaus, GPU costs stabilize. Frontier model prices hold near 2026 levels with only incremental cuts. Claude Sonnet-class stays at $2.50–3.00 through 2028.

Historical data (2023–2026) tracks closest to the bull case. The prudent budget model is base case for committed spend, bull case for on-demand optimization, bear case for vendor contract negotiations (the downside scenario justifies requiring price-reduction clauses).

Budget Structure to Benefit from Price Cuts

Avoid: Annual committed-rate contracts. Even with a 15% EDP discount, a committed rate signed at June 2026 prices will likely be above-market by early 2028 if price cuts continue at base-case pace.

Prefer: Spend commitments with market-rate pricing. Structure your EDP as a spend volume guarantee (“we will spend $3M on AWS AI services this year”) rather than a rate guarantee (“we will pay $3/1M input tokens for Sonnet”). The spend commitment earns the discount tier; pricing tracks the market.

Hedge with open-weight runway. Maintain deployed capability to run Llama 4 or equivalent open-weight models on your own infrastructure (SageMaker, Vertex AI Model Garden, or bare-metal GPU). This is negotiating leverage — a credible threat to migrate high-volume workloads off proprietary APIs if prices don’t move favorably.

Use LiteLLM routing to capture price drops automatically. Configure model routing to “always use the current cheapest Pareto-efficient model for classification tasks” rather than hardcoding a model name. When a new model launches at lower cost with equivalent quality, your routing config updates; contracts don’t need to.

Budget Template for FinOps Teams

For a 12-month AI inference budget:

  1. Classify workloads by tier:

    • Tier 1 (frontier): agentic coding, complex reasoning, synthesis tasks → use Claude Opus 4.8 or o3
    • Tier 2 (production): document processing, extraction, Q&A → use Claude Sonnet 4.6 or GPT-4o
    • Tier 3 (commodity): classification, routing, triage → use Gemini 2.5 Flash or Llama 3.1 8B
    • Tier 4 (batch offline): data enrichment, bulk tagging → use batch API at 50% discount
  2. Model token volumes at P50 and P90 for each tier.

  3. Apply base-case price trajectory (20% annual reduction) for your planning horizon — budget the lower number, not 2026 list.

  4. Add 20% buffer for reasoning token overrun (o-series) and agentic loop cost overrun.

  5. Build in a quarterly review trigger: if actual spend is <70% of budget at Q2, revisit whether a committed-spend upgrade earns a better discount tier.


7. Third-Party Benchmarking and Price Monitoring Resources

Primary Resources

Artificial Analysis (artificialanalysis.ai)
The most rigorous quality/price/speed leaderboard. Tracks 356+ models on the Artificial Analysis Intelligence Index (composite of 10 benchmarks), provider-reported pricing, and live latency metrics. Updated continuously. As of June 2026, Claude Opus 4.8 tops the Intelligence Index at 61.4. Best used for: initial model selection, quarterly reviews, and contract negotiation evidence. The site publishes free data; an API is available for programmatic monitoring.

LMSYS Chatbot Arena (lmarena.ai)
ELO-based ranking from human preference comparisons. The most widely cited quality proxy for real-world user satisfaction. Frontier ELO scores in June 2026 range from 1,450 (strong mid-tier) to 1,561 (current top-1). Use as a sanity check on Artificial Analysis composite scores — the two rankings should be directionally consistent; large divergences signal benchmark gaming.

Vellum LLM Leaderboard (vellum.ai/llm-leaderboard)
Focused on enterprise use cases: instruction following, structured output, long-context recall. More practically oriented than pure benchmark composites. Good complement to Artificial Analysis for procurement decision-making.

OpenRouter (openrouter.ai)
Aggregates pricing from 100+ models and providers with live pricing data. Useful for checking third-party hosting prices for open-weight models (Llama, Mistral, DeepSeek) across Groq, Together AI, DeepInfra, and others simultaneously.

PricePerToken / AI Pricing Guru (pricepertoken.com, aipricing.guru)
Price aggregators with historical pricing data. AI Pricing Guru maintains per-provider pages updated in near-real-time. Useful for contract negotiations where you need to demonstrate that a provider’s rates have moved since your last negotiation.

LLM Stats (llm-stats.com)
Tracks 300+ models across speed, cost, and capability. Maintains a live leaderboard updated as new models ship.

Automating Price Monitoring

For enterprise procurement teams, manual monitoring is insufficient given 90-day average price drop cycles. Recommended automation:

  1. Scrape provider pricing pages weekly. Anthropic’s pricing page, OpenAI’s pricing API endpoint, Google AI Studio pricing, Vertex pricing calculator. Parse for the relevant model rows. Most providers expose pricing in structured HTML or JSON that can be scraped reliably.

  2. Subscribe to Artificial Analysis’s model updates. They publish structured pricing data that can be ingested via their API or scraped from their model pages.

  3. Configure alerts in your LiteLLM gateway. LiteLLM can be configured to alert when a new model is added that beats your current model on a defined value score. This catches launches that would update your routing immediately.

  4. Subscribe to provider changelogs. Anthropic’s changelog (anthropic.com/news), OpenAI’s platform changelog, Google’s AI Studio release notes. Price changes are announced here before updating pricing pages — advance notice of 24–72 hours is typical.

  5. Benchmark quarterly. Run your workload-specific benchmark (Section 3) on a 90-day cycle. New model launches between benchmark runs may not affect your routing immediately, but the quarterly benchmark ensures you don’t stay on a dominated model for more than one cycle.


Appendix: Key Numbers for Negotiation

Quick reference for procurement conversations:

Metric Value
GPT-4 input price, Mar 2023 $30/1M
Equivalent capability input price, Jun 2026 $2.50–3.00/1M
Price decline, 3 years ~90%
Average price-cut cycle (major providers) ~90 days
Recommended max rate lock duration 12 months
Batch API discount (Anthropic, OpenAI) 50%
Prompt cache discount (Anthropic) 90% on cache reads
Bedrock EDP discount at $5M commit ~18–25%
Bedrock EDP discount at $20M commit ~30–43%
o-series reasoning token multiplier 3–6x vs. visible output
Gemini 2.5 Pro pricing step threshold 200K tokens per prompt
LiteLLM routing cost saving (typical) 40–60% vs. single-provider
FOCUS 1.4 ratification date June 4, 2026

Sources