Enterprises deploying large language models in 2026 face a messy, consequential choice: how to pay. Cloud APIs still sell by tokens, cloud and specialist vendors offer GPU-second or inference-second billing, and newer SaaS bundles package unlimited usage with feature tiers. Each approach changes incentives, forecasting, vendor lock-in and the true total cost of ownership (TCO) for workplace AI.
Why pricing model matters beyond the sticker price
Pricing is not just about unit cost. The billing model affects architecture (centralized vs edge), product design (concise prompts vs long-context memory), procurement cadence (cap-ex vs op-ex), and governance (auditability and data routing). A model that looks cheap per request can rapidly balloon when concurrency, retrains, or heavy context windows are required. Conversely, a high subscription fee can be cheaper at scale for high-volume, predictable workloads.
Three dominant models explained
- Per-token (or per-character) API billing: Common with early cloud LLM APIs. Price scales with input+output tokens processed. Encourages concise prompts and low-response lengths.
- Per-GPU-second / per-inference-second: Charges for actual compute time used to run an inference on hardware (GPU/TPU), often reported as GPU-seconds or vCPU-seconds. Better ties cost to model size and latency but makes batching and concurrency important cost levers.
- Subscription / seat / feature-tier: Fixed monthly or annual fee for a service bundle (unlimited queries, SLA, integrations). Predictable but can be expensive for low-usage teams and may hide per-tenant scaling costs.
Scenario modeling: three enterprise assistants
To compare models concretely, we use three hypothetical enterprise assistants. These are illustrative scenarios—change any assumption and results will shift materially.
-
Helpdesk Triage Assistant (High QPS, short responses)
Volume: 200,000 queries/month. Average context: 150 input tokens, 50 output tokens. Latency target: 2s average inference.
-
Knowledge-Worker Research Assistant (Medium volume, long context)
Volume: 25,000 queries/month. Average context: 1,500 input tokens (long RAG context), 400 output tokens. Latency tolerance: up to 10s per request.
-
Executive Briefing Generator (Low volume, heavy compute)
Volume: 1,200 requests/month. Average context: 5,000 tokens (multi-doc synthesis), 1,000 output tokens. Each request runs multiple reasoning chains and a second-pass summarization.
We make no claims about real vendor prices. Below are conceptual comparisons based on measurement units that procurement teams can apply to live quotes.
How pricing shapes cost per workload
Key internal metrics to calculate:
- Tokens per request (input + output)
- Average inference time per request (seconds) on selected model
- Concurrent inferences and batching efficiency
- Monthly volume
Example: For the Helpdesk assistant, tokens/request = 200. Under a per-token API, cost = tokens/request × token-price × volume. Under GPU-second pricing, cost = inference-seconds/request × GPU-second-price × volume. Under subscription, cost = fixed monthly fee.
Implication: High QPS, short-token workloads usually favor per-token if token prices are low; but if the vendor charges extra for context windows or retrievals (RAG calls), costs can grow. For long-context, compute-heavy tasks, GPU-second billing or subscription with unlimited inferences becomes more attractive.
Operational and commercial trade-offs
- Predictability vs efficiency: Subscriptions maximize budgeting predictability; per-unit models incentivize engineering optimization (shorter prompts, caching, batching).
- Scaling economics: GPU-second has first-order link to model size and latency. Bigger models cost more per second; however, faster GPUs and batching can lower per-request seconds dramatically.
- Vendor lock-in: Subscriptions and productized assistants often include proprietary RAG layers and UI integrations, increasing switching costs. Per-token and GPU-second models are more modular but still risk model-specific embedding formats or fine-tuning artifacts.
- Governance and compliance: Subscription bundles may limit data export or require enterprise contracts for data residency. Per-inference billing on private infra shifts control (and capex) back to the customer.
- Hidden add-ons: Many vendors add charges for streaming, real-time SLAs, guardrails, or private model hosting. Ensure total cost of service is included in quotes.
Market dynamics in mid-2026 that shape TCO
Three trends change the arithmetic today:
- Hardware differentiation: Purpose-built inference chips and optimized runtimes (sparsity, quantization, compilation) reduce GPU-seconds for many models, lowering compute-based bills.
- Hybrid contract models: Vendors increasingly offer blended contracts—base subscription + per-use overage—aimed at mid-size customers to capture predictable revenue while scaling with usage.
- Edge and on-prem options: For regulated industries, on-prem inference (hardware capex + managed services) can be cheaper at sustained high volume and avoids per-request egress and data residency fees common in cloud subscriptions.
Procurement checklist: questions to force apples-to-apples comparisons
- Ask vendors for a unitized quote: token price, GPU-second price, or per-seat price, plus any charges for context windows, retrievals, or streaming.
- Request representative performance numbers (ms/request) on the model variant you’ll use and the batching assumptions. Benchmark on your prompts if possible.
- Clarify what “unlimited” means in subscriptions—are there soft caps, fair-use policies, or throttles?
- Include integration and guardrail costs: will you pay extra for private endpoints, audit logs, or fine-tuning jobs?
- Model migration: what artifacts (embeddings, fine-tuned weights) can you export? Are there egress charges for moving to a new vendor?
- Negotiate pilot-to-production runway: tiered pricing that drops per-unit costs as volume grows incentivizes adoption but requires clear volume milestones.
How engineering teams should respond
Engineering and product teams must instrument cost metrics as rigorously as performance metrics. Track:
- Cost per completed user task (not per call)
- Tokens per task and inference-seconds per task
- Cache hit rates for RAG retrieves and rate of reuse for generated artifacts
- Concurrency and utilization of inference pools
Small changes—shortening prompts by 10–20%, batching requests, caching embeddings—can move TCO substantially under per-unit models. Under subscription models, focus on increasing adoption and measuring ROI per user to justify fixed fees.
Recommendations: matching model to workload
- High-volume, short-response services (chatbots): Start with per-token if pricing is competitive and you can optimize prompts aggressively. Consider subscription only if vendor adds business value (SLA, analytics) you can’t replicate.
- Long-context, research, and synthesis workflows: Prefer GPU-second or subscription bundles that include heavy compute—these workloads incur high token counts but more critically, long inference times and multiple passes.
- Regulated or predictable enterprise suites: Evaluate on-prem or managed subscription with private hosting. The fixed cost and compliance controls can justify capex.
Outlook: convergence, not winner-takes-all
Expect more hybrid pricing and tooling to emerge. Vendors will offer blended models, usage discounts, or conversion pathways (e.g., credits for committed spend to convert per-token charges into subscription allowances). For procurement teams, the task is less picking a “best” model and more instrumenting workloads, benchmarking real costs in production, and negotiating contracts that reflect your usage profile and compliance needs.
The most important metric is not token-price or GPU-second price alone—it’s the cost to deliver business value: cost per resolved ticket, cost per analyst-hour saved, or cost per compliant automation. Treat pricing models as levers to optimize that metric, and build the telemetry and contract terms that let you pull those levers without surprises.