Skip to main content

Overview

Inference pricing for models is designed to be straightforward and predictable. Instead of relying on complex token-based pricing (which doesn’t make sense for non-text-generation models), we calculate costs based on Inference Meter Price and Time to First Inference.

Formula

Key Features

Instance-Based Pricing

  • Models run on instances optimized for RAM usage.
  • Instances are categorized by size (e.g., Micro, Small, Super).
  • LLMs (Large Language Models) have their own specific pricing meters.

Transparent API Response Metadata

Each API response includes:
  • Inference Meter
  • Inference Meter Price
  • Inference Time
  • Inference Cost

Prices

Language Models

All other models

Example Pricing

A developer runs an LLM on a Micro instance with an Inference Meter Price of $0.0000872083/sec. They configure their cluster to shut down after 1 minute of inactivity. They perform non-stop streaming inference for 9 minutes, then stop. Since the cluster shuts down at 10 minutes, total cost is:

Real-World Savings

GPT-4o Cost Breakdown

  • ✅ Our pricing is significantly cheaper than GPT-4o for continuous inference.
  • ✅ For real-time AI workloads, our GPU-based pricing provides better cost efficiency.