Overview
Inference pricing for models is designed to be straightforward and predictable. Instead of relying
on complex token-based pricing (which doesn’t make sense for non-text-generation models), we
calculate costs based on
Inference Meter Price and Time to First Inference.Formula
Key Features
Instance-Based Pricing
- Models run on instances optimized for RAM usage.
- Instances are categorized by size (e.g.,
Micro,Small,Super). - LLMs (Large Language Models) have their own specific pricing meters.
Transparent API Response Metadata
Each API response includes:Inference MeterInference Meter PriceInference TimeInference Cost
Prices
Language Models
All other models
Example Pricing
A developer runs an LLM on aMicro instance with an Inference Meter Price of $0.0000872083/sec. They configure their cluster to shut down after 1 minute of inactivity. They perform non-stop streaming inference for 9 minutes, then stop. Since the cluster shuts down at 10 minutes, total cost is:
Real-World Savings
GPT-4o Cost Breakdown
- ✅ Our pricing is significantly cheaper than GPT-4o for continuous inference.
- ✅ For real-time AI workloads, our GPU-based pricing provides better cost efficiency.