Model catalog
30+ models, one page
Real-time latency, throughput and price — filter by region, sort by any column.
Models
7-day requests
Updated
30+ models — latency, throughput and price at a glance
Official full-precision weights, no distillation. Prices are our wholesale cost (+5% to cover payment fees).
Prices are reference cost values, loading latency/throughput…
| Model | Latency (p50) | Throughput | Input / 1M | Output / 1M |
|---|---|---|---|---|
🐋 deepseek-v4-flashRecommendedvs GPT-4o: 30× cheaper | — | — | $0.15 | $0.29 |
🌐 qwen3-maxvs GPT-4o: 6.9× cheaper | — | — | $0.38 | $1.51 |
🐋 deepseek-v3.1vs GPT-4o: 5.4× cheaper | — | — | $0.60 | $1.81 |
🌙 kimi-k2.5vs GPT-4o: 3.5× cheaper | — | — | $0.63 | $3.15 |
🧩 glm-5vs GPT-4o: 3.0× cheaper | — | — | $1.05 | $3.36 |
🌙 kimi-k2.6vs GPT-4o: 2.5× cheaper | — | — | $1.00 | $4.20 |
🎵 Media models · Embeddings / ASR / TTS / Image / Video
text-embedding-3-smalltext-embedding-3-largetext-embedding-ada-002whisper-1tts-1tts-1-hddall-e-3dall-e-3-hdsoraveo-2vidu-2.0vidu-q1-t2v