Model catalog

30+ models, one page

Real-time latency, throughput and price — filter by region, sort by any column.

Models

7-day requests

Updated

30+ models — latency, throughput and price at a glance

Official full-precision weights, no distillation. Prices are our wholesale cost (+5% to cover payment fees).

Prices are reference cost values, loading latency/throughput…
ModelLatency (p50)ThroughputInput / 1MOutput / 1M
🐋deepseek-v4-flashRecommended

vs GPT-4o: 30× cheaper

$0.15$0.29
🌐qwen3-max

vs GPT-4o: 6.9× cheaper

$0.38$1.51
🐋deepseek-v3.1

vs GPT-4o: 5.4× cheaper

$0.60$1.81
🌙kimi-k2.5

vs GPT-4o: 3.5× cheaper

$0.63$3.15
🧩glm-5

vs GPT-4o: 3.0× cheaper

$1.05$3.36
🌙kimi-k2.6

vs GPT-4o: 2.5× cheaper

$1.00$4.20

🎵 Media models  ·  Embeddings / ASR / TTS / Image / Video

text-embedding-3-smalltext-embedding-3-largetext-embedding-ada-002whisper-1tts-1tts-1-hddall-e-3dall-e-3-hdsoraveo-2vidu-2.0vidu-q1-t2v