Token-based pricing, at our cost
Prices are our wholesale cost — just +5% to cover payment fees — below each provider's public rate card. Per 1M tokens. Pay with any credit card via Stripe.
🇨🇳 Chinese value models
Production-grade, 1/10 – 1/100 the cost of GPT-4o
| Model | Input / 1M | Output / 1M |
|---|---|---|
deepseek-v3.120% off · full≈ gpt-4o quality at 1/10 the price | $0.22 | $0.88 |
deepseek-v4-flash20% off · fullDeepSeek ultra-cheap | $0.11 | $0.22 |
kimi-k2.5full modelMoonshot, long context | $0.60 | $2.99 |
kimi-k2.6full modelKimi latest flagship | $0.95 | $3.99 |
sonarfull modelPerplexity real-time search | $0.25 | $0.25 |
🌎 Western frontier
Full access, no cross-border hassle
| Model | Input / 1M | Output / 1M |
|---|---|---|
gpt-5.5~10% off · fullOpenAI flagship | $4.46 | $26.78 |
gpt-4.1~10% off · fullOpenAI strong | $1.79 | $7.14 |
gpt-4o~10% off · fullOpenAI balanced | $2.23 | $8.93 |
claude-opus-4-7full modelAnthropic flagship | $4.99 | $74.81 |
claude-sonnet-4-6full modelAnthropic general | $2.99 | $14.96 |
claude-haiku-4-5-20251001full modelAnthropic budget | $1.00 | $4.99 |
gemini-2.5-profull modelGoogle flagship | $1.25 | $9.98 |
gemini-2.5-flashfull modelGoogle fast/cheap | $0.30 | $2.49 |
grok-4.3~10% off · fullxAI latest | $1.12 | $2.23 |
Every model is the provider's official full version (no distillation/quantization).20% / ~10% off= genuinely below list price,full model= official full model at our cost (≈ list price).
Image · Audio · Video
Billed at our cost price. Video is async (create → poll → download); polling and downloading are free.
🖼 Image
dall-e-31024² standard
dall-e-3-hd1024² HD
🎙 Audio
whisper-1speech-to-text
tts-1text-to-speech
tts-1-hdHD TTS
🎬 Video
sora-2OpenAI 720p
veo-3.1-fastGoogle fast
veo-3.1Google flagship + audio
vidu2.0Vidu value (estimated)
viduq3-proVidu Q3 flagship 1080p (estimated)
Video is billed per generated second; seconds is 4 / 8 / 12. Rates marked "≈" (Vidu) are estimates, final billed cost. Full model list & examples in the docs.
No per-seat. No minimums. No markup.
Pay for forwarded traffic at our cost. Balance never expires.
FAQ
Are these the real, production Chinese models?+
Yes — DeepSeek v3.1, Qwen Plus, GLM 4.6, Kimi K2 are the same versions the labs deploy in their own consumer apps. No distillations, no fine-tunes.
How reliable is this?+
We run redundant contracts with multiple upstream providers. If a model goes down, requests failover transparently to an equivalent model (opt-in per tenant).
Do I need a Chinese entity, phone number, or visa?+
No. Sign up with any email. We handle provider-side compliance. Pay via Stripe in USD.
Can I still use GPT-4o / Claude / Gemini if I need them?+
Yes — same API. All models billed at our cost price (wholesale + 5%). Use the `model` parameter to pick, or set `model="auto"` and let our router choose the cheapest.
Data retention?+
Request bodies are forwarded, not persisted. Only anonymized metadata (token counts, model used, latency) is stored for billing and analytics.
How do I migrate from OpenAI?+
Change `base_url` in your OpenAI SDK to `https://promptoll.com/v1`. That's it. Your existing code keeps working. Switch `model` to a cheaper one when you're ready.