DeepSeek V4
Chat completions on DeepSeek V4 through one OpenAI-compatible endpoint: Flash for everyday work, Pro for hard reasoning.
POST
/v3/llm/deepseek-v4/chat/completionsTry it- Price0.01 / 0.02 ICper 1k tokens, in / out
- Served19Kcalls
- v4
- Active
Request
JSON, in the OpenAI chat-completions shape, so an OpenAI client works with the base URL changed.
Headers
apikeystringrequiredYour API key, created under API Keys in the control panel.
Body
modelstringrequireddeepseek-v4-flash or deepseek-v4-pro. Any other value returns 400, so a typo cannot bill you the dearer tier.messagesarrayrequiredThe conversation: objects with
role, one of system, user, assistant, and content.max_tokensintegerThe most tokens to generate. Default 4096.
streambooleantrue for server-sent events, a token at a time. Default false.temperaturefloat0 to 2, default 0.7. Near 0 for code and maths, 1.0 for analysis, 1.3 for conversation and translation, 1.5 for creative writing.
top_pfloatNucleus sampling, default 0.9.
Price
| Model | For | Input | Output |
|---|---|---|---|
deepseek-v4-flash | Chat, questions over documents, classification, drafting | 0.01 IC per 1K tokens | 0.02 IC per 1K tokens |
deepseek-v4-pro | Hard reasoning, multi-step agents, olympiad-level maths and code | 0.20 IC per 1K tokens | 0.40 IC per 1K tokens |
Use Flash by default: Pro costs twenty times as much per token, so send it only the requests that need the deeper reasoning.
Response
The OpenAI chat-completion object, as in the panel: the answer is
choices[0].message.content and usage holds the billed token counts. With stream: true the same arrives as data: lines, each a chunk with choices[0].delta, ending in data: [DONE].The OpenAI chat-completion object, as in the panel: the answer is choices[0].message.content and usage holds the billed token counts. With stream: true the same arrives as data: lines, each a chunk with choices[0].delta, ending in data: [DONE].