LLM API Peak/Off-Peak Pricing: How to Schedule Workloads for Maximum Cost Savings
DeepSeek's peak/off-peak API pricing turns time into a real cost lever. This guide shows which LLM workloads can move, how to schedule them safely, and where a router helps without promising automatic cheapest routing.
A 30-second answer: peak/off-peak LLM API pricing is worth acting on when the work can wait, be retried, and be reconciled later. DeepSeek now publishes peak windows of 01:00-04:00 and 06:00-10:00 UTC, with off-peak rates at half of peak rates and weekend off-peak billing under its August 23 Beijing-time rule. Move offline evals, document enrichment, embeddings, nightly summaries, and non-urgent agent traces into cheaper windows. Keep live chat, customer-visible tool calls, payment decisions, and incident workflows on latency-first routes.
The important distinction is ownership. A routing layer can help you select configured providers and model fallbacks through an OpenAI-compatible path. It should not be treated as magic time-of-day arbitrage unless that exact feature is verified in the live product path. In practice, your queue or scheduler decides when to run; TheRouter decides which configured OpenAI-compatible route receives the call.
- Swap three values, not three SDKs. Change
api_key,base_url, andmodelin the existing OpenAI client. Keep your request/response code unchanged. - Map model IDs explicitly. The target provider's model id is almost never identical to the OpenAI id. Keep a single dict of
{ openai_id: target_id }outside business logic. - Verify streaming format. SSE chunks must follow the OpenAI
data: {...}+data: [DONE]contract. Test one streaming call before moving production traffic. - Check rate-limit headers. Some providers omit
x-ratelimit-*headers. Add a wrapper that defaults safely when headers are absent. - Keep a rollback path. Ship the swap behind a feature flag, run both endpoints in shadow for 24 hours, then cut over.
What changed with DeepSeek pricing
DeepSeek's August 13 V4-Pro GA note announced flexible reasoning effort, native OpenAI Responses API support, and a pricing update that introduced peak and off-peak API rates. Its pricing page now lists DeepSeek V4-Flash, V4-Pro, and V4-Flash-Vision-Exp with off-peak prices exactly half of peak prices. The same page states that peak hours are 01:00-04:00 and 06:00-10:00 UTC, with all other hours off-peak, and that weekends in Beijing time become all-day off-peak after the August 23 billing-rule adjustment.
That makes time a first-class cost dimension. A V4-Pro output token costs $3.96 per 1M tokens during peak and $1.98 per 1M tokens off-peak. V4-Flash output moves from $1.32 to $0.66 per 1M tokens. Cache-miss input and cache-hit input follow the same 2x peak/off-peak ratio.
This is not a universal LLM market rule. OpenAI and Anthropic primarily document batch discounts for async work rather than public time-of-day price windows. OpenAI's Batch API guide says batch jobs are for requests that do not need immediate responses, with 50% lower costs and a 24-hour completion window. Anthropic's Message Batches docs describe results becoming available after completion or after 24 hours, whichever comes first, and official pricing snippets describe a 50% discount for batch processing.
Workloads that can move safely
The safest candidates have three properties: the answer is not user-visible in real time, the job can be replayed, and the output can be joined back to a stable ID.
| Workload | Move to off-peak? | Why |
|---|---|---|
| Offline eval suites | Yes | Usually compare output quality later, not during a user session |
| Document enrichment | Yes | Can run from a queue and write reviewed metadata |
| Embedding backfills | Yes | High volume, low urgency, easy to checkpoint |
| Dataset classification | Yes | Natural batch shape with row-level IDs |
| Nightly summaries | Yes | Already scheduled work |
| Customer chat | No | Waiting for a price window breaks product latency |
| Agent tool calls inside a live session | Usually no | Delays change the agent's behavior and user experience |
| Payment, safety, or abuse decisions | No | Freshness and audit timing matter more than token discount |
A useful test is simple: if delaying the request by six hours would create a customer-support ticket, do not move it.
The scheduling pattern we use
Treat off-peak execution as a lane, not a global default. A minimal production design looks like this:
- Classify the job. Add a
latency_classfield such asrealtime,deferred, orbatch_window. - Attach a deadline. A deferred job still needs a latest acceptable completion time.
- Record the pricing rule. Store the provider, timezone, peak windows, weekend rule, retrieval date, and source URL.
- Enqueue with a stable key. Use a durable job ID so retries do not duplicate writes.
- Run during the window. A worker releases jobs when the provider window is off-peak and the deadline still allows delay.
- Route and fallback. Send the actual OpenAI-compatible request through your configured provider or router path.
- Reconcile outcomes. Compare accepted outputs, retries, stale jobs, and actual billable tokens.
const deepseekOffPeak = (date: Date) => {
const utcHour = date.getUTCHours();
const isPeak = (utcHour >= 1 && utcHour < 4) || (utcHour >= 6 && utcHour < 10);
return !isPeak;
};
export function shouldRelease(job: { latencyClass: string; deadline: Date }, now = new Date()) {
if (job.latencyClass === "realtime") return true;
if (now > job.deadline) return true;
return deepseekOffPeak(now);
}
The code above is intentionally incomplete. It does not encode Beijing-time weekend billing, holidays, provider policy changes, or model-specific exceptions. In production, store these rules as data and refresh them from official pricing pages instead of burying them in application code.
Where TheRouter fits
TheRouter routes OpenAI-compatible requests through configured providers and supports provider/model routing and fallback where the live product path supports it. That is useful once the scheduler has decided a job should run.
For example, a deferred eval worker can call one stable OpenAI-compatible base_url, select a configured DeepSeek target for the off-peak lane, and keep a fallback chain for provider errors. The scheduler still owns time. The router owns the provider/model path.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.THEROUTER_API_KEY,
baseURL: "https://api.therouter.ai/v1",
});
const response = await client.chat.completions.create({
model: "deepseek/deepseek-v4-flash",
messages: [{ role: "user", content: "Classify this support ticket." }],
});
Use DeepSeek when the pricing window and model quality fit. Keep DeepSeek V4-Pro for hard reasoning or agent tasks, and DeepSeek V4-Flash for high-volume jobs where speed and unit cost matter more. For adjacent cost patterns, compare this guide with our LLM batch processing guide, LLM API cost optimization guide, and DeepSeek V4-Pro GA pricing guide.
Failure modes to design for
The first failure mode is stale pricing. Provider pricing pages change, and a scheduler that assumes last month's windows can become quietly expensive. Put the retrieved source and timestamp next to the rule, then alert when it has not been refreshed.
The second is deadline inversion. If every deferred job waits for the same cheap window, the queue can burst into rate limits. DeepSeek lists concurrency limits of 500 for V4-Pro and 2,500 for V4-Flash. Treat those as capacity constraints, not a promise that every queued job will clear instantly.
The third is unsafe retry. A classification row can retry. A customer email send, billing action, or database mutation needs idempotency and review before replay.
The fourth is quality drift. Moving a job to a cheaper model or a cheaper window should not change acceptance criteria. Track cost per accepted output, not only cost per token.
Decision matrix
| If your workload is... | Use this pattern |
|---|---|
| Live and user-visible | Route for latency and reliability; ignore peak/off-peak discounts |
| Offline and high-volume | Queue for off-peak, use row IDs, reconcile outputs |
| Offline but deadline-bound | Wait for off-peak until the deadline threshold, then run immediately |
| Quality-sensitive | Keep the stronger model and use off-peak timing before downgrading model quality |
| Provider-risk-sensitive | Sample alternate providers and keep fallback ready; do not assume time discount fixes outages |
The best savings usually come from combining levers carefully: off-peak timing for movable DeepSeek work, provider batch APIs for true async bulk jobs, prompt caching for repeated prefixes, and routing/fallback for reliability boundaries.
Sources
- DeepSeek V4-Pro GA release, retrieved 2026-08-22: https://api-docs.deepseek.com/news/news260813/
- DeepSeek models and pricing, retrieved 2026-08-22: https://api-docs.deepseek.com/quick_start/pricing/
- OpenAI Batch API guide, retrieved 2026-08-22: https://developers.openai.com/api/docs/guides/batch
- Anthropic batch processing docs, searched 2026-08-22: https://docs.anthropic.com/en/docs/build-with-claude/batch-processing
- Anthropic pricing docs, searched 2026-08-22: https://platform.claude.com/docs/en/about-claude/pricing