← All articles

LLM API Peak/Off-Peak Pricing: How to Schedule Workloads for Maximum Cost Savings

DeepSeek's peak/off-peak API pricing turns time into a real cost lever. This guide shows which LLM workloads can move, how to schedule them safely, and where a router helps without promising automatic cheapest routing.

· TheRouter

A 30-second answer: peak/off-peak LLM API pricing is worth acting on when the work can wait, be retried, and be reconciled later. DeepSeek now publishes peak windows of 01:00-04:00 and 06:00-10:00 UTC, with off-peak rates at half of peak rates and weekend off-peak billing under its August 23 Beijing-time rule. Move offline evals, document enrichment, embeddings, nightly summaries, and non-urgent agent traces into cheaper windows. Keep live chat, customer-visible tool calls, payment decisions, and incident workflows on latency-first routes.

The important distinction is ownership. A routing layer can help you select configured providers and model fallbacks through an OpenAI-compatible path. It should not be treated as magic time-of-day arbitrage unless that exact feature is verified in the live product path. In practice, your queue or scheduler decides when to run; TheRouter decides which configured OpenAI-compatible route receives the call.

  1. Swap three values, not three SDKs. Change api_key, base_url, and model in the existing OpenAI client. Keep your request/response code unchanged.
  2. Map model IDs explicitly. The target provider's model id is almost never identical to the OpenAI id. Keep a single dict of { openai_id: target_id } outside business logic.
  3. Verify streaming format. SSE chunks must follow the OpenAI data: {...} + data: [DONE] contract. Test one streaming call before moving production traffic.
  4. Check rate-limit headers. Some providers omit x-ratelimit-* headers. Add a wrapper that defaults safely when headers are absent.
  5. Keep a rollback path. Ship the swap behind a feature flag, run both endpoints in shadow for 24 hours, then cut over.

What changed with DeepSeek pricing

DeepSeek's August 13 V4-Pro GA note announced flexible reasoning effort, native OpenAI Responses API support, and a pricing update that introduced peak and off-peak API rates. Its pricing page now lists DeepSeek V4-Flash, V4-Pro, and V4-Flash-Vision-Exp with off-peak prices exactly half of peak prices. The same page states that peak hours are 01:00-04:00 and 06:00-10:00 UTC, with all other hours off-peak, and that weekends in Beijing time become all-day off-peak after the August 23 billing-rule adjustment.

That makes time a first-class cost dimension. A V4-Pro output token costs $3.96 per 1M tokens during peak and $1.98 per 1M tokens off-peak. V4-Flash output moves from $1.32 to $0.66 per 1M tokens. Cache-miss input and cache-hit input follow the same 2x peak/off-peak ratio.

This is not a universal LLM market rule. OpenAI and Anthropic primarily document batch discounts for async work rather than public time-of-day price windows. OpenAI's Batch API guide says batch jobs are for requests that do not need immediate responses, with 50% lower costs and a 24-hour completion window. Anthropic's Message Batches docs describe results becoming available after completion or after 24 hours, whichever comes first, and official pricing snippets describe a 50% discount for batch processing.

Workloads that can move safely

The safest candidates have three properties: the answer is not user-visible in real time, the job can be replayed, and the output can be joined back to a stable ID.

WorkloadMove to off-peak?Why
Offline eval suitesYesUsually compare output quality later, not during a user session
Document enrichmentYesCan run from a queue and write reviewed metadata
Embedding backfillsYesHigh volume, low urgency, easy to checkpoint
Dataset classificationYesNatural batch shape with row-level IDs
Nightly summariesYesAlready scheduled work
Customer chatNoWaiting for a price window breaks product latency
Agent tool calls inside a live sessionUsually noDelays change the agent's behavior and user experience
Payment, safety, or abuse decisionsNoFreshness and audit timing matter more than token discount

A useful test is simple: if delaying the request by six hours would create a customer-support ticket, do not move it.

The scheduling pattern we use

Treat off-peak execution as a lane, not a global default. A minimal production design looks like this:

  1. Classify the job. Add a latency_class field such as realtime, deferred, or batch_window.
  2. Attach a deadline. A deferred job still needs a latest acceptable completion time.
  3. Record the pricing rule. Store the provider, timezone, peak windows, weekend rule, retrieval date, and source URL.
  4. Enqueue with a stable key. Use a durable job ID so retries do not duplicate writes.
  5. Run during the window. A worker releases jobs when the provider window is off-peak and the deadline still allows delay.
  6. Route and fallback. Send the actual OpenAI-compatible request through your configured provider or router path.
  7. Reconcile outcomes. Compare accepted outputs, retries, stale jobs, and actual billable tokens.
const deepseekOffPeak = (date: Date) => {
  const utcHour = date.getUTCHours();
  const isPeak = (utcHour >= 1 && utcHour < 4) || (utcHour >= 6 && utcHour < 10);
  return !isPeak;
};

export function shouldRelease(job: { latencyClass: string; deadline: Date }, now = new Date()) {
  if (job.latencyClass === "realtime") return true;
  if (now > job.deadline) return true;
  return deepseekOffPeak(now);
}

The code above is intentionally incomplete. It does not encode Beijing-time weekend billing, holidays, provider policy changes, or model-specific exceptions. In production, store these rules as data and refresh them from official pricing pages instead of burying them in application code.

Where TheRouter fits

TheRouter routes OpenAI-compatible requests through configured providers and supports provider/model routing and fallback where the live product path supports it. That is useful once the scheduler has decided a job should run.

For example, a deferred eval worker can call one stable OpenAI-compatible base_url, select a configured DeepSeek target for the off-peak lane, and keep a fallback chain for provider errors. The scheduler still owns time. The router owns the provider/model path.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.THEROUTER_API_KEY,
  baseURL: "https://api.therouter.ai/v1",
});

const response = await client.chat.completions.create({
  model: "deepseek/deepseek-v4-flash",
  messages: [{ role: "user", content: "Classify this support ticket." }],
});

Use DeepSeek when the pricing window and model quality fit. Keep DeepSeek V4-Pro for hard reasoning or agent tasks, and DeepSeek V4-Flash for high-volume jobs where speed and unit cost matter more. For adjacent cost patterns, compare this guide with our LLM batch processing guide, LLM API cost optimization guide, and DeepSeek V4-Pro GA pricing guide.

Failure modes to design for

The first failure mode is stale pricing. Provider pricing pages change, and a scheduler that assumes last month's windows can become quietly expensive. Put the retrieved source and timestamp next to the rule, then alert when it has not been refreshed.

The second is deadline inversion. If every deferred job waits for the same cheap window, the queue can burst into rate limits. DeepSeek lists concurrency limits of 500 for V4-Pro and 2,500 for V4-Flash. Treat those as capacity constraints, not a promise that every queued job will clear instantly.

The third is unsafe retry. A classification row can retry. A customer email send, billing action, or database mutation needs idempotency and review before replay.

The fourth is quality drift. Moving a job to a cheaper model or a cheaper window should not change acceptance criteria. Track cost per accepted output, not only cost per token.

Decision matrix

If your workload is...Use this pattern
Live and user-visibleRoute for latency and reliability; ignore peak/off-peak discounts
Offline and high-volumeQueue for off-peak, use row IDs, reconcile outputs
Offline but deadline-boundWait for off-peak until the deadline threshold, then run immediately
Quality-sensitiveKeep the stronger model and use off-peak timing before downgrading model quality
Provider-risk-sensitiveSample alternate providers and keep fallback ready; do not assume time discount fixes outages

The best savings usually come from combining levers carefully: off-peak timing for movable DeepSeek work, provider batch APIs for true async bulk jobs, prompt caching for repeated prefixes, and routing/fallback for reliability boundaries.

Sources

Models covered in this article

Help & contact