008 / tool ·
Off-peak batch cost sheet + style-guide rewrite prompt
Price any batch AI job at peak vs off-peak in under a minute, then check the output before you ship it
1. Price the job in under a minute
Open cost_sheet.csv in Google Sheets or Excel. The formulas calculate when the file is imported. Fill in three numbers per job and read off the peak and off-peak cost.
Formula (per job):
cost = (input_cache_miss_tokens / 1,000,000 × input_miss_rate)
+ (input_cache_hit_tokens / 1,000,000 × input_hit_rate)
+ (output_tokens / 1,000,000 × output_rate)DeepSeek V4.1 Flash rates, per 1M tokens (API model name deepseek-flash, from DeepSeek's pricing page):
| Peak | Off-peak | |
|---|---|---|
| Input, cache miss | $0.30 | $0.15 |
| Input, cache hit | $0.006 | $0.003 |
| Output | $1.20 | $0.60 |
The reel's job, worked through (token counts are a demo assumption: 200,000 in with no cache hits, 100,000 out):
| Model / time | Input | Output | Total |
|---|---|---|---|
| V4.1 Flash, peak | 0.2 × $0.30 = $0.06 | 0.1 × $1.20 = $0.12 | $0.18 |
| V4.1 Flash, off-peak | 0.2 × $0.15 = $0.03 | 0.1 × $0.60 = $0.06 | $0.09 |
| GPT-6.1 Sol, standard ($2 / $10) | $0.40 | $1.00 | $1.40 |
| GPT-6.1 Sol, Batch (50% off standard) | $0.20 | $0.50 | $0.70 |
To keep the comparison fair: OpenAI also halves GPT-6.1 Sol's price for Batch and Flex requests (OpenAI announcement, SitePoint summary). Check OpenAI's pricing page before you quote it. The sheet's "Compare model" column is editable, so you can enter whatever model you actually use.
Get real token counts instead of guessing
- Rough estimate: in English, one token is roughly ¾ of a word, so tokens ≈ words ÷ 0.75. This is a rule of thumb. DeepSeek's tokenizer differs from OpenAI's, and JSON wrapping adds overhead.
- Better: run the 20-string pilot from step 3. Read the token counts from the
usageobject in each API response, or from the usage page in your DeepSeek console. Then scale up: full-job tokens ≈ pilot tokens × (total strings ÷ 20). - Enter the scaled numbers in the sheet. The pilot gives you a token count and a quality check from a single run.
Cut the input bill further with cache hits
If every request starts with the same style guide, that repeated prefix can bill at the cache-hit rate ($0.003 off-peak, against $0.15 for a miss). Put the style guide first and the strings last, and keep the style-guide text byte-for-byte identical between requests. Check the cache-hit and cache-miss counts in usage to confirm it's working.
2. Know when off-peak actually is
DeepSeek defines peak hours as 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays. Every other hour, including all weekend hours, is off-peak (source).
"Off-peak" doesn't mean night where you live. Here are the peak windows in local time for mid-October 2026:
| Your zone | Peak window 1 | Peak window 2 |
|---|---|---|
| UTC | 01:00–04:00 | 06:00–10:00 |
| India (IST, UTC+5:30) | 06:30–09:30 | 11:30–15:30 |
| London (BST, UTC+1, until 25 Oct) | 02:00–05:00 | 07:00–11:00 |
| Berlin/Paris (CEST, UTC+2, until 25 Oct) | 03:00–06:00 | 08:00–12:00 |
| New York (EDT, UTC−4, until 1 Nov) | 21:00–00:00 (previous day) | 02:00–06:00 |
| San Francisco (PDT, UTC−7, until 1 Nov) | 18:00–21:00 (previous day) | 23:00–03:00 |
What this means in practice:
- US teams: "after dark" is peak for you. Your evening overlaps both windows, and Sunday evening in the US is Monday in UTC, so it's peak too. Your working day (roughly 06:00–18:00 PT) is off-peak.
- India and Europe: start batches after 15:30 IST or after 12:00 CEST and they run off-peak until the next morning.
- Safest daily slot: 10:15 UTC to 00:45 UTC (about 14.5 hours) on weekdays, or any time on Saturday and Sunday in UTC.
- Unknown: DeepSeek's page doesn't say whether a request is priced by its start time or its finish time. Leave a 15-minute buffer at each edge of a window and check the first bill.
- Clocks change on 25 Oct (Europe) and 1 Nov (US). After that, your local times for the windows shift by an hour. UTC doesn't change.
3. Rewrite the strings, then check them
The reel's test sends the same strings through the model and checks each rewrite against three style rules. Both prompts are below and in prompts.txt.
Prompt A: batch rewrite to a style guide
You rewrite UI strings to match a style guide. Return JSON only.
STYLE GUIDE
[STYLE_RULES]
HARD CONSTRAINTS
- Keep every placeholder exactly as written, including braces: {count}, {name}, %s, {{var}}.
- Never exceed the max_chars value given for each string.
- Do not add information the source string doesn't contain.
- Product and feature names stay exactly as written: [PROTECTED_TERMS]
- If a string can't meet every rule without losing meaning, keep your best rewrite and explain in "flag".
OUTPUT FORMAT (json)
{"results": [
{"id": "<same id as input>",
"rewrite": "<new string>",
"rules_applied": ["<rule number>", ...],
"flag": null or "<one-sentence reason>"}
]}
STRINGS (json)
[STRINGS_JSON]Example inputs
[STYLE_RULES]:
1. Sentence case: capitalise only the first word and proper nouns.
2. No "please", "oops", "click here" or exclamation marks.
3. Lead buttons with a verb. Address the user as "you".[PROTECTED_TERMS]: [Your product name], Workspace, Pro plan
[STRINGS_JSON]:
[
{"id": "btn_save", "text": "Click here to save your changes", "max_chars": 20},
{"id": "err_save", "text": "Oops! Something went wrong, please try again.", "max_chars": 40},
{"id": "badge_notif", "text": "You have {count} New Notifications", "max_chars": 40}
]Example output (illustrative, not a recorded result):
{"results": [
{"id": "btn_save", "rewrite": "Save changes", "rules_applied": ["2","3"], "flag": null},
{"id": "err_save", "rewrite": "Couldn't save. Try again.", "rules_applied": ["2"], "flag": "Source didn't say what failed; assumed a save action from context."},
{"id": "badge_notif", "rewrite": "You have {count} new notifications", "rules_applied": ["1"], "flag": null}
]}Settings when calling the API:
- Base URL:
https://api.deepseek.com(OpenAI-compatible) - Model:
deepseek-flash - JSON output is supported. The prompt above contains the word "json", which JSON modes usually require.
- 25–50 strings per request keeps each response short enough to check, and a failed request only loses one chunk.
The err_save flag shows the prompt doing its job. The model admits it guessed, so a person checks that row before it ships.
Prompt B: second-pass rule check
Run this in a separate request, and ideally with a different model, so the rewriter isn't grading its own work.
You are checking UI string rewrites against a style guide. Do not rewrite anything.
STYLE GUIDE
[STYLE_RULES]
For each pair, return one row per rule: PASS or FAIL, with the exact offending word or character for every FAIL.
Also check: placeholders identical to source (PASS/FAIL), meaning unchanged (PASS/FAIL + reason).
OUTPUT FORMAT (json)
{"checks": [{"id": "...", "rule_1": "PASS|FAIL: ...", "rule_2": "...", "rule_3": "...",
"placeholders": "PASS|FAIL: ...", "meaning": "PASS|FAIL: ..."}]}
PAIRS (json)
[PAIRS_JSON] // [{"id": "...", "source": "...", "rewrite": "..."}]The 20-row human check
Open style_check.csv. Paste in 20 source strings picked at random (not the first 20) and their rewrites. Formula columns flag:
- Length OK:
LEN(rewrite) <= max_chars - Placeholders kept: the number of
{in the source matches the number in the rewrite - Banned words free: no "please", "oops" or "click here"
- Starts capitalised: a basic sentence-case check. It can't catch Title Case Mid-String, which is why the human column exists.
- Human: meaning kept (Y/N): you read the row
Is the batch good enough? Before running the full job off-peak, you want 20 of 20 on placeholders, at least 18 of 20 overall, and every flagged row read by a person. A broken {count} is a production bug, so one placeholder failure means you fix the prompt and rerun the pilot. Don't treat it as a style miss.
4. Schedule it
Decision rule: can this job wait until tomorrow morning without blocking anyone?
| Live (pay peak if needed) | Off-peak queue | Never send to a hosted API |
|---|---|---|
| Chat while prototyping, a one-off string in a review, anything blocking a teammate | Alt text passes, microcopy audits, design-token renames, translation-memory cleanup, changelog drafts | Unreleased product names, client work under NDA, anything your contract keeps in-house |
The third column is the honest limit. Your strings go to DeepSeek's servers unless you run the MIT-licensed open weights yourself (Hugging Face: deepseek-ai/DeepSeek-V4.1-Flash lists it as the base model).
Cron (Linux, cronie): start at 10:15 UTC on weekdays and 09:00 UTC on weekends.
CRON_TZ=UTC
15 10 * * 1-5 /usr/bin/python3 /path/to/run_batch.py
0 9 * * 6,0 /usr/bin/python3 /path/to/run_batch.pyGitHub Actions schedules always run in UTC:
on:
schedule:
- cron: "15 10 * * 1-5"Peak guard (Python sketch, untested; run it on your 20-string pilot first). It checks the clock before every request and waits if a peak window is close:
import os, time, json, datetime as dt
from openai import OpenAI
client = OpenAI(api_key=os.environ["DEEPSEEK_API_KEY"], base_url="https://api.deepseek.com")
PEAK_UTC = [(1, 4), (6, 10)] # hours, Mon–Fri
BUFFER = dt.timedelta(minutes=15)
def in_peak(t):
return t.weekday() < 5 and any(a <= t.hour < b for a, b in PEAK_UTC)
def safe_now():
now = dt.datetime.now(dt.timezone.utc)
return not (in_peak(now) or in_peak(now + BUFFER))
def run_chunk(prompt):
while not safe_now():
time.sleep(300) # check again in 5 minutes
r = client.chat.completions.create(
model="deepseek-flash",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"},
)
print(r.usage) # log tokens for the cost sheet
return json.loads(r.choices[0].message.content)Chinese public holidays are off-peak too. The guard ignores them and treats those days as normal weekdays, so at worst it waits when it didn't need to.
5. When it goes wrong
| Symptom | Most likely cause | Fix |
|---|---|---|
| Bill shows peak rates | The job ran from local time, or a US Sunday-evening start was already Monday in UTC | Set CRON_TZ=UTC, use the guard, check the local-time table |
| Bill is about 2× the sheet's estimate | Output was longer than the pilot suggested (JSON keys, flags, rules_applied) | Scale from the pilot's measured usage, not from word counts |
| Few or no cache hits | The style guide changes between requests, or the strings come before it | Keep the style guide first and byte-identical |
{count} became {Count} or disappeared | The model "fixed" casing inside the placeholder | Keep the hard constraint in Prompt A, and stop if the placeholder column isn't 20/20 |
| Rewrites look fine but read badly in the UI | Character limits were checked, but rendered width wasn't | Paste five rewrites into the actual Figma component before approving the batch |
Sources: DeepSeek models & pricing · TechBriefly on the V4.1 Flash launch, 11 Sep 2026 · NVIDIA NVFP4 checkpoint card (MIT, base model id) · Introducing GPT-6.1 Sol · SitePoint: GPT-6.1 Sol pricing