Your AI feature is one bad month away from killing your gross margin.
The Token Economics Workbook is the planning-side discipline most teams discover three months too late. A forecasting calculator, a model routing matrix, cache-hit patterns, a gross margin protection playbook, and 15 production teardowns with cost-per-invocation math. One afternoon to install, every month afterward to compound.
Illustrative figures from a representative teardown. The workbook ships with 15 anonymized teardowns and the calculator to model your own numbers.
The prototype costs $40 a month. The production feature costs $14,000.
Almost every team ships their first AI feature the same way: the most capable model on the menu, default prompt structure, no caching, no routing logic, no per-tenant ceilings. It works. It impresses the demo. It also produces a cost-per-invocation that looks innocuous until usage compounds, and then a single Tuesday afternoon spike turns into a panicked Slack thread from the CFO.
By that point the options are bad. Rip out the feature. Raise prices. Cap usage and watch customers churn. Or stand up an emergency optimization sprint that ships a week late and breaks something else along the way. The teams that avoid that fork aren’t smarter — they just installed cost discipline before shipping, not after.
Share of input tokens in Teardown #07 that were the same on every ticket (1,200 of ~1,500), billed at full price before caching.
How long the AI line item in the worked example grew quietly before the CFO asked why.
Illustrative cost-per-invocation reduction range across the 15 teardowns. Each one names the catch that would have cost quality.
Clear about the lane. No inflated promises.
- A planning-side discipline you install before (and after) shipping.
- A forecasting calculator that lets you model unit economics in an afternoon.
- A model routing matrix and a set of cache patterns that survive price changes.
- A gross margin playbook a CFO will actually read, with a one-page margin memo template.
- 15 anonymized production teardowns, each before → change → after, with illustrative figures.
- An observability platform. Helicone, Langfuse, Vellum, and OpenLLMetry already do that well.
- A code library you install. The forecasting sheet is yours; it runs offline.
- A vendor pitch. The routing matrix names the cases where Claude loses to GPT or Gemini.
- A theoretical paper. Every teardown is a production feature's architecture moves, with the catch nobody saw coming.
- A subscription. One-time $59 with 12 months of rate-card updates.
Five deliverables. One install afternoon.
Drop in your token shape, cache hit rate, and per-model rates for up to three models from any vendor, plus users, usage and price. Output: cost per task and per user side by side, then gross margin, break-even price, price for a 70% margin, and monthly and annual spend. Live-formula .xlsx (Excel or Google Sheets) + offline HTML.
Routing by job and capability tier: when the small, fast tier (Haiku, GPT mini, Gemini Flash) clears the bar, when the largest context window wins, when the top reasoning tier is worth it, and when a small open-weight model on your own GPU is the right answer. Markdown; imports into Notion or Google Docs.
A Markdown guide with copy-paste TypeScript and Python patterns for prompt caching, response caching, semantic cache layers, and per-tenant cache isolation. With cost math on each pattern.
Written for a CFO to read in one sitting. Covers unit economics framing, the margin lever priority order, responses by situation (including pricing), the three KPIs to put on the dashboard week one, and a one-page margin memo template. Markdown.
Markdown, plus one teardown worked end to end. Each teardown: feature description, original architecture, specific changes applied, cost-per-invocation reduction, deployment context, and the catch nobody saw coming. Figures are illustrative of format and magnitude.
When Anthropic drops Haiku pricing or OpenAI launches a new tier, you get the updated rate card. Update one cell in the forecasting sheet; every projection re-flows.
What one teardown actually looks like.
Below is an abridged version of Teardown #07 from the workbook — an AI customer support triage feature at a B2B SaaS company handling roughly 50,000 tickets a month, with illustrative figures. The worked example in the workbook walks it end to end with every number shown: the per-token math before and after, the four changes in priority order, and the eval gate that held quality.
- Frontier-tier model on every ticket, default settings.
- ~1,500 input tokens (full ticket + KB context) · ~200 output tokens.
- No prompt caching. No response caching.
- No routing — the easy 80% of tickets ran the same expensive path as the hard 20%.
- Two-stage: a small-tier classifier triages the easy 80% of tickets; a mid tier only on the hard 20%.
- Prompt caching on the system prompt and KB context (~1,200 of the 1,500 input tokens).
- Response cache on a 4-week semantic window for the top recurring issue patterns.
- Output-token ceiling enforced; structured response schema prevents runaway generation.
When each model is the right answer — and when it isn’t.
The full matrix in the workbook is a routing framework by capability tier across three providers, not a price list. Below is the top-level grid you walk first. Tier names describe capability classes; map them to the current generation and put your own verified rates in the calculator.
The full matrix walks every row and pairs it with five sub-questions: output token ceiling, latency budget, structured-output requirements, eval-quality threshold, and per-tenant cost ceiling. The workbook also includes the failure-mode notes — the spots where a cheaper model looks identical on quick spot-checks and silently degrades on production traffic.
Every teardown ships with the catch.
Each row below is an anonymized production AI feature. The delta column is the cost-per-invocation reduction after applying the routing, caching, and prompt-structure changes documented in the teardown. Deltas are illustrative of format and magnitude; they depend on token volumes and on pricing at the time.
The integrity moat.
Exactly what you get for $59, and what you don’t.
- Cost-per-invocation modeling and forecasting.
- Model selection, routing, and provider trade-off analysis.
- Prompt caching, response caching, and semantic cache patterns.
- Output-token discipline and structured-response patterns.
- Gross margin framing for non-engineering stakeholders.
- Latency engineering. Important, separate problem.
- Eval framework design. Covered in the Prompt Evaluation & Versioning System ($49).
- Agent orchestration patterns. Covered in the Agent Orchestration Cookbook ($79).
- Building observability infrastructure. Use Helicone, Langfuse, or OpenLLMetry.
- Tax, financial, or fundraising advice. Talk to your CPA and your board.
Model the cost per token, then watch the whole spend.
This models cost at the token level; the rest of the AI-spend stack zooms out. The AI Cost-Per-Task Calculator rolls tokens up to the price of a finished task, the AI Burn-Rate & Budget Blowout Forecaster projects that spend forward to the month you'd blow the budget, and the AI Spend Runaway & Billing-Safeguard Gate checks a leaked key or retry loop can't blow it overnight.
The questions teams actually ask before they trust the model.
Those are observability tools. They tell you what your AI feature spent AFTER you spent it — usually after a CFO has already noticed. The Token Economics Workbook is the planning-side discipline that lives upstream of observability: a forecasting calculator, a routing matrix, and cache-hit patterns you apply BEFORE you ship. The two layers are complementary. The playbook's three KPIs start as a spreadsheet cell each; once volume justifies it, an observability tool reads them for you.
All three. The routing matrix works by job and capability tier across Anthropic (Haiku / Sonnet / Opus), OpenAI (mini / standard / reasoning) and Google (Flash / Pro): when the small, fast tier clears the bar, when the largest context window wins long-context synthesis, when the top reasoning tier is worth it, and when paid hosted models lose to a small open-weight model on your own GPU. Tier names describe capability classes; you map them to the current generation. The forecasting calculator is provider-agnostic — enter your token shape and the current rates for up to three models and it shows cost per task and per user side by side.
Anonymized production teardowns, with illustrative numbers. Each teardown documents the feature, the original architecture, the specific changes applied (model routing, caching, output capping, pre-filtering, etc.), the cost-per-invocation reduction, the deployment context, and the catch nobody saw coming. The workbook states that the dollar figures and deltas are illustrative of format and magnitude, not quoted results: they depend on token volumes and on model pricing at the time of each deployment. The architecture moves are the durable part; you run your own numbers in the calculator.
A spreadsheet, plus an offline HTML version of the same calculator. The .xlsx opens in Excel and imports into Google Sheets, with live formulas and a tab that shows every formula. The decision was deliberate: a $59 product that becomes a SaaS dependency for a team's planning is a worse product, not a better one. The sheet is yours, it runs offline, and you can fork it into your own internal tooling without asking us for an export.
The pricing inputs are one input rate and one output rate per model — when Anthropic drops Haiku pricing or OpenAI launches a new tier, you update those cells and every projection in the workbook re-flows. The structural advice (when to route, what to cache, how to think about cost-per-invocation as a margin lever) outlasts any specific price point. You also get free updates to the rate cards for 12 months.
Three audiences. (1) Founders or engineering leads about to ship an AI feature, who want to model unit economics before launch instead of after. (2) Teams already in production whose AI line item just got CFO-flagged and need a 30-day plan to cut burn without breaking the feature. (3) Agencies and consultants building AI features for clients who need defensible cost models in proposals. If you're none of these, the workbook is probably overkill.
A spreadsheet, an offline HTML tool, and Markdown. The forecasting calculator ships as a live-formula .xlsx and as a self-contained HTML page that runs in your browser. The routing matrix, the cache-hit patterns (copy-paste TypeScript and Python snippets), the gross margin playbook, the 15 teardowns and the worked example are Markdown files, which import into Notion, Google Docs or Word without reformatting.
Yes. 30-day no-questions refund. If the workbook doesn't pay for itself in identified savings within 30 days of you actually opening it, email RedHub AI support and we refund. We can count on one hand the refunds requested across the catalog to date — the math tends to be self-evident once people open the forecasting sheet.
One afternoon to install.
Every month afterward to compound.
Open the forecasting sheet. Drop in your token volumes. Pick a teardown that matches your feature. The playbook turns the lever order into a 30-day plan to cut burn.
Sold by RedHub AI LLC · Secured by Stripe · redhub.ai