Tokens are cheap right now. Build as if they will not stay that way.
The pricing bet
Current token prices are a promotional moment, not a plateau. Providers are buying market share, and everything underneath them pushes the other way: compute is expensive, competition will not stay this generous forever, and each model generation is more complex to run than the last. I would not bet an architecture on today’s rate card holding.
A price rise should be a line item, not a rewrite
If your system only pencils out at current rates, a price change forces a redesign, and it forces it under pressure on someone else’s timeline. Build tight now and the same increase is a number you adjust in a budget. Nothing structural has to move.
The two dials
Every AI call has two token costs: what you send in, and what comes back. Output typically runs 3 to 5 times the price of input, depending on model and tool. You control both, but the levers are different, and they deserve separate attention.
Controlling input
Input is where the volume lives. Most of the savings on this side come from keeping a model out of the loop entirely.
Be careful when you use AI
Delegate anything repetitive and predictable to plain Python. Cron jobs make the point concrete. Best case is pure Python. Middle case is Python that calls a model through the API for the parts that genuinely need one. Worst case is Python that fires an agent, or a scheduled agent run, and that last one should be rare.
Don’t re-digest
Processing work should happen once. A large raw file gets read a single time and its output saved in smaller, structured formats that everything downstream reads instead. On a team, one person or process owns ingestion, and everyone else works from digested markdown files that follow a shared template. Sharing the digest is the whole point.
Clean with old-school tools before ingesting
A lot of file noise strips out with Python or a regex, no model involved. Converting XLSX to CSV saves a substantial share of tokens on its own. Transcripts often shed 10 to 20 percent to a small regex pass. Docx and similar formats carry metadata the model never needed. This is free savings, so take it first.
API-based ingestion instead of chat or agentic
API calls give you control over memory. You force the model to work memory-less and feed it exactly what it needs for the output you want. Chat and agentic tools carry state you can’t inspect, which makes tuning toward a predictable output much harder.
The side benefits add up. Model deprecation is less scary, because you can see every place a model is called. Switching billable accounts (client A versus client B, personal versus business) is trivial. And prompt caching is available on some models, which pays off on exactly the repetitive tasks this approach is built for.
Controlling output
Output is the expensive meter. The goal on this side is a model that thinks less and writes less.
Prompt engineering
Overused term, simple goal: get the model thinking less and outputting less. If you routinely get two or three pages back and only read a couple of paragraphs, the prompt needs work. Naming the exact output format helps. Naming an exact length is unreliable, because models take length instructions literally in ways you may not want.
API-based interaction
For repetitive tasks, the API gives you precise control over output shape. The absence of chat memory is a feature here: every call starts from the same baseline and produces the same shape of result.
Sub-agent delegation
Route different task complexities to different models. Automatic effort switching (a light model for classification, a heavier one for actual reasoning) is a substantial saving on volume workloads. Derrick Hicks has a good talk on this pattern, worth chasing down if it is new to you.
Output formatting
Naming the output type (JSON, CSV, markdown, a specific schema) shrinks output dramatically. Structured formats push the model away from wordy prose and toward exactly the fields you asked for.
Team-level savings
The biggest saving on a team is not making the same thing twice. If two teammates each summarize the same meeting, you paid for it twice and now have two slightly different summaries to reconcile. Every generation task should start with one question: what’s already written?
A shared context layer with a shared digest format turns individual work into team infrastructure. One person’s ingestion becomes everyone’s reference.
Close every session with a write-out
End a working session with a /write-out command, or something like it, that captures what you learned into the shared layer. The next session, yours or a teammate’s, starts from that instead of rediscovering it at full price.
What I’m still working on
None of this is finished. These are the three places the approach is weakest right now, and I would rather name them than let the diagrams imply otherwise.
Open problems
- Team sharing and access. Harder than expected, mostly because of permissioning. I am working through it now.
- Distilling agent flows into Python. Rebuilding a working agent flow as a script is slow and carries none of the dopamine of building something new. When the agent version already works, finding the ROI to rewrite it is a discipline problem, not a technical one.
- Indexing beyond folder structure. Folder trees work, then plateau. Topic clustering has promise. Cross-folder search and tagging is the current weakness, and a tagging system backed by a content database is the plan. It is not built yet.
The through-line
Token discipline is not about saving money today. At current rates the savings are real but modest. It is about building an architecture that survives whatever the pricing curve does next, where a price change moves a number in a budget and nothing in the system.
