You already have one of these. These are the three calls I made building mine that I would actually defend.
Most of the interesting decisions in a context layer have nothing to do with the model. They are about where you spend tokens, what machine the thing lives on, and how much autonomy you hand to an agent.
What follows is three positions, the reasoning behind each, and the parts that are honestly still unfinished. I have marked the unfinished parts, because a talk that only shows the working bits is not much use to anyone building the same thing.
Token Discipline as a Design Constraint
Token budget is not an optimization I got to after the system worked. It is the constraint the structure was built around, and it is the reason the structure looks the way it does.
Do not design around today’s rates
The per-token price of any given model tends to fall over time. My monthly bill does not. Every year I reach for stronger models, longer contexts, and more passes per item, so the spend goes up even as the unit price comes down. A pipeline that only pencils out at the cheap tier is a pipeline you rewrite the moment you move up a model class.
So the system is already tight at rates I am not paying yet. That is a deliberate margin, not premature optimization.
Pay for digestion once, not per query
Every item gets digested at ingestion into progressively smaller tiers. That costs real money, once, on the way in. Every query afterward reads the smallest tier that answers the question and only opens heavier material when it genuinely needs to.
The alternative, where the model re-reads source material on every query, has a lower setup cost and an unbounded running cost. Mine has the opposite shape, and the crossover arrives faster than you would guess.
Navigate by folder, not by dumping the tree
Each level of the tree carries a small index describing what is in it and what the subfolders are for. The AI reads that index, decides which branch is relevant, and only then opens anything deeper. Walking down four levels costs a few hundred tokens of navigation rather than a full directory listing plus a pile of speculative reads.
This is the cheapest structural win in the whole system and it required no AI at all to build.
Cross-folder taggingWork in progress
Folder trees are good at “everything about this client” and bad at “everything where pricing came up, across every client.” Right now that second question is answered by search over the index, which works but is not precise.
The plan is a real tagging system backed by a content database, with the folder tree kept as the human-readable view. That is designed and not built. If you are solving this one, I want to hear how.
A Real Server, Not a Machine You Also Use
The context layer stopped being a toy the week it moved off my working machine. Not because the hardware was better, but because the thing became available whether or not I had a laptop open.
Dedicated, always on, with a deliberate storage split
Bulk archive sits on cheap spinning disk, where capacity matters and speed does not. The pipeline working set and the index sit on SSD, where every scheduled run touches them repeatedly. That split is the only hardware decision that has measurably mattered.
Treat the box as infrastructure. A machine you also use for work is a machine that gets rebooted mid-run, sleeps during a nightly job, and has its disk filled by something unrelated.
Remote access means the context follows you
Reaching the server over SSH turns out to be the whole multi-device story. Laptop, desktop, tablet, a borrowed machine: the client is whatever is in front of me and the context is always the same context. Nothing syncs, because there is only one copy.
This is less work than it looks
A used mini PC or an aging all-in-one, plus a couple of external drives, gets you most of the way there. The expensive part of this project was never the hardware. If the server is the thing stopping you, it should not be.
The corporate version needs real permissioningWork in progress
Single user is easy: everything on the box is mine. The moment a team shares one of these, you need role-based scoping, per-user context slices, and actual auth, and you need them before anyone loads a client folder they should not see.
I am working through exactly this with a client now. The shape is clear and the implementation is not finished. Treat anything I say about multi-user as a design sketch rather than a running system.
Python Everywhere, AI on Rails
This is the decision I argue about most, so here is the blunt version: do not let an agent auto-run your pipeline. Write Python that runs on a schedule, and let that Python call an AI API at the specific points where intelligence is actually required.
The AI is a subroutine. It is not the runtime.
What you get for the extra typing
Control over what runs and when. Cost you can predict before the run instead of reconstructing after it. Prompts that live in version control and get diffed like anything else. Models you can swap without touching logic. And no black box explaining after the fact why it decided to reprocess four hundred files.
The worked example
Meeting ingest is the pipeline I have run the longest, so it is the honest one to show. A scheduled Python script pulls new recordings. Python converts the raw export into clean markdown. Then, and only then, three AI passes produce a summary, a structured brief, and a one-card gloss, each driven by its own versioned prompt file. Routing items into the right project folder and rebuilding the index are pure Python again.
Seven stages, three of which cost tokens. The other four are deterministic, debuggable, and free.
Prompts are files, not string literals
Every AI call in the pipeline loads its instructions from a markdown file on disk. Not an f-string three functions deep in a processor. This sounds like a small thing and it is not: a prompt in a file gets versioned, diffed, reviewed, and reverted. A prompt buried in code drifts, and you find out six weeks later when the output shape quietly changed.
# Meeting Gloss Prompt — v1.0
# Flag definitions live in prompts/flags/meeting_flags.md
# Last updated: 2026-06-27
## What a Gloss is
A Gloss is a catalog card. It is NOT a summary of content. It does
not narrate what was discussed. It is descriptive of THE FILE, not
THE EVENTS within the file.
Example:
GOOD: "Internal 1:1 between two team members covering team
management, service initiative planning, and deal pipeline."
BAD: "A strategic discussion where the lead approved the service
initiative and aligned on team role changes."
The first describes the file. The second describes what happened.
Glosses are the first kind.
## Accuracy requirements
- Never add a last name if only a first name was used in the source
- Never assign a title, role, or company affiliation not stated
- Never name a person, company, or product that was not in the source
- When in doubt, omit rather than invent
Prompt caching pays for the structure
Because the system prompt and the shared flag definitions are identical across every call in a run, they go up marked as cacheable. The first call writes the cache at a small premium. Every call after that reads it at roughly a tenth of the input price. Three tier passes over a day’s worth of meetings means the fixed instructions get billed properly once and discounted for the rest of the run.
This only works because the prompt is stable and reused, which is a direct consequence of it being a file rather than something assembled per call.
Multi-client work without login juggling
Different work bills to different accounts. Rather than signing in and out, the client resolves an API key profile from the environment and the scheduled run passes whichever profile matches the context it is operating in. Adding an account is an environment variable, not a code change.
# Profiles let one pipeline bill several accounts without
# signing in and out. The scheduled run passes whichever
# profile matches the context it is operating in.
import os
import anthropic
def resolve_key(profile: str = "default") -> str | None:
"""Return the API key for a profile, falling back to the default."""
if profile == "default":
return os.getenv("ANTHROPIC_API_KEY")
return (
os.getenv(f"ANTHROPIC_API_KEY_{profile.upper()}")
or os.getenv("ANTHROPIC_API_KEY")
)
class AIClient:
def __init__(self, model="claude-sonnet-4-6", key_profile="default"):
api_key = resolve_key(key_profile)
if not api_key:
raise EnvironmentError(f"No API key for profile {key_profile!r}")
self.model = model
self._client = anthropic.Anthropic(api_key=api_key)
What a processor actually looks like
Load the prompt. Read the input tier. Call the API. Write the output tier. That is the entire shape, and it is the same shape for all three AI stages. The interesting logic, which file to process, whether it has already been done, where the result belongs, lives in Python on either side of the call.
from pathlib import Path
from shared.ai_client import AIClient
from shared.prompt_loader import load_prompt
STAGING = Path("<context-layer>/_staging/summarized")
system_prompt = load_prompt("summarize_meeting.md")
flags = load_prompt("flags/meeting_flags.md")
client = AIClient(key_profile=args.key_profile)
def process_file(stem: str, path: Path) -> dict:
cleaned = path.read_text(encoding="utf-8")
result = client.generate(
system_prompt=system_prompt,
user_content=f"# Meeting To Process\n{cleaned}",
cache_system=True, # stable across the whole run
cacheable_prefix=f"# Flag Definitions\n{flags}\n",
)
(STAGING / f"{stem}.md").write_text(result["text"], encoding="utf-8")
return {"cost_usd": result["cost_usd"], "tokens": result["input_tokens"]}
The method: let AI find the process, then make it a script
This is the part I would most want you to take away. Do not try to write the pipeline cold. Work through the task conversationally with an AI until you have a manual pass that produces output you would actually keep. You will learn what the right summary length is, which fields matter, and where the judgment calls really are.
Then have it convert that working process into a Python script, calling an AI API only at the steps that genuinely need a model. The conversation is where you discover the process. The script is where you make it repeatable, cheap, and boring.
The Through-Line
Own the glue. Pay tokens on purpose. Treat the server like infrastructure rather than a side effect of owning a laptop.
None of these three decisions is novel on its own. Taken together, they are the difference between a context layer you built and a context layer you rent.
