How AI cost is calculated, which model to pick for what you're doing, and the habits that keep spend under control. Built for every team — you don't need to know how to code.
Each tool shows usage in a different place. Worth a look before the rest of this guide.
Daily spend by model, a running table, and how it moved versus the previous period. This is the same screen for a personal account or a team seat.
You can see the model and the running session cost the whole time you're working, turn by turn.
Tip: set up Claude Code's status line once, and it keeps the model and cost in view without asking.
$ /statusline Set up Claude Code's status line UI C:/Projects/AIEficient | Sonnet 5 | effort:med | ctx 14% used (86% left) | ~$0.420 ▶▶ auto mode on (shift+tab to cycle)
If you're on the company account, usage lives in the workspace admin panel, not your personal profile.
These screens change often. Later in this guide: where we're headed is one shared cost dashboard for the whole company.
A token is, roughly, three-quarters of a word. Models don't bill per message — they bill per token, both what you send and what comes back.
This is a rough estimate. The real count depends on the model, but the ratio holds: a 5,000-line codebase can easily turn into tens of thousands of tokens.
Every model has two meters running. The second one is always the expensive one.
The prompt, the conversation history, attached documents, pasted code. Billed every time you send it — even the third time you send the same thing.
The reply you get back. Usually 5 to 6 times more expensive per token than input. If the model "reasons" before answering, that reasoning also bills as output.
Every new message resends the entire conversation so far. That's why a long chat gets expensive without you noticing.
Message number 20 fully reprocesses the 19 before it. Nobody flags it in the moment — you notice it at the end of the month.
Context that piles up in a long session inflates that first term every single turn. A 150-person team using a coding assistant daily can rack up a monthly bill in the thousands from this alone.
Documents and history are billed every time they're resent.
Useful for hard tasks, wasteful for easy ones.
Fractions of a cent per action, times thousands of actions, times hundreds of people.
Because it sounds better — not because the task needs it.
The useful equivalence between providers is by task tier, not by exact model name. There's no perfect one-to-one match.
Pick your area and see how its everyday tasks split across the three main tiers.
Pick the model for what the task needs, not for how important it feels. A tense reply to an unhappy customer is still just a clear, well-written reply. Figuring out why churn moved is genuinely hard, even with no deadline in sight.
"Reply to this angry customer within the hour" → still a clear, empathetic reply. Default tier.
"Check why churn ticked up 2% last quarter" → conflicting signals, real reasoning, no single obvious answer. Deep tier.
It's the best balance of quality and cost for day-to-day work.
If the result isn't convincing, it's almost always the instructions, not the model.
Only if, with a good prompt, the output still isn't good enough.
Try it: was the result good enough?
Rule: use the smallest approved model that succeeds reliably.
The model pays more attention to what's at the start and the end of your message. In the middle, attention drops.
Put the general framing at the start: role, format, rules. Right before your actual question, repeat the important part in one short line. That's where the model pays the most attention at the moment it answers.
The model can lose details from the middle, give more generic answers, and make more mistakes. On top of that, every new message costs more because it drags the whole history along. If a long session starts giving worse answers, that's a signal to start over, not to push harder.
One session per topic. If you switch subjects, open a new chat instead of continuing the same thread. If you need continuity, paste a two or three sentence summary into the new chat instead of dragging the whole history along.
It's automatic. Claude picks up relevant details from your conversations on its own — your role, recurring projects, preferences you've mentioned — and brings them into new chats without you doing anything.
You can also step in. Ask directly to remember something ("remember I lead the QA team") or to forget something ("forget the topic we discussed about X"). You can review or edit what's stored in Settings → Memory.
For anything that has to carry across sessions or days — a decision made, a naming convention, an open question — don't rely on memory alone. Close a session with a short recap ("here's what we decided: X, Y, Z") and paste it into the next chat. Memory is good at knowing who you are and what you generally work on. It's not a substitute for carrying forward the specific facts a task depends on.
Goal — use a Fast-tier model to draft or critique the prompt you're about to spend Deep-tier tokens on.
Before running a complex task on an expensive model, ask a cheap one to draft the prompt, list what context is missing, or critique your first attempt. The Fast tier is good at exactly this kind of mechanical, well-specified work — so the Deep tier's expensive tokens go toward solving the actual problem, not toward you iterating on phrasing.
Tip: on a complex task, try
/model opusplan — it plans with a stronger model, then hands execution to a lighter one. Plan
deep, run cheap.
State the goal, the format and the length in your first message. "Summarize this in three bullets, for the team channel" works better than "summarize this." Give an example if the format matters. The more specific you are the first time, the fewer rounds you need afterward.
Ask for concise, clear answers. It's the cheapest optimization there is, and it makes the output easier to act on too.
A skill stores the prompt, context and format of a task you repeat, so you're not
rebuilding it from scratch every time. Synapse already has skills ready for several common workflows, and
some are built specifically to cut token use — for example rtk, caveman or
graphify, which compress, restructure or pre-filter input before it ever reaches the model.
If you're doing something similar every week, check whether a skill already exists before writing the
prompt from zero.
Had a session that nailed a task you know you'll repeat? Call skill-creator
right there and let it turn that conversation into a reusable skill — the prompt, the context and the
format, packaged so next time it's one step instead of rebuilding it from scratch.
An MCP connector is powerful, but every tool definition, every call and every result becomes part of your context — it's billed like any other input. A CLI tool, called directly, doesn't need a full schema loaded into every request, and lets you filter what comes back before it ever reaches the model. For a simple, well-defined job, prefer a CLI over an MCP connector when both are on the table — same result, a fraction of the context. Reach for an MCP when the model genuinely needs to explore or act across a live system, not just fetch one known thing.
In Claude (chat/app): open the + menu on a
message, go to Connectors, and toggle each one on or off. Only turn on what a given task
actually needs — every enabled connector's tools are loaded into context whether you use them or not.
In Claude Code: run /mcp to see which MCP
servers are configured and connected, without leaving the terminal.
$ /mcp MCP Server Status ✓ jira-mcp connected 12 tools ✓ slack-mcp connected 8 tools ○ notion-mcp disabled ✓ incident-io-mcp connected 5 tools Use /mcp <server> to toggle a server on or off
Model choice won't sit on every individual's shoulders forever. Three pieces are already in motion.
Automatically decides which model to use for a task, so you don't have to pick it yourself every time.
Some work will run on our own or private infrastructure, some on public providers, depending on the case.
Visibility into how much is spent, and on what, by team and by use case.
This is in progress, not a finished product announcement. Until it lands, the habits in this guide are the biggest lever you personally have.
Not sure which model fits your case? Ask in the AI Champions channel. That's what it's there for.