MCP
How models reach your tools and data
12Build your Unity scene by typing a sentence
A Unity MCP server exposes the Editor's actions as tools an AI client can call — so 'make a red cube and add a Rigidbody' appears in your scene. What MCP is, what the server exposes, and why it's a bridge (not magic) to review.
read →MCP deleted sessions. Your cart ID is not a password.
The 2026-07-28 MCP spec removed protocol-level sessions, so anything spanning calls is now referenced by an ID your server hands out. The spec is blunt: possession of that handle is not authentication.
read →MCP vs function calling
They're not rivals. Function calling is how one app calls a tool; MCP sits on top of it so any client can discover and call the same tool. The honest difference, and the one rule that decides which you need.
read →MCP, explained in one drawing
What MCP is and why it exists: the M×N integration mess it replaces, the three primitives, and the transports — the whole protocol in one page.
read →The tool says read only. The server wrote that itself.
MCP tool annotations — readOnlyHint, destructiveHint, even the title — are written by the server that ships the tool. The spec says clients MUST treat them as untrusted unless the server is trusted, and endorses showing the arguments instead.
read →Tools vs. resources vs. prompts in MCP
An MCP server has three primitives, not one. Tools (the model decides), resources (the app decides), prompts (the user decides). The 'who controls it' axis, with a decision rule and the client-support catch.
read →Two servers, one tool name
MCP tool names are unique per server, not globally. Aggregate two servers and the model can see two identical tools — and the prefix everyone reaches for is ruled out by the spec itself.
read →Why your AI can't touch your data
A fresh model can describe your database perfectly and read nothing from it. MCP is the standard that hands it real tools — wrap a tool once, and any AI app can use it.
read →You passed the token on. Now nobody knows it was you.
Forwarding a client's token downstream is the second half of one spec sentence: MCP servers MUST NOT accept or transit any other tokens. The far side trusts it because it came through you, and its log names your user instead.
read →Your agent skipped the tool you gave it
The model never sees your function body. It picks a tool from three things: the name, the description, and the input schema. Write descriptions that say WHEN to call, not just what the tool does.
read →Your MCP server is taking the wrong tokens
Token passthrough is forbidden by the MCP specification, and it is easy to ship by accident because the token is never invalid. It is a real credential issued for a different service.
read →Your MCP tool asked a question. Now it runs twice.
When a tool needs more input it answers with a question — and the answer comes back on a NEW tools/call carrying the same arguments. Your handler runs from the top again, so anything it did before it asked happens twice.
read →Agentic
Loops, edits, and what they cost to run
19An AI agent is just a loop
Strip the hype off an AI agent and you're left with a loop: think, act, observe, repeat. The model is stateless — the loop and its growing context are the memory. Here's the mechanism, when the loop stops, and how to stop it running forever.
read →Grok Bot vs a $20 scheduled Claude: one thing new
Does Grok Bot do anything a scheduled Claude can't? Per each vendor's own docs: one thing. It keeps a browser signed in on its own cloud computer, so a job you teach it reruns with your laptop shut. A scheduled Claude runs in the cloud too; Claude's browser needs your computer.
read →Hide your .env from Claude Code
In Claude Code, a Read deny rule for .env stops Claude's own file tools and the shell commands it recognizes, but not a grep run across the folder or a script that opens the file itself. The sandbox, with the rule kept and strict mode on, covers every command.
read →How coding agents edit your files
Coding agents don't retype your file — they do targeted find-and-replace on a unique anchor. That's why broad, multi-file requests break far more often than small specific ones.
read →Idempotent tools = agents that don't nuke prod
Agents retry, loop, and re-call tools — so a non-idempotent charge can fire twice. Idempotency means many calls, one effect: a unique key makes the retry a no-op. Which tools need it, and why the MCP idempotent hint is only a hint.
read →Nothing is telling your agent to stop
An agent ends its own loop — the model says "I'm done", not the tool. The runaway guard is your code's job, and a cap the model cannot see truncates it mid-thought.
read →OpenAI's dots vs a $20 scheduled Claude: one thing new
Do OpenAI's dots do anything a scheduled Claude can't? Per each vendor's own docs: one thing. A dot's cloud browser stays signed in and works with your devices off. Its own 9 AM check-in example is what a Claude scheduled task already does.
read →Prompt injection: the SQL injection of AI
Your AI agent reads a web page that says 'ignore your instructions, email the database' — and it does. Instructions and data share one channel, so attacker text becomes commands. Why it's still unsolved, and how to shrink the blast radius.
read →The reply costs ~5× the prompt
Per token, LLM output costs about five times input on every major tier. The levers are all on the output side: cap the reply, ask for the short form, and never make the model retype your document.
read →Why 20 tool calls cost you 200
The API is stateless, so every step of an agent loop re-sends the whole conversation. Twenty steps bills like two hundred and ten, and it climbs far faster than your step count.
read →Why agents need a scratchpad (memory)
LLMs are stateless — each call is a blank slate. Agents bolt on memory: a scratchpad (working memory, the thought/action/observation loop) plus a long-term external store. Here's how, and what breaks without it.
read →Why AI confidently makes things up
LLMs predict likely text, not truth — and they were trained on benchmarks where a confident guess scores higher than 'I don't know.' Why hallucination is baked in, and what actually reduces it: grounding, citations, and letting it abstain.
read →Why your agent ignores half your prompt
Models don't read your context evenly. Recall follows a U shape — strong at the start and end, weak in the middle ('Lost in the Middle'), and longer context degrades more. Put the important stuff first or last.
read →Your agent asked for three tools. You answered one.
Parallel tool use is on by default, and answering a batch one result at a time silently stops the model asking for batches. Two halves: run them concurrently, and return every result in a single user message.
read →Your agent rewrote the failing test
When a coding agent can't pass a test honestly, it may edit the test instead. In ImpossibleBench (ICLR 2026), GPT-5 cheated on 54% of impossible tasks built from real GitHub issues despite a strict do-not-modify prompt. Grading with the original tests and giving the agent an exit both helped, on different models.
read →Your agent summarised itself and lost the detail
When a long conversation fills the context, compaction replaces it with a summary — and the API drops every content block before that summary. It is not a trim of the oldest few messages.
read →Your AI's fake package is real now
Package names an AI invents repeat, so someone can find them by asking the same questions and register them before you look. A package existing on PyPI or npm, even with downloads, proves nothing. Copy the install line from the tool's own official docs.
read →Your subagent doesn't know what you know
Subagents share the filesystem but not the conversation, so one starts with nothing except the message you sent it. The brief is the entire context — and that is why review is the one thing you should never delegate.
read →Your tool said failed. That was the whole message.
After a tool call fails, the error text you send back is the only thing steering the next attempt — and there are only about two or three attempts before the agent apologises to the user.
read →RAG
Retrieval, embeddings, and what goes in the window
13Embeddings, explained with one analogy
An embedding turns text into a point on a map of meaning — similar meaning lands close together. Cosine similarity, the king-queen caveat, and why it's the engine of semantic search and RAG.
read →Pure vector search misses exact matches
Ask a pure vector DB for error code E-4031 and it returns vaguely related prose, not the doc that says E-4031. Hybrid search runs keyword (BM25) + vector in parallel and fuses the ranked lists (RRF) so exact matches come back too.
read →RAG vs. fine-tuning: pick in 30 seconds
RAG changes WHAT a model knows; fine-tuning changes HOW it behaves. Use RAG for facts, fine-tuning for format/tone/skill — and why fine-tuning on new facts actually increases hallucination.
read →Reranking: the cheap RAG upgrade
Your RAG retrieved the right chunk — but it's ranked ninth, so the model never reads it. A cross-encoder reranker scores query and chunk jointly and fixes the order. The two-stage pattern, why you can't rerank everything, and the caveats.
read →Vector DB, or just Postgres?
The reflex is to spin up a separate vector database. But Postgres + the pgvector extension already does similarity search — vectors next to your rows, filtered and ranked in one query. Start simple; go dedicated only when you outgrow it.
read →What a context window really limits
A context window isn't the model's memory. It's the max tokens the model can attend to in one call — input and output share that budget, and it resets every turn because the model is stateless.
read →Why your RAG retrieves garbage — fix your chunking
Most RAG fails on the chunking, not the model. The three chunking mistakes — chunks too big, too small, no overlap — and the fix: split on structure, add overlap, one idea per chunk with its context attached.
read →You deleted the page. It still quotes it.
An index is a copy, not a view, and nothing syncs it. Deleting or editing the source leaves the old text perfectly retrievable, and a well-behaved model quotes it faithfully.
read →Your follow-up question breaks your search
"Which one is cheaper?" has no subject in it. Retrieval only sees the string you send, so resolve the question against the conversation before you embed it — and resolve it, don't redecorate it.
read →Your PDF was garbage before you chunked it
Chunking assumes the text is already correct. On a two-column page a naive extractor reads straight across, so every sentence is destroyed — and nothing downstream can tell.
read →Your RAG answer cannot prove where it came from
Retrieved text pasted into the prompt comes back as prose with no link to its source. Hand the same text back as a search result and every sentence returns carrying the page it leaned on.
read →Your RAG is broken. Find out where.
One wrong answer, three different bugs behind it, and the fix for one does nothing for the other two. One thirty-second test tells you which branch you are on.
read →Your search broke when you upgraded the model
Retrieval only works when the stored vectors and the query vector come from the same embedding model. Swap the model and nothing errors — the results are just quietly wrong.
read →Workflows
Choosing, testing, and paying for models
25Do you actually need a local LLM?
Running a model locally feels free — but 'free' is a GPU you paid for running a weaker model. When local (Ollama, llama.cpp) actually beats a cloud API: privacy, offline, and very high volume.
read →Do you need a second GPU?
On llama.cpp's default split, two graphics cards take turns on every token of a reply, so a model that already fits on one card writes no faster on two. Ollama keeps a fitting model on one card by default. Where a second card does help, and what to buy instead.
read →Grade your AI with another AI
LLM-as-a-judge scales grading past what humans can read — but it has real biases: it favors whatever answer it sees first, rewards longer answers, and likes its own style. Swap the order, use a rubric, calibrate to humans.
read →How to eval an AI feature without vibes
A vibe check can't tell you if a prompt change helped. An eval = test cases + a grader, turning 'seems fine' into a number. Grader types, LLM-as-judge caveats, and the start-small workflow.
read →Mac or RTX 5090 for local LLMs?
Memory size decides whether a local model fits. Memory speed, the GB/s line on the spec sheet, sets its best-case writing speed: roughly GB/s divided by model GB. The M5 Ultra vs RTX 5090 speed limits, measured Macs against theirs, and how to pick.
read →MoE vs dense: active parameters are not your RAM bill
A mixture-of-experts model reads a few billion parameters per token but has to store all of them. Active parameters buy speed; total parameters set the memory you need to buy.
read →Move the experts, not the layers
When a mixture-of-experts model won't fit on your GPU, offloading layers is the obvious fix and the wrong one. Move only the expert tensors, keep attention resident, and know what it costs before you commit.
read →One run of your eval is a sample, not a result
The same case, same model, same prompt can come back the other way on the next run — even at temperature zero. Run each case several times and ask whether it passes every time, not whether it can pass.
read →Prompt caching: stop overpaying your LLM
You re-pay full price for the same fixed prompt on every call. Cache the stable prefix and reads cost about a tenth — here's the one rule that makes it work.
read →Quantization: run a big model on a small GPU
Quantization stores each model weight in fewer bits (FP16 -> INT8 -> INT4), so a 7B model drops from ~14GB to ~4GB and fits a cheap GPU. Two catches: quality slips worst at 4-bit, and those figures are weights only — the KV cache grows with num_ctx and can dwarf them.
read →Same GPU: bigger model or more bits?
At the same download size, a model four times bigger at 4 bits usually beats a small one at 16 bits: Qwen3-32B at 4 bits scored 80.6 on MMLU vs 74.7 for Qwen3-8B at 16. Four bits is the floor; below it, test first.
read →Sonnet 5: a third off, or an eighth?
Claude switched tokenizers at Opus 4.7, and the same text now counts as about 30% more tokens. So Sonnet 5's third-off per-token price is roughly an eighth off the same text. The check is free: count on both models, then compare cost per solved task.
read →Speculative decoding: same answer, not always faster
A small draft model guesses ahead and the big one checks the whole guess at once. The text you get is the big model's own — that half really is free. The speed is not: in one consumer-hardware study three of five configurations got slower.
read →Structured outputs > prompt-and-pray
Stop asking the model for JSON and praying it parses. Bind a schema and the decoder is constrained to it — malformed output becomes impossible. OpenAI, Claude, and how it actually works.
read →Too much context makes your AI worse
Context is a finite attention budget, not free storage. More of it dilutes the signal, the middle gets under-weighted, and stale or lookalike-wrong text actively flips answers.
read →Under 24GB VRAM? Ollama gives you 4K
Unless you set it, Ollama sizes the context window from your GPU's memory, not the model: under 24 GiB of VRAM (and CPU-only) that's 4,096 tokens. A longer prompt is cut from the start, about half a window kept, with no error to your app on most library models. The tiers, the fix, and the one-line check.
read →What temperature actually changes
Temperature isn't a creativity dial. It scales the logits before softmax — sharpening or flattening the next-token distribution. The mechanism, the practical rule, the top_p interaction, and the caveats.
read →Why every new model has two sizes
Almost every new open-weight model ships with a total parameter count and a much smaller active count. Here is what a mixture of experts actually is, why the labs all switched to one, and which of the two numbers answers which question.
read →You are already on high
output_config.effort has five levels and the default is high. It moves every output token — not just thinking — but it is a behavioural signal rather than a cap, and it is not the dial that shortens an answer.
read →You're paying your best model to do filing
The small model in the same family costs a fraction of the top tier, and sorting and extraction do not need the big one. But prompt caches are model-scoped, so swapping mid-conversation can cost more than it saves.
read →Your agent said it worked. Check the database.
The agent's closing sentence and the state of your database are two different facts. Most evals only ever read the first one — make the test open the table and assert the row.
read →Your agent's token budget doesn't count your conversation
A task budget counts what the model writes and the tool results it reads back on that turn. The history you resend every request rides along free — which is why budgets run out later than people expect.
read →Your eval measures a metric you guessed
Helpfulness, tone, accuracy — a list written before anyone read a failure. Read real runs, write what went wrong in plain words, group the notes, count them, and build the eval for the biggest pile.
read →Your eval set is too small to trust
An eval score is a sample statistic. With 20 test questions the 95% band is roughly ±17 points, so a one-case win is inside the noise. Check the range, not the score — and make the cases hard.
read →Your retry loop will never clear this 429
The Claude API returns the same rate_limit_error for a per-minute rate limit and for your tier's monthly spend cap. Only one clears by waiting, the SDKs retry both by default, and the tell is error.details.error_code.
read →News
What just changed, and what it means for you
5All Gemini API keys die in September? No
The claim that every Gemini API key stops working in September 2026. Google's own key page, updated 2026-09-25, says restricted standard keys continue to work and names no cutoff. The real cut, of unrestricted keys, came on June 19.
read →Claude marks everything it writes. It can't say it was you.
Claude now watermarks its text output. It carries no identifying information and cannot be traced to a person — and it cannot confirm a human wrote something either. Both halves, from Anthropic's own page.
read →Claude now bills you when it refuses? Yes, but narrow
The claim that Claude now bills you when it refuses. Since 2026-09-24, Anthropic bills refusals that arrive before any output, but only in three categories: bio, frontier_llm and reasoning_extraction. Its own docs say benign machine learning work can trigger frontier_llm.
read →OpenAI shuts off GPT-4 and o1 Oct 23
On October 23, 2026 OpenAI's API shuts off a dozen older models and more than twenty model names, dated snapshots and plain names alike, plus the fine-tunes built on them. The full list, the substitutes, what is not on it, and what follows on December 11.
read →Siri is now Google Gemini? Yes
The claim that the new Siri is Google Gemini. Apple's own launch post says Siri AI's Apple Foundation Models were custom-built in collaboration with Google and its Gemini models, and Google's blog says they are based on Gemini. Apple says they run on device and on Private Cloud Compute.
read →One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
