Blog
- What Mixture of Experts Had to Fix Before Mixtral Worked: In 2020 Google trained a 600 billion parameter MoE model in four days, then left its trillion parameter sibling out of the paper because it kept hitting numerical trouble. Read in order, the MoE papers are a chain of problems and fixes, ending with a finding about what experts actually learn.
- The Cost Ledger of an LLM Task, Line by Line: In 2024 a Princeton team re-ran published coding agents and found that a simple retry loop was as accurate as the best of them and far cheaper than most. Here is the cost of one task built up term by term from papers, then the cascades, routers and caches that measurably cut it.
- Postmortem of a Silent Regression: What to Log and Alert On for an LLM Feature: Between March and June 2023, GPT-4 went from 84% to 51% on a prime number task and from 52% to 10% on code that runs as written, while every request still came back as normal text. Read as an incident report, the drift studies and ML monitoring papers say which signals would have caught it.
- Anatomy of One Traced LLM Request, Using Google's Dapper as the Guide: Google built Dapper because one search query touched thousands of machines and nobody could say which one was slow. An agent request has the same shape on a smaller scale. This post dissects one span tree, field by field, and then asks whether Dapper's sampling rules still hold when every span is a model call.
- Why a Passing Eval Suite Still Needs Production Logs: In a 2024 Berkeley study, developers kept rewriting their own grading rules as they read more model outputs. That finding, plus research on LLM judges, benchmark leakage, and real user traffic, makes a case that offline evaluation can only be a draft of what production observability later corrects.
- Timing an LLM Cold Start, Phase by Phase: In the ServerlessLLM paper, KServe took 128 seconds to return the first token from a cold OPT-6.7B replica, and 114 of those seconds went to downloading the weights. This post times each phase of a GPU cold start with numbers from the papers, then shows what each paper does to shrink it and when those fixes stop helping.
- Choosing an LLM Serving Stack by the Idea, Not the Benchmark: One A100 served a 13B model at 1.6 requests per second within its latency targets. Split across three GPUs by phase, the same model reached 3.3 per GPU. The frameworks differ mostly in which of three scheduling ideas they implement, and each idea helps a different workload.
- Counting Idle Slots: Where vLLM Gets Its Throughput: When the vLLM team profiled existing LLM servers, as little as 20.4% of the memory reserved for the KV cache held actual token states. Walk four requests through two schedulers by hand, then follow the memory, and the 2 to 4 times throughput gain stops looking like magic.
- Six Web-Service Assumptions an LLM Endpoint Breaks: On Azure production traces, a 1,500-token prompt on BLOOM-176B took as long as generating six output tokens. That one measurement, and a few others from Splitwise and Google's PaLM inference paper, knock over most of what ordinary web-service capacity planning assumes.
- Tracing a Wrong RAG Answer to the Stage That Broke: A Deakin University team built three RAG systems and catalogued seven ways they failed. Those failures sort into four branches, from "never retrieved" to "used but misstated", and each branch has its own metric and its own record of agreeing, or not, with human judges.
- Reranking, Read Off the MS MARCO Leaderboard: In January 2019 a BERT reranker took the top of the MS MARCO passage leaderboard with a 27% relative jump in MRR@10, and, as a later paper timed it, about 33 seconds of GPU time per query. The reranking papers that followed are a fight over that bill.
- CAG vs RAG, Read Off the Tables of Three Papers: Preloading 85k tokens of HotPotQA into a precomputed KV cache cut answer time from 92 seconds to 2.3 on Llama 3.1 8B. On that same test set, plain BM25 retrieval scored higher. A head-to-head readout of what three papers measured, and where each approach won.
- HyDE: Searching With an Answer That Might Be Wrong: In December 2022 a CMU and Waterloo team trained nothing and still nearly matched a retriever fine-tuned on MS MARCO. Their trick was to embed a made-up answer instead of the question. Here is one query followed through every step, then the results, the Query2doc variant, and where the papers say it breaks.
- Who Decides When to Retrieve? Four Retrieval Policies, From Fixed to Learned: Handed the top retrieved passages, ChatGPT got worse on a health fact-checking benchmark, falling from 70.1% to 54.7%. The fix the research tried was to let something other than the pipeline decide when to search. Here are four answers to that question, each with its trigger rule, its results, and what it costs in retrieval calls.
- From Brute Force to HNSW: How Vector Search Gets Fast: In 2017 FAISS searched a billion image vectors at 0.0133 ms per query while storing each one in 20 bytes. Getting there took three ideas, each one a fix for what the previous one still cost: inverted lists, product quantization, and layered graphs.
- One Turn of the Agent Improvement Loop, Measured: In 2023 an agent built on gpt-3.5-turbo went from 40% to 59% on ALFWorld household tasks with no weight updates, only by studying its own past runs. This post follows that loop one stage at a time (collect, label, distill, redeploy) using the papers that measured each stage.
- 61% Once, Under 25% Eight Times: Deploying Agents That Act: In 2024 the tau-bench authors ran gpt-4o on the same customer service tasks again and again. It solved about 61% on average, but fewer than 25% of tasks on all eight tries. Here is the math behind that drop, what ToolEmu measured about risky actions, and what the guardrails in these papers actually bought.
- One A2A Handoff, Message by Message, and Where Each Message Can Be Abused: A 2026 analysis found 11 attacks on the Agent2Agent protocol that work without breaking a single rule of the spec. Follow one task from Agent Card to artifact and you can see where each one lives.
- Four Ways to Wire LLM Agents, and What Each One Measured: Three copies of gpt-3.5-turbo that read and revise each other's answers solved 81.8% of an arithmetic test the same model solved 67.0% of alone. Pipelines, managers, group chats and debates each come from a paper that measured them. Here is the wiring, the numbers, and the bill.
- Splitting an Agent: An Evidence Ledger: A Berkeley team annotated 1,642 traces from seven multi-agent frameworks and found failure rates between 41% and 86.7%. Here is what the controlled studies say about when splitting a task across agents helps, when it hurts, and whether the gains survive an equal token budget.
- Who Holds the Off Switch: A Ladder of Agent Autonomy: Three papers rank AI agents by how much the human still decides. Walked rung by rung, they show what the user actually does at each level, which risk arrives with it, and why one of them argues the top rung should never be built.
- What an MCP Server Says on the Wire, and What Studies Found in Real Ones: A walk through one stdio session of the Model Context Protocol, one JSON-RPC message at a time, with the measured problems researchers found in real open-source servers placed at the exact message where each one enters.
- More Tools, Worse Picks: What the Tool Selection Papers Measured: When RAG-MCP buried one correct MCP server among thousands of distractors, success held above 90% for small pools and collapsed past about 100. Here is the curve, the reasons the papers give for it, and what retrieval, two-stage routing, and fine-tuning on API docs actually recovered.
- Who pulls the trigger: MCP tools, resources and prompts, and what breaks when the wrong party does: MCP gives each server primitive an owner: the model fires tools, the application picks resources, the user picks prompts. Here is what the specification says, what attack studies measured against each primitive, and why no paper can yet say how often real servers use the last two.
- Host, Client, Server: Where MCP Draws Its Trust Boundaries: A security benchmark pointed a web page at a local MCP server and reached it every time, on Claude Desktop, OpenAI and Cursor alike. A tour of MCP's three roles, what the spec makes each one responsible for, and what the research found breaking at each boundary.
- Function Calling, MCP and the API Underneath: A Stack Read Bottom Up: In the Berkeley Function Calling Leaderboard paper, o1 scored 91.5 on parallel calls when the tools were described in its prompt and 0.0 when they went through its native tools field. Same model, same functions. That gap only makes sense once you see function calling, MCP and a plain API as separate layers.
- Before MCP there was LSP: the M times N problem and what studies found: The argument for MCP, that M apps times N tools should become M plus N, was first made for code editors. Here is what research found when the Language Server Protocol tried it, and what ecosystem measurements show about MCP so far.
- Same Retriever, Different Driver: Fixed RAG Pipelines vs Models That Search: Search-R1 kept the retriever, the corpus and the three passages per call fixed, and changed only who writes the queries. Exact match on seven QA sets went from 0.304 to 0.431. The papers that measured the bill, and the cases where the fixed pipeline still won, tell the rest.
- Context Rot: Three Ways a Longer Prompt Loses the Answer: With 20 documents in its prompt, GPT-3.5-Turbo answered 75.8% of questions when the right one came first and 53.8% when it came tenth, below its score with no documents at all. Four papers explain why more context can make answers worse: position, sheer length, and questions that share no words with their answers.
- Context Budgeting: A Worksheet for What to Keep, Drop, or Compress: LLMLingua cut a 2,366-token math prompt to 117 tokens and GPT-3.5-Turbo lost 1.5 points of accuracy. Here is a token budget worksheet, and what three papers measured for each way of making things fit: dropping, summarizing, and deleting tokens.
- What an agent reads on every request: four memories and a window: Generative Agents lost believability each time a memory component was switched off. CoALA gives the vocabulary for why: working, episodic, semantic and procedural memory, plus the tool results that flow back in. Here is each one as the papers define it, and how MemGPT and Generative Agents build and measure it.
- Prompt Wording vs Window Contents: Where the Accuracy Moves: Changing only the separators and spacing of a prompt swung LLaMA-2-7B from 3.6% to 80.4% accuracy on one task. Wording clearly matters. The evidence on what it can and cannot fix, set against what changing the information in the window does, points to where the effort belongs.
- Valid JSON, wrong values: a failure catalog for document extraction: On 20-page FDA device reviews, a single-prompt LLM extractor skipped about 27.5% of the fields a human would fill and added values the documents never state. The papers that measured this show four distinct ways extraction goes wrong, and none of them breaks the JSON.
- Auditing a leaderboard number: what MMLU hides and what to measure instead: An audit of MMLU found errors in 57% of sampled Virology questions and an estimated 6.49% error rate overall. Add HELM's other metrics and Chatbot Arena's preference votes, and a single leaderboard score says little about which model fits your application.
- From CheckList tests to error bars: building an eval pipeline: In 2019, three commercial sentiment APIs failed nearly every sentence like "I thought the plane would be awful, but it wasn't." Build an evaluation pipeline in order: capability tests from CheckList, per-case scores, standard errors and paired comparisons from Miller, and a case count large enough to see the difference you care about.
- BloombergGPT's ledger: what training your own model cost, and what adapting one scored: Bloomberg spent a 1.3 million GPU hour budget training a 50 billion parameter finance model. Months later, GPT-4, prompted and never retrained, beat it on four of the five public benchmarks it had reported. A line by line ledger of what each path costs a team, from the papers' own figures.
- Jailbroken: the two ways safety training loses, and what defenses actually buy: In 2023 Berkeley researchers found, for every one of 32 red-team prompts, an attack that got GPT-4 and Claude v1.3 to answer it. Their two failure modes still explain automated suffix attacks, many-shot prompts, and why every defense measured since moves the numbers without closing the gap.
- The measured biases of LLM judges, one at a time: In 2023, GPT-4 judged the same 80 pairs of answers twice with only their order swapped and contradicted itself on 37 of them. A catalog of the biases papers have actually measured in LLM judges, how big each one was, which fixes were tested, and when the grades came close to human ones.
- When the tests got bigger: how thin test suites overstate code correctness: EvalPlus gave each HumanEval problem about 80 times more tests, and GPT-4's greedy pass rate fell from 88.4% to 76.2%. What the original tests missed, how pass@k is actually computed, and what changes when the task is a whole repository.
- Detection without a detector: from a closed list to a sentence: OWL-ViT found LVIS rare categories at 31.2 AP after every box for them was removed from its training data. How image-text training turns class names into vectors, how OWL-ViT and Grounding DINO turn those vectors into boxes, and where the papers say it still breaks.
- What Specializing a Model for Medicine Measurably Bought: Flan-PaLM passed a medical licensing exam benchmark, yet clinicians judged only 61.9% of its answers to match scientific consensus. Med-PaLM, Med-PaLM 2 and Medprompt show what specialization changes, and which yardstick each result was measured with.
- The price of a sentence: why models do worse outside English: The tokenizer behind ChatGPT and GPT-4 needs about 15 times as many tokens for a sentence in Shan as for the same sentence in English. What that premium is made of, what it does to cost, context and latency, how accuracy falls with a language's share of the training data, and which fixes the papers measured.
- One distribution, three dials, and a temperature zero that still wobbles: Researchers sent ChatGPT the same coding prompt five times and, on 75.76% of CodeContests problems, got five programs with no test output in common. Where that variation comes from: temperature, top-k, nucleus sampling, and the batching arithmetic that keeps temperature 0 from being deterministic.
- The most probable sentence is often empty: decoding after beam search: Widening the beam dropped a translation model's BLEU from 36.42 to 14.66, and exact search picked the empty string for most sentences. Why the most probable output is not the best one, and how locally typical sampling, contrastive search and min-p each answer that, with their measured results and the dispute over min-p.
- The KV cache and three ways to shrink it: DeepSeek-V2 reported a KV cache 93.3% smaller than its predecessor's. Why generation stores keys and values at all, how to size that store, and what multi-query, grouped-query and latent attention each gave up to make it smaller.
- BM25 and dense retrieval, two scoring functions side by side: Dense Passage Retrieval beat BM25 by 19 points of top-20 accuracy on Natural Questions. On BEIR's 18 unseen datasets, the same model averaged 47.7% worse. What each formula actually computes, term by term, why each wins where it does, and what the papers report when you fuse them.
- How small should a chunk be? What retrieval papers measured: Re-indexing Wikipedia as one-fact propositions lifted an unsupervised retriever's Recall@5 on EntityQuestions from 36.3% to 51.7% with no retraining. What the research found about chunk size, semantic boundaries, late chunking and RAPTOR's summary trees, including where each one did not pay.
- How 2,048 Tokens Became 32,768: A History of the Context Window: In 2023 a Meta team stretched LLaMA from 2,048 to 32,768 tokens with about 1,000 fine-tuning steps. The number had held because of how models encode position. This is the story of RoPE, ALiBi, position interpolation and YaRN, told through the perplexity each one measured.
- Chain of thought: the 17.9 to 56.9 result and its fine print: Eight worked examples took PaLM 540B from 17.9% to 56.9% on grade-school math. The papers that followed added four conditions: it needs scale, one sentence can trigger it, voting over many chains helps, and the written reasoning may not be why the model answered.
- Six questions to ask a leaderboard, learned from the Chatbot Arena dispute: In early 2025 researchers counted 27 private Meta variants on Chatbot Arena in the run-up to Llama 4, and showed that picking the best of several tries lifts a score. The maintainers pushed back. The argument is a good checklist for reading any benchmark claim.
- Three generations of embeddings, and what each one measured: Pairwise BERT needed about 65 hours to find the closest pair among 10,000 sentences; SBERT embeddings did it in about 5 seconds. How word2vec, Sentence-BERT and MTEB changed what a good text vector means, what the geometry also encodes, and how the scores are measured.
- Shannon's guessing game and the loss language models minimize: In 1951 Shannon had people guess English text one letter at a time and put its entropy at roughly 0.6 to 1.3 bits per letter. The same arithmetic defines cross-entropy, perplexity and KL divergence, and it is why a language model is also a file compressor.
- Test-time compute: the curve, the verifier, and the bill: Sampling DeepSeek-Coder 250 times took it from 15.9% to 56% on SWE-bench Lite. Coverage rises along a fitted curve, but picking the right sample is a separate problem, and a FLOPs-matched study shows when extra inference beats a 14 times larger model and when it loses badly.
- Structured outputs: does forcing JSON make a model worse at thinking?: Constrained decoding guarantees output that parses. A 2024 study said it also hurts reasoning, and a rebuttal said the study was measuring bad prompts. What each side measured, and the condition both end up agreeing on.
- SFT, DPO and PPO on the same model: what controlled comparisons found: Tülu 2 13B averaged 56.8 after supervised fine-tuning, 61.0 after DPO and 62.2 after PPO, and PPO took 3 days where DPO took 9 hours. The objectives written out term by term, what mattered more than the algorithm, and Anthropic's measured tension between helpful and harmless.
- Emergent abilities: what a bigger model buys, and what the metric invents: In 2022, GPT-3 looked useless at arithmetic until 13 billion parameters, then suddenly wasn't. A year later a Stanford team argued the jump came from the scoring rule, and a Microsoft team got a 28M model to write coherent stories. Read together, the three papers give a sharper answer to what size buys.
- Mapping the AI stack by what each layer discloses: In October 2023 Stanford researchers scored ten foundation model developers on 100 transparency indicators. The best scored 54, the worst 12, and the average 37. Sorted by layer, the scores show which parts of the stack you can see into and which you have to take on trust.
- What changed when engineers stopped training the model: A 2019 Microsoft study that surveyed 551 ML practitioners and a 2023 interview study of 26 engineers building LLM copilots, read side by side: where the data work went, why every test became flaky, and what "the model" means when you rent it.
- The bill for a guardrail pipeline: what input, output and dialog checks catch and cost: Anthropic's classifier guards survived an estimated 3,000+ hours of paid red teaming, but the first version refused about 44% of real traffic. Reading Constitutional Classifiers, Llama Guard and NeMo Guardrails as one pipeline, check by check, with the catch rate and the price each paper reports.
- One tool call, traced through four hops: On API-Bank's 73 working APIs, GPT-4 made 63.66% of calls correctly when told which API to use and 37.04% when it had to find one first. How function calling moves a request from schema to call to runtime to answer, and which failures researchers measured at each hop.
- Four serving metrics and the stalls they hide: In one benchmark, 60% of vLLM requests queued for over 25 seconds, yet its normalized latency was within a few hundred milliseconds of a system that never queued past 15. TTFT, TPOT, throughput and cost per token each have a case they cannot see. Here is each one defined, and what it misses.
- When the test labels are wrong: label errors, data cascades, and datasheets: In 2021 a team checked ten widely used test sets and estimated that at least 3.3% of their labels are wrong on average, about 6% for ImageNet. Correct them and smaller models start beating bigger ones. How those errors were found, how data problems compound in deployed systems, and what documenting a dataset asks of you.
- How few bits can a language model survive?: Quantization stores a model's numbers with fewer bits. A 35,000-experiment study found 4 bits was the best trade almost everywhere, and a model trained from scratch with weights of only -1, 0 and +1 pushed below that. Here is the mapping, the arithmetic, and the conditions attached to each result.
- Fine-tuning or retrieval for new knowledge: what controlled tests found: In 2023 a Microsoft team fine-tuned three 7B models on news the models had never seen, and retrieval beat fine-tuning on every one; for Llama 2, fine-tuning made the score worse. Three controlled studies on putting new facts into a model, what each measured, and why training on unknown facts raises hallucination.
- An agent is a loop, and the loop still loses to people: On WebArena's 812 realistic web tasks, the best GPT-4 agent succeeded 14.41% of the time and humans 78.24%. What the ReAct loop of thought, action and observation is, what it gained over acting or reasoning alone, and where it measurably breaks.
- What RAG is, and why it helps with obscure facts more than famous ones: On 4,000 long-tail questions GPT-3 davinci-003 was right 19% of the time, and a 2.7B model with one retrieved paragraph beat it. On famous subjects the same retrieval made answers worse. Where retrieval-augmented generation came from (REALM and RAG), what its equation says, and when to skip the search.
- Prompt injection: who writes the prompt, and what defenses measurably do: In 2023 researchers hid instructions in a web page and Bing Chat started coaxing its user into giving up their name. The measurements since then show how often injected text wins, and what each defense costs in accuracy.
- Prompting Folklore, Checked Against What Papers Measured: An optimizer found that "Take a deep breath and work on this problem step-by-step." lifted PaLM 2-L to 80.2% on GSM8K. Four popular prompting claims, each set against what a paper actually measured, including where the effect was tiny or belonged to one model.
- Two bent rulers: grading text that has no answer key: In 2021, paid readers asked to tell GPT-3 stories, news and recipes from human ones were right 49.9% of the time, a coin flip. Word-overlap metrics do no better at tracking human ratings. Why open-ended generation is harder to evaluate than classification, measured on both rulers we use for it.
- Varghese never existed: the statistics of confident falsehoods: In 2023 a federal judge sanctioned lawyers for citing court opinions ChatGPT invented, then vouched for. Two papers explain why calibrated models must make some facts up and why benchmarks reward guessing over "I don't know." TruthfulQA and FActScore measure how often it happens.
- How one pretrained model replaced a model per task: In 2018 BERT set new results on eleven language tasks by adding one output layer to a single pretrained model. Read with GPT-1 before it and GPT-2 after, it shows how the foundation model idea formed: pretrain once on raw text, then adapt, with fine-tuning or with none.