<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>Barakaeli Lawuo: blog</title>
    <link>https://www.princetekki.com/blog</link>
    <description>Technical writing on large language models, with sources cited.</description>
    <language>en</language>
    <item>
      <title>What Mixture of Experts Had to Fix Before Mixtral Worked</title>
      <link>https://www.princetekki.com/blog/mixture-of-experts</link>
      <guid>https://www.princetekki.com/blog/mixture-of-experts</guid>
      <pubDate>Sat, 26 Sep 2026 12:00:00 GMT</pubDate>
      <description>In 2020 Google trained a 600 billion parameter MoE model in four days, then left its trillion parameter sibling out of the paper because it kept hitting numerical trouble. Read in order, the MoE papers are a chain of problems and fixes, ending with a finding about what experts actually learn.</description>
    </item>
    <item>
      <title>The Cost Ledger of an LLM Task, Line by Line</title>
      <link>https://www.princetekki.com/blog/cost-token-observability</link>
      <guid>https://www.princetekki.com/blog/cost-token-observability</guid>
      <pubDate>Thu, 24 Sep 2026 12:00:00 GMT</pubDate>
      <description>In 2024 a Princeton team re-ran published coding agents and found that a simple retry loop was as accurate as the best of them and far cheaper than most. Here is the cost of one task built up term by term from papers, then the cascades, routers and caches that measurably cut it.</description>
    </item>
    <item>
      <title>Postmortem of a Silent Regression: What to Log and Alert On for an LLM Feature</title>
      <link>https://www.princetekki.com/blog/what-to-log-and-alert</link>
      <guid>https://www.princetekki.com/blog/what-to-log-and-alert</guid>
      <pubDate>Mon, 21 Sep 2026 12:00:00 GMT</pubDate>
      <description>Between March and June 2023, GPT-4 went from 84% to 51% on a prime number task and from 52% to 10% on code that runs as written, while every request still came back as normal text. Read as an incident report, the drift studies and ML monitoring papers say which signals would have caught it.</description>
    </item>
    <item>
      <title>Anatomy of One Traced LLM Request, Using Google's Dapper as the Guide</title>
      <link>https://www.princetekki.com/blog/tracing-llm-requests</link>
      <guid>https://www.princetekki.com/blog/tracing-llm-requests</guid>
      <pubDate>Sat, 19 Sep 2026 12:00:00 GMT</pubDate>
      <description>Google built Dapper because one search query touched thousands of machines and nobody could say which one was slow. An agent request has the same shape on a smaller scale. This post dissects one span tree, field by field, and then asks whether Dapper's sampling rules still hold when every span is a model call.</description>
    </item>
    <item>
      <title>Why a Passing Eval Suite Still Needs Production Logs</title>
      <link>https://www.princetekki.com/blog/evaluation-vs-observability</link>
      <guid>https://www.princetekki.com/blog/evaluation-vs-observability</guid>
      <pubDate>Wed, 16 Sep 2026 12:00:00 GMT</pubDate>
      <description>In a 2024 Berkeley study, developers kept rewriting their own grading rules as they read more model outputs. That finding, plus research on LLM judges, benchmark leakage, and real user traffic, makes a case that offline evaluation can only be a draft of what production observability later corrects.</description>
    </item>
    <item>
      <title>Timing an LLM Cold Start, Phase by Phase</title>
      <link>https://www.princetekki.com/blog/autoscaling-cold-starts</link>
      <guid>https://www.princetekki.com/blog/autoscaling-cold-starts</guid>
      <pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate>
      <description>In the ServerlessLLM paper, KServe took 128 seconds to return the first token from a cold OPT-6.7B replica, and 114 of those seconds went to downloading the weights. This post times each phase of a GPU cold start with numbers from the papers, then shows what each paper does to shrink it and when those fixes stop helping.</description>
    </item>
    <item>
      <title>Choosing an LLM Serving Stack by the Idea, Not the Benchmark</title>
      <link>https://www.princetekki.com/blog/serving-frameworks-compared</link>
      <guid>https://www.princetekki.com/blog/serving-frameworks-compared</guid>
      <pubDate>Fri, 11 Sep 2026 12:00:00 GMT</pubDate>
      <description>One A100 served a 13B model at 1.6 requests per second within its latency targets. Split across three GPUs by phase, the same model reached 3.3 per GPU. The frameworks differ mostly in which of three scheduling ideas they implement, and each idea helps a different workload.</description>
    </item>
    <item>
      <title>Counting Idle Slots: Where vLLM Gets Its Throughput</title>
      <link>https://www.princetekki.com/blog/vllm-continuous-batching</link>
      <guid>https://www.princetekki.com/blog/vllm-continuous-batching</guid>
      <pubDate>Wed, 09 Sep 2026 12:00:00 GMT</pubDate>
      <description>When the vLLM team profiled existing LLM servers, as little as 20.4% of the memory reserved for the KV cache held actual token states. Walk four requests through two schedulers by hand, then follow the memory, and the 2 to 4 times throughput gain stops looking like magic.</description>
    </item>
    <item>
      <title>Six Web-Service Assumptions an LLM Endpoint Breaks</title>
      <link>https://www.princetekki.com/blog/why-llm-deployment-is-different</link>
      <guid>https://www.princetekki.com/blog/why-llm-deployment-is-different</guid>
      <pubDate>Sun, 06 Sep 2026 12:00:00 GMT</pubDate>
      <description>On Azure production traces, a 1,500-token prompt on BLOOM-176B took as long as generating six output tokens. That one measurement, and a few others from Splitwise and Google's PaLM inference paper, knock over most of what ordinary web-service capacity planning assumes.</description>
    </item>
    <item>
      <title>Tracing a Wrong RAG Answer to the Stage That Broke</title>
      <link>https://www.princetekki.com/blog/rag-evaluation</link>
      <guid>https://www.princetekki.com/blog/rag-evaluation</guid>
      <pubDate>Fri, 04 Sep 2026 12:00:00 GMT</pubDate>
      <description>A Deakin University team built three RAG systems and catalogued seven ways they failed. Those failures sort into four branches, from &quot;never retrieved&quot; to &quot;used but misstated&quot;, and each branch has its own metric and its own record of agreeing, or not, with human judges.</description>
    </item>
    <item>
      <title>Reranking, Read Off the MS MARCO Leaderboard</title>
      <link>https://www.princetekki.com/blog/reranking</link>
      <guid>https://www.princetekki.com/blog/reranking</guid>
      <pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate>
      <description>In January 2019 a BERT reranker took the top of the MS MARCO passage leaderboard with a 27% relative jump in MRR@10, and, as a later paper timed it, about 33 seconds of GPU time per query. The reranking papers that followed are a fight over that bill.</description>
    </item>
    <item>
      <title>CAG vs RAG, Read Off the Tables of Three Papers</title>
      <link>https://www.princetekki.com/blog/cag-vs-rag</link>
      <guid>https://www.princetekki.com/blog/cag-vs-rag</guid>
      <pubDate>Sun, 30 Aug 2026 12:00:00 GMT</pubDate>
      <description>Preloading 85k tokens of HotPotQA into a precomputed KV cache cut answer time from 92 seconds to 2.3 on Llama 3.1 8B. On that same test set, plain BM25 retrieval scored higher. A head-to-head readout of what three papers measured, and where each approach won.</description>
    </item>
    <item>
      <title>HyDE: Searching With an Answer That Might Be Wrong</title>
      <link>https://www.princetekki.com/blog/hyde</link>
      <guid>https://www.princetekki.com/blog/hyde</guid>
      <pubDate>Thu, 27 Aug 2026 12:00:00 GMT</pubDate>
      <description>In December 2022 a CMU and Waterloo team trained nothing and still nearly matched a retriever fine-tuned on MS MARCO. Their trick was to embed a made-up answer instead of the question. Here is one query followed through every step, then the results, the Query2doc variant, and where the papers say it breaks.</description>
    </item>
    <item>
      <title>Who Decides When to Retrieve? Four Retrieval Policies, From Fixed to Learned</title>
      <link>https://www.princetekki.com/blog/agentic-rag</link>
      <guid>https://www.princetekki.com/blog/agentic-rag</guid>
      <pubDate>Tue, 25 Aug 2026 12:00:00 GMT</pubDate>
      <description>Handed the top retrieved passages, ChatGPT got worse on a health fact-checking benchmark, falling from 70.1% to 54.7%. The fix the research tried was to let something other than the pipeline decide when to search. Here are four answers to that question, each with its trigger rule, its results, and what it costs in retrieval calls.</description>
    </item>
    <item>
      <title>From Brute Force to HNSW: How Vector Search Gets Fast</title>
      <link>https://www.princetekki.com/blog/vector-databases-explained</link>
      <guid>https://www.princetekki.com/blog/vector-databases-explained</guid>
      <pubDate>Sat, 22 Aug 2026 12:00:00 GMT</pubDate>
      <description>In 2017 FAISS searched a billion image vectors at 0.0133 ms per query while storing each one in 20 bytes. Getting there took three ideas, each one a fix for what the previous one still cost: inverted lists, product quantization, and layered graphs.</description>
    </item>
    <item>
      <title>One Turn of the Agent Improvement Loop, Measured</title>
      <link>https://www.princetekki.com/blog/agent-improvement-flywheel</link>
      <guid>https://www.princetekki.com/blog/agent-improvement-flywheel</guid>
      <pubDate>Thu, 20 Aug 2026 12:00:00 GMT</pubDate>
      <description>In 2023 an agent built on gpt-3.5-turbo went from 40% to 59% on ALFWorld household tasks with no weight updates, only by studying its own past runs. This post follows that loop one stage at a time (collect, label, distill, redeploy) using the papers that measured each stage.</description>
    </item>
    <item>
      <title>61% Once, Under 25% Eight Times: Deploying Agents That Act</title>
      <link>https://www.princetekki.com/blog/agent-deployment</link>
      <guid>https://www.princetekki.com/blog/agent-deployment</guid>
      <pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate>
      <description>In 2024 the tau-bench authors ran gpt-4o on the same customer service tasks again and again. It solved about 61% on average, but fewer than 25% of tasks on all eight tries. Here is the math behind that drop, what ToolEmu measured about risky actions, and what the guardrails in these papers actually bought.</description>
    </item>
    <item>
      <title>One A2A Handoff, Message by Message, and Where Each Message Can Be Abused</title>
      <link>https://www.princetekki.com/blog/a2a-protocol</link>
      <guid>https://www.princetekki.com/blog/a2a-protocol</guid>
      <pubDate>Sat, 15 Aug 2026 12:00:00 GMT</pubDate>
      <description>A 2026 analysis found 11 attacks on the Agent2Agent protocol that work without breaking a single rule of the spec. Follow one task from Agent Card to artifact and you can see where each one lives.</description>
    </item>
    <item>
      <title>Four Ways to Wire LLM Agents, and What Each One Measured</title>
      <link>https://www.princetekki.com/blog/multi-agent-orchestration</link>
      <guid>https://www.princetekki.com/blog/multi-agent-orchestration</guid>
      <pubDate>Wed, 12 Aug 2026 12:00:00 GMT</pubDate>
      <description>Three copies of gpt-3.5-turbo that read and revise each other's answers solved 81.8% of an arithmetic test the same model solved 67.0% of alone. Pipelines, managers, group chats and debates each come from a paper that measured them. Here is the wiring, the numbers, and the bill.</description>
    </item>
    <item>
      <title>Splitting an Agent: An Evidence Ledger</title>
      <link>https://www.princetekki.com/blog/single-vs-multi-agent</link>
      <guid>https://www.princetekki.com/blog/single-vs-multi-agent</guid>
      <pubDate>Mon, 10 Aug 2026 12:00:00 GMT</pubDate>
      <description>A Berkeley team annotated 1,642 traces from seven multi-agent frameworks and found failure rates between 41% and 86.7%. Here is what the controlled studies say about when splitting a task across agents helps, when it hurts, and whether the gains survive an equal token budget.</description>
    </item>
    <item>
      <title>Who Holds the Off Switch: A Ladder of Agent Autonomy</title>
      <link>https://www.princetekki.com/blog/levels-of-agentic-autonomy</link>
      <guid>https://www.princetekki.com/blog/levels-of-agentic-autonomy</guid>
      <pubDate>Fri, 07 Aug 2026 12:00:00 GMT</pubDate>
      <description>Three papers rank AI agents by how much the human still decides. Walked rung by rung, they show what the user actually does at each level, which risk arrives with it, and why one of them argues the top rung should never be built.</description>
    </item>
    <item>
      <title>What an MCP Server Says on the Wire, and What Studies Found in Real Ones</title>
      <link>https://www.princetekki.com/blog/build-an-mcp-server</link>
      <guid>https://www.princetekki.com/blog/build-an-mcp-server</guid>
      <pubDate>Wed, 05 Aug 2026 12:00:00 GMT</pubDate>
      <description>A walk through one stdio session of the Model Context Protocol, one JSON-RPC message at a time, with the measured problems researchers found in real open-source servers placed at the exact message where each one enters.</description>
    </item>
    <item>
      <title>More Tools, Worse Picks: What the Tool Selection Papers Measured</title>
      <link>https://www.princetekki.com/blog/mcp-tool-overload</link>
      <guid>https://www.princetekki.com/blog/mcp-tool-overload</guid>
      <pubDate>Sun, 02 Aug 2026 12:00:00 GMT</pubDate>
      <description>When RAG-MCP buried one correct MCP server among thousands of distractors, success held above 90% for small pools and collapsed past about 100. Here is the curve, the reasons the papers give for it, and what retrieval, two-stage routing, and fine-tuning on API docs actually recovered.</description>
    </item>
    <item>
      <title>Who pulls the trigger: MCP tools, resources and prompts, and what breaks when the wrong party does</title>
      <link>https://www.princetekki.com/blog/mcp-primitives</link>
      <guid>https://www.princetekki.com/blog/mcp-primitives</guid>
      <pubDate>Fri, 31 Jul 2026 12:00:00 GMT</pubDate>
      <description>MCP gives each server primitive an owner: the model fires tools, the application picks resources, the user picks prompts. Here is what the specification says, what attack studies measured against each primitive, and why no paper can yet say how often real servers use the last two.</description>
    </item>
    <item>
      <title>Host, Client, Server: Where MCP Draws Its Trust Boundaries</title>
      <link>https://www.princetekki.com/blog/mcp-architecture</link>
      <guid>https://www.princetekki.com/blog/mcp-architecture</guid>
      <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
      <description>A security benchmark pointed a web page at a local MCP server and reached it every time, on Claude Desktop, OpenAI and Cursor alike. A tour of MCP's three roles, what the spec makes each one responsible for, and what the research found breaking at each boundary.</description>
    </item>
    <item>
      <title>Function Calling, MCP and the API Underneath: A Stack Read Bottom Up</title>
      <link>https://www.princetekki.com/blog/mcp-vs-function-calling</link>
      <guid>https://www.princetekki.com/blog/mcp-vs-function-calling</guid>
      <pubDate>Sun, 26 Jul 2026 12:00:00 GMT</pubDate>
      <description>In the Berkeley Function Calling Leaderboard paper, o1 scored 91.5 on parallel calls when the tools were described in its prompt and 0.0 when they went through its native tools field. Same model, same functions. That gap only makes sense once you see function calling, MCP and a plain API as separate layers.</description>
    </item>
    <item>
      <title>Before MCP there was LSP: the M times N problem and what studies found</title>
      <link>https://www.princetekki.com/blog/what-mcp-solves</link>
      <guid>https://www.princetekki.com/blog/what-mcp-solves</guid>
      <pubDate>Thu, 23 Jul 2026 12:00:00 GMT</pubDate>
      <description>The argument for MCP, that M apps times N tools should become M plus N, was first made for code editors. Here is what research found when the Language Server Protocol tried it, and what ecosystem measurements show about MCP so far.</description>
    </item>
    <item>
      <title>Same Retriever, Different Driver: Fixed RAG Pipelines vs Models That Search</title>
      <link>https://www.princetekki.com/blog/manual-rag-vs-agentic-context</link>
      <guid>https://www.princetekki.com/blog/manual-rag-vs-agentic-context</guid>
      <pubDate>Tue, 21 Jul 2026 12:00:00 GMT</pubDate>
      <description>Search-R1 kept the retriever, the corpus and the three passages per call fixed, and changed only who writes the queries. Exact match on seven QA sets went from 0.304 to 0.431. The papers that measured the bill, and the cases where the fixed pipeline still won, tell the rest.</description>
    </item>
    <item>
      <title>Context Rot: Three Ways a Longer Prompt Loses the Answer</title>
      <link>https://www.princetekki.com/blog/context-rot</link>
      <guid>https://www.princetekki.com/blog/context-rot</guid>
      <pubDate>Sat, 18 Jul 2026 12:00:00 GMT</pubDate>
      <description>With 20 documents in its prompt, GPT-3.5-Turbo answered 75.8% of questions when the right one came first and 53.8% when it came tenth, below its score with no documents at all. Four papers explain why more context can make answers worse: position, sheer length, and questions that share no words with their answers.</description>
    </item>
    <item>
      <title>Context Budgeting: A Worksheet for What to Keep, Drop, or Compress</title>
      <link>https://www.princetekki.com/blog/context-budgeting</link>
      <guid>https://www.princetekki.com/blog/context-budgeting</guid>
      <pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate>
      <description>LLMLingua cut a 2,366-token math prompt to 117 tokens and GPT-3.5-Turbo lost 1.5 points of accuracy. Here is a token budget worksheet, and what three papers measured for each way of making things fit: dropping, summarizing, and deleting tokens.</description>
    </item>
    <item>
      <title>What an agent reads on every request: four memories and a window</title>
      <link>https://www.princetekki.com/blog/six-types-of-agent-context</link>
      <guid>https://www.princetekki.com/blog/six-types-of-agent-context</guid>
      <pubDate>Mon, 13 Jul 2026 12:00:00 GMT</pubDate>
      <description>Generative Agents lost believability each time a memory component was switched off. CoALA gives the vocabulary for why: working, episodic, semantic and procedural memory, plus the tool results that flow back in. Here is each one as the papers define it, and how MemGPT and Generative Agents build and measure it.</description>
    </item>
    <item>
      <title>Prompt Wording vs Window Contents: Where the Accuracy Moves</title>
      <link>https://www.princetekki.com/blog/context-vs-prompt-engineering</link>
      <guid>https://www.princetekki.com/blog/context-vs-prompt-engineering</guid>
      <pubDate>Sat, 11 Jul 2026 12:00:00 GMT</pubDate>
      <description>Changing only the separators and spacing of a prompt swung LLaMA-2-7B from 3.6% to 80.4% accuracy on one task. Wording clearly matters. The evidence on what it can and cannot fix, set against what changing the information in the window does, points to where the effort belongs.</description>
    </item>
    <item>
      <title>Valid JSON, wrong values: a failure catalog for document extraction</title>
      <link>https://www.princetekki.com/blog/information-extraction-prompts</link>
      <guid>https://www.princetekki.com/blog/information-extraction-prompts</guid>
      <pubDate>Sun, 05 Jul 2026 12:00:00 GMT</pubDate>
      <description>On 20-page FDA device reviews, a single-prompt LLM extractor skipped about 27.5% of the fields a human would fill and added values the documents never state. The papers that measured this show four distinct ways extraction goes wrong, and none of them breaks the JSON.</description>
    </item>
    <item>
      <title>Auditing a leaderboard number: what MMLU hides and what to measure instead</title>
      <link>https://www.princetekki.com/blog/capabilities-that-matter</link>
      <guid>https://www.princetekki.com/blog/capabilities-that-matter</guid>
      <pubDate>Sat, 04 Jul 2026 12:00:00 GMT</pubDate>
      <description>An audit of MMLU found errors in 57% of sampled Virology questions and an estimated 6.49% error rate overall. Add HELM's other metrics and Chatbot Arena's preference votes, and a single leaderboard score says little about which model fits your application.</description>
    </item>
    <item>
      <title>From CheckList tests to error bars: building an eval pipeline</title>
      <link>https://www.princetekki.com/blog/evaluation-pipeline</link>
      <guid>https://www.princetekki.com/blog/evaluation-pipeline</guid>
      <pubDate>Fri, 03 Jul 2026 12:00:00 GMT</pubDate>
      <description>In 2019, three commercial sentiment APIs failed nearly every sentence like &quot;I thought the plane would be awful, but it wasn't.&quot; Build an evaluation pipeline in order: capability tests from CheckList, per-case scores, standard errors and paired comparisons from Miller, and a case count large enough to see the difference you care about.</description>
    </item>
    <item>
      <title>BloombergGPT's ledger: what training your own model cost, and what adapting one scored</title>
      <link>https://www.princetekki.com/blog/build-vs-buy-model</link>
      <guid>https://www.princetekki.com/blog/build-vs-buy-model</guid>
      <pubDate>Thu, 02 Jul 2026 12:00:00 GMT</pubDate>
      <description>Bloomberg spent a 1.3 million GPU hour budget training a 50 billion parameter finance model. Months later, GPT-4, prompted and never retrained, beat it on four of the five public benchmarks it had reported. A line by line ledger of what each path costs a team, from the papers' own figures.</description>
    </item>
    <item>
      <title>Jailbroken: the two ways safety training loses, and what defenses actually buy</title>
      <link>https://www.princetekki.com/blog/jailbreaking</link>
      <guid>https://www.princetekki.com/blog/jailbreaking</guid>
      <pubDate>Wed, 01 Jul 2026 12:00:00 GMT</pubDate>
      <description>In 2023 Berkeley researchers found, for every one of 32 red-team prompts, an attack that got GPT-4 and Claude v1.3 to answer it. Their two failure modes still explain automated suffix attacks, many-shot prompts, and why every defense measured since moves the numbers without closing the gap.</description>
    </item>
    <item>
      <title>The measured biases of LLM judges, one at a time</title>
      <link>https://www.princetekki.com/blog/ai-as-a-judge</link>
      <guid>https://www.princetekki.com/blog/ai-as-a-judge</guid>
      <pubDate>Tue, 30 Jun 2026 12:00:00 GMT</pubDate>
      <description>In 2023, GPT-4 judged the same 80 pairs of answers twice with only their order swapped and contradicted itself on 37 of them. A catalog of the biases papers have actually measured in LLM judges, how big each one was, which fixes were tested, and when the grades came close to human ones.</description>
    </item>
    <item>
      <title>When the tests got bigger: how thin test suites overstate code correctness</title>
      <link>https://www.princetekki.com/blog/functional-correctness</link>
      <guid>https://www.princetekki.com/blog/functional-correctness</guid>
      <pubDate>Mon, 29 Jun 2026 12:00:00 GMT</pubDate>
      <description>EvalPlus gave each HumanEval problem about 80 times more tests, and GPT-4's greedy pass rate fell from 88.4% to 76.2%. What the original tests missed, how pass@k is actually computed, and what changes when the task is a whole repository.</description>
    </item>
    <item>
      <title>Detection without a detector: from a closed list to a sentence</title>
      <link>https://www.princetekki.com/blog/detection-without-a-detector</link>
      <guid>https://www.princetekki.com/blog/detection-without-a-detector</guid>
      <pubDate>Mon, 29 Jun 2026 12:00:00 GMT</pubDate>
      <description>OWL-ViT found LVIS rare categories at 31.2 AP after every box for them was removed from its training data. How image-text training turns class names into vectors, how OWL-ViT and Grounding DINO turn those vectors into boxes, and where the papers say it still breaks.</description>
    </item>
    <item>
      <title>What Specializing a Model for Medicine Measurably Bought</title>
      <link>https://www.princetekki.com/blog/domain-specific-models</link>
      <guid>https://www.princetekki.com/blog/domain-specific-models</guid>
      <pubDate>Sun, 28 Jun 2026 12:00:00 GMT</pubDate>
      <description>Flan-PaLM passed a medical licensing exam benchmark, yet clinicians judged only 61.9% of its answers to match scientific consensus. Med-PaLM, Med-PaLM 2 and Medprompt show what specialization changes, and which yardstick each result was measured with.</description>
    </item>
    <item>
      <title>The price of a sentence: why models do worse outside English</title>
      <link>https://www.princetekki.com/blog/multilingual-quality</link>
      <guid>https://www.princetekki.com/blog/multilingual-quality</guid>
      <pubDate>Sat, 27 Jun 2026 12:00:00 GMT</pubDate>
      <description>The tokenizer behind ChatGPT and GPT-4 needs about 15 times as many tokens for a sentence in Shan as for the same sentence in English. What that premium is made of, what it does to cost, context and latency, how accuracy falls with a language's share of the training data, and which fixes the papers measured.</description>
    </item>
    <item>
      <title>One distribution, three dials, and a temperature zero that still wobbles</title>
      <link>https://www.princetekki.com/blog/probabilistic-nature</link>
      <guid>https://www.princetekki.com/blog/probabilistic-nature</guid>
      <pubDate>Fri, 26 Jun 2026 12:00:00 GMT</pubDate>
      <description>Researchers sent ChatGPT the same coding prompt five times and, on 75.76% of CodeContests problems, got five programs with no test output in common. Where that variation comes from: temperature, top-k, nucleus sampling, and the batching arithmetic that keeps temperature 0 from being deterministic.</description>
    </item>
    <item>
      <title>The most probable sentence is often empty: decoding after beam search</title>
      <link>https://www.princetekki.com/blog/next-word-sampling</link>
      <guid>https://www.princetekki.com/blog/next-word-sampling</guid>
      <pubDate>Thu, 25 Jun 2026 12:00:00 GMT</pubDate>
      <description>Widening the beam dropped a translation model's BLEU from 36.42 to 14.66, and exact search picked the empty string for most sentences. Why the most probable output is not the best one, and how locally typical sampling, contrastive search and min-p each answer that, with their measured results and the dispute over min-p.</description>
    </item>
    <item>
      <title>The KV cache and three ways to shrink it</title>
      <link>https://www.princetekki.com/blog/kv-cache</link>
      <guid>https://www.princetekki.com/blog/kv-cache</guid>
      <pubDate>Wed, 24 Jun 2026 12:00:00 GMT</pubDate>
      <description>DeepSeek-V2 reported a KV cache 93.3% smaller than its predecessor's. Why generation stores keys and values at all, how to size that store, and what multi-query, grouped-query and latent attention each gave up to make it smaller.</description>
    </item>
    <item>
      <title>BM25 and dense retrieval, two scoring functions side by side</title>
      <link>https://www.princetekki.com/blog/retrieval-algorithms</link>
      <guid>https://www.princetekki.com/blog/retrieval-algorithms</guid>
      <pubDate>Wed, 17 Jun 2026 12:00:00 GMT</pubDate>
      <description>Dense Passage Retrieval beat BM25 by 19 points of top-20 accuracy on Natural Questions. On BEIR's 18 unseen datasets, the same model averaged 47.7% worse. What each formula actually computes, term by term, why each wins where it does, and what the papers report when you fuse them.</description>
    </item>
    <item>
      <title>How small should a chunk be? What retrieval papers measured</title>
      <link>https://www.princetekki.com/blog/chunking-for-rag</link>
      <guid>https://www.princetekki.com/blog/chunking-for-rag</guid>
      <pubDate>Wed, 10 Jun 2026 12:00:00 GMT</pubDate>
      <description>Re-indexing Wikipedia as one-fact propositions lifted an unsupervised retriever's Recall@5 on EntityQuestions from 36.3% to 51.7% with no retraining. What the research found about chunk size, semantic boundaries, late chunking and RAPTOR's summary trees, including where each one did not pay.</description>
    </item>
    <item>
      <title>How 2,048 Tokens Became 32,768: A History of the Context Window</title>
      <link>https://www.princetekki.com/blog/context-length</link>
      <guid>https://www.princetekki.com/blog/context-length</guid>
      <pubDate>Wed, 03 Jun 2026 12:00:00 GMT</pubDate>
      <description>In 2023 a Meta team stretched LLaMA from 2,048 to 32,768 tokens with about 1,000 fine-tuning steps. The number had held because of how models encode position. This is the story of RoPE, ALiBi, position interpolation and YaRN, told through the perplexity each one measured.</description>
    </item>
    <item>
      <title>Chain of thought: the 17.9 to 56.9 result and its fine print</title>
      <link>https://www.princetekki.com/blog/chain-of-thought</link>
      <guid>https://www.princetekki.com/blog/chain-of-thought</guid>
      <pubDate>Wed, 27 May 2026 12:00:00 GMT</pubDate>
      <description>Eight worked examples took PaLM 540B from 17.9% to 56.9% on grade-school math. The papers that followed added four conditions: it needs scale, one sentence can trigger it, voting over many chains helps, and the written reasoning may not be why the model answered.</description>
    </item>
    <item>
      <title>Six questions to ask a leaderboard, learned from the Chatbot Arena dispute</title>
      <link>https://www.princetekki.com/blog/reading-benchmarks</link>
      <guid>https://www.princetekki.com/blog/reading-benchmarks</guid>
      <pubDate>Wed, 20 May 2026 12:00:00 GMT</pubDate>
      <description>In early 2025 researchers counted 27 private Meta variants on Chatbot Arena in the run-up to Llama 4, and showed that picking the best of several tries lifts a score. The maintainers pushed back. The argument is a good checklist for reading any benchmark claim.</description>
    </item>
  </channel>
</rss>
