Valid JSON, wrong values: a failure catalog for document extraction
Barakaeli Lawuo, Jul 5, 2026
In 2023 a Stanford and Cornell team pointed text-davinci-003 at FDA 510(k) reviews, the roughly 20-page PDFs a device maker files before selling a medical device. They gave it one fixed prompt: list the attributes in this document and their values. On that set the model missed an average of 4.4 of the 16 attributes a human annotator had marked, 27.5% of them, in every document. It also produced an average of 9.7 attributes or values per document that the document did not explicitly mention. And it named things inconsistently: across a sample of 10 documents, the device classification came back as "classification", "device classification", "regulatory information", or not at all.1
The paper's system, Evaporate, asks for a simple list of attribute and value pairs that it turns into a table, and none of the three failures it lists is about that format.1 They are about content. The authors' own summary of the failure is blunt: "Since the error modes are quite varied, it is unclear how to improve quality."1 This post takes that sentence as a challenge. Later papers measured the variety more carefully, on forms, receipts and invoices, and the errors fall into four families. Each one has a paper that caught it and at least one mitigation someone measured.
- Information extraction (IE)
- Turning unstructured text into structured records, such as entities, relations and events. A recent survey defines it that way and frames the LLM version as generation: the model writes out the target structure token by token, given the text and a prompt.4
- Schema and field
- The schema is the list of slots to fill, each with a name and a data type, like file_date as a date or registration_num as digits. A field (papers also say attribute, key or entity type) is one slot.5
- OCR
- Optical character recognition, the step that turns a scanned page into text lines with bounding boxes. Text-only LLM pipelines see whatever OCR produced, errors included.2
- Hierarchical entity
- A field made of grouped sub-fields, like an invoice line item made of description, dates and price. Getting it right means getting the grouping right too.3
- Grounding
- Checking that an extracted value actually appears in the source document, at a place you can point to.2
What the score counts as correct
Before the catalog, it helps to see how these papers score an extraction, because the scoring rule decides which failures you can even see. Evaporate uses Pair F1. Every cell in the output table becomes a tuple of document, attribute and value, and a predicted tuple counts only if it exactly matches a tuple in the hand-built ground truth.1
is the gold set: one tuple per filled cell, where is the document, the attribute and the value. is what the system produced. Precision is the share of predicted tuples that are right, so every invented value drags it down. Recall is the share of gold tuples the system recovered, so every skipped field drags it down. is their harmonic mean, which stays low if either one is low. On the FDA reports, direct prompting scored 45.5 Pair F1.1
The quiet part is the : what counts as a match. Under exact matching, "$ 40,000" and "40,000" are different values, and so are "July 1, 2022" and "07/01/2022". The VRDU benchmark, built from political ad-buy invoices filed with the FCC and foreign-agent registration forms, rejects that rule. Its evaluation tool matches by data type: price values are converted to numbers before comparison, dates are parsed and compared as dates, and addresses stay strict, since "4, Main St." and "40 Main St." are not the same place.3 Another receipt benchmark reports two scores side by side. Exact Match gives a key and value pair credit only if both match after lowercasing and whitespace cleanup, while token-level Value F1 gives partial credit.6 Keep that gap in mind. It comes back in the number section.
Failure 1: a field that is there comes back empty
The FDA result is the clearest case. Of the gold attributes the model missed in a given document, every one was extracted in at least one other document.1 The model could find them. It just did not do it every time, so this is a consistency failure, not a capability gap. Evaporate's comparison across model providers found another flavor: asked for one attribute value, Claude-V1 sometimes replied in chatbot style, "I'm not sure, please give me more information," instead of giving a value.1 Either way, the field comes back empty.
Empty is ambiguous, and that ambiguity is what the mitigations target. Evaporate's code-generation variant writes many small extraction functions and has to decide whether a function's empty output means "this document has no such field" or "this function could not handle this document". A function written for a lowercase "k" product code returns nothing on documents that use an uppercase "K", for instance.1 The system estimates how often the attribute is present by asking the LLM on up to 10 sample documents, then treats empty outputs accordingly. That abstention handling added 1.9 Pair F1 on average and 7.8 on the FDA setting, on top of filtering out bad functions.1
Google's LMDX work measured a cheaper fix at the prompt level. Their completions list every schema field in order and write null for a missing single field or [] for a missing repeated one. When they trained the model to skip absent fields instead, micro-F1 on the ad-buy invoices fell from 54.35 to 47.58, a 6.77 point drop.2 Their explanation is a hypothesis, labeled as one: with explicit nulls the model copies the next key from the schema and makes a present-or-absent call, while skipping forces it to pick which of the remaining keys comes next.2
Failure 2: a value the document never states
The 9.7 unmentioned attributes or values per FDA document are the first half of the Evaporate result.1 The same pattern shows up on receipts. A 2026 benchmark of six open 7B to 8B models on the FUNSD, SROIE and CORD datasets used a prompt that explicitly discourages hallucinated fields. The models still sometimes invented fields such as "subtotal" or "invoice number" on receipts that did not contain them. On long receipts they also over-extracted, trying to label nearly every number on the page.6 Telling the model not to invent things did not make it stop.
Some hallucinations are hard to spot because they sit one digit away from the truth. In a manual review of LLM extractions from VRDU registration forms, Colakoglu and colleagues give an example of a hallucinated date: the model wrote "1992-04-24" where the form says "1992-04-21".5 That output has the right key, a valid date format and a plausible value. Nothing in the JSON structure flags it.
The measured mitigation is grounding. LMDX puts a coordinate token after every OCR line in the prompt, such as "Apple Store 38|05", and asks the model to copy that token next to each extracted value. Decoding then looks up the line by its coordinates and checks that the extracted text really appears on it. If it does not, the value is thrown away.2 On the ad-buy invoices with no target-domain training, 0.59% of completions contained such a mismatch, and the check discarded them. Invalid JSON, by comparison, showed up in only 0.18% of completions.2 (Zero-shot here means no ad-buy training documents. The model had been fine-tuned on other forms first, to learn the task and the output syntax.2) LMDX also samples 16 completions per chunk and takes a majority vote. Dropping to a single completion cost 1.5 micro-F1, mostly because repeat samples let the system recover from a malformed or ungrounded answer.2
Put those two numbers next to each other. Under one in five hundred completions failed to parse, yet the same zero-shot system scored 39.74 micro-F1 on those invoices.2 Almost all of the missing quality sat in values that parsed fine.
Failure 3: a number is misread or loses its unit
Numbers get corrupted before the model ever sees them. The receipt benchmark ran each model twice: once on clean text taken from the human annotations, and once on text from real OCR engines. It reports digit corruption such as "193.00" becoming "19300", along with broken decimal points.6 Its observation about what that does to scores is the most useful sentence in the paper for this topic. These errors often keep partial token overlap, so they earn moderate Value F1 while failing Exact Match completely.6 A lenient metric can make a wrong number look close to right.
On CORD, the best zero-shot model kept a Value F1 of 0.83 on PaddleOCR text, but its Exact Match fell from 0.77 on clean text to 0.45.6 Under Tesseract, the noisiest engine in the study, Value F1 dropped to 0.27.6 The authors conclude that once OCR noise enters, the main source of error shifts from reasoning to input corruption, and that strong semantic modeling cannot compensate for degraded input.6 One caution on this source: it is a 2026 preprint that tests small open models on text alone, so its absolute numbers say less about large multimodal models than its pattern does.
Units and formats fail in a quieter way. The model returns "40,000" for a total printed as "$ 40,000", or a date in a different format from the one in the ground truth. VRDU's type-aware matching exists precisely so those cases score as correct.3 Colakoglu and colleagues measured the same thing from the pipeline side. After GPT-3.5, GPT-4o and LLaMA3-70B extracted from VRDU registration forms, a data-cleaning step reformatted each value using a regular expression for its field type. That lifted average exact-match F1 from 0.650 to 0.734. A schema-mapping step that fixed misspelled keys changed nothing, because the models already returned the right keys.5 So the keys were right and the values needed repair.
Tolerant scoring has its own risk. The same study scored with fuzzy string matching at a 0.8 similarity threshold, then had people check 91 pairs that failed exact match but passed fuzzy match. Fuzzy precision came out at 0.984, not 1.0.5 Their example of a "wrong info" error is a model that returned "2016-10-31" as the file date instead of "2016-10-08".5 A string metric sees two dates that share most of their characters. A person filing a form sees the wrong day. My reading of these results: normalization is safe for money and dates only when it compares parsed values, as VRDU does, and fuzzy matching should never be the check on a numeric field.
Money fields are not always the weak spot, though. On the ad-buy invoices, LMDX scored 98.86 F1 on gross_amount, and removing all layout information barely moved it (98.47). The paper's explanation is that a total can be found from cues like "
quot; or "USD" without reading its label.2 The dates on the same invoices were much harder: 67.74 F1 for the flight start date.2Failure 4: table rows come apart
Line items are where extraction scores collapse. VRDU's authors found that across training-set sizes, the FormNet model's micro-F1 on hierarchical entities trailed its score on other entities by 60 to 70 points. They call proper extraction of hierarchical entities "an open question".3 LLM extractors show the same gap. LMDX's zero-shot comparison on the ad-buy invoices reports line-item F1 separately from overall micro-F1, and the gap is large for every model.
The coordinates are what the LMDX rows add, and the paper tests how much they matter. In an ablation fine-tuned on 10 ad-buy documents, replacing the coordinate tokens with plain line numbers cut line-item F1 from 39.35 to 18.35. Eight of the nine single fields lost less than 10 points.2 The paper's explanation: line-item parts sit in tables, so the model needs horizontal and vertical alignment to group a description with its own dates and price.2

The figure shows the two error patterns the LMDX authors call common. In the first, OCR grouped the Channel and Description columns into one line, so the extracted description carried the channel code along with it. In the second, "Invoice Period" and "Flight Dates" landed on the same OCR line, and the model returned the invoice dates as the flight dates.2 Both answers passed the grounding check, because the text really is on that line. The receipt benchmark describes the same failure in CORD: models attach the right key, such as total, to an item-level amount, which costs Exact Match even when the model found every key.6 VRDU's own annotators made a milder version of this mistake, sometimes confusing the flight dates with other periods on the invoice, such as the invoice period.3
Three mitigations have numbers behind them. Layout in the prompt is one: the 21 point line-item difference above.2 Examples from the same template are another. On CORD, LMDX with in-context examples retrieved by nearest-neighbor search matched its best random-example score using a single example, and matched its fine-tuned score at 10 examples, because the retrieval found receipts from the same merchant.2 The third is the image itself. In the Colakoglu study, GPT-4o-vision and Qwen2.5-vision, given the page image, were the best performers, at about 0.90 F1 against at most 0.80 for the tuned text-only pipelines. GPT-4o-vision used about twice the tokens and cost more than 10 times as much at November 2024 prices.5
Where the grounding check stops
The survey of generative IE lists the "misalignment between natural language output and structured form" as an open challenge next to hallucination.4 The papers above make that concrete. Parse errors are rare and cheap to catch. The hard errors are values that parse, match the schema, and even appear in the document, while belonging to a different field.
LMDX, which built the strongest check in this catalog, says where that check ends. Its input is OCR text lines, so it inherits OCR's mistakes: wrong reading order, incorrect line grouping, undetected text and misrecognized characters.2 And its verification works at the level of a line. It confirms that the extracted text is present on the line the model pointed to. "If the entity text appears multiple times on the line," the authors write, "we don't have a definitive way to choose the correct text."2
Sources
- Arora, Yang, Eyuboglu, Narayan, Hojel, Trummer, Ré. Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes (Evaporate). PVLDB 2023 / arXiv 2304.09433
- Perot et al. LMDX: Language Model-based Document Information Extraction and Localization. arXiv 2309.10952
- Wang, Zhou, Wei, Lee, Tata. VRDU: A Benchmark for Visually-rich Document Understanding. KDD 2023 / arXiv 2211.15421
- Xu et al. Large Language Models for Generative Information Extraction: A Survey. Frontiers of Computer Science 2024 / arXiv 2312.17617
- Colakoglu, Solmaz, Fürst. Problem Solved? Information Extraction Design Space for Layout-Rich Documents using LLMs. arXiv 2502.18179
- Anvari, Athitsos. From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings. arXiv 2609.17538 (preprint, 2026)