Files
duocthu/docs/pdf-parsing-outlier-catalog.md
T

436 lines
25 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PDF Parsing Outlier Catalog
A generalized checklist of structural risks found while parsing
`duoc-thu-quoc-gia-viet-nam-2018.pdf` (1668 pages). Every item here was
**confirmed with real evidence** (bounding-box inspection, cross-tool
comparison, or a whole-document scan) — not assumed. The goal of this
document is reuse: if this project (or a future one) needs to parse another
structured reference PDF — another national formulary, a different
government-published multi-part document, any dense print-layout book —
this is the checklist of "things that go wrong that a small page sample
won't reveal," and how to actually check for each one cheaply (most checks
here run over the whole 1668-page book in under a minute).
For the narrative investigation and drug-formulary-specific numbers, see
`docs/adr/0003-pdf-parsing-strategy.md`. This document is the distilled,
reusable checklist form of the same findings, plus items found afterward.
---
## Structural discovery risks (before you even parse content)
### 1. No bookmarks/TOC
**What it looks like:** `doc.get_toc()` (PyMuPDF) returns an empty list.
**Why it matters:** the obvious, easiest structural signal for section
boundaries simply doesn't exist — don't design a pipeline that assumes it
will.
**Check:** one line, `len(doc.get_toc())`. Do this first, always, before
assuming a bookmark-based approach.
**Generalizes:** yes, directly — always check this before designing around
bookmarks, for any PDF.
### 2. Shallow/unusable tagged-PDF structure tree
**What it looks like:** the PDF has a `/StructTreeRoot` (looks promising —
"tagged PDF"), but it only covers a handful of generic `/H1`/`/P` elements
for a fraction of the document (here: ~29 elements for 1668 pages).
**Why it matters:** easy to assume "tagged PDF = rich semantic structure
available"; in practice many tagging tools produce a minimal
compliance-only tree that covers almost nothing.
**Check:** walk the struct tree (`doc.xref_object` on `/StructTreeRoot`,
recurse into `/K`) and count real leaf elements vs. total page count. If the
ratio is tiny, it's not a usable data source.
**Generalizes:** yes — always verify depth/coverage before trusting a
struct tree, don't just check for its existence.
---
## Page layout risks
### 3. Multi-column body layout
**What it looks like:** body pages are genuinely two-column (confirmed via
bounding boxes: left column x≈44-299, right column x≈308-562, page width
≈595). Front-matter pages that *look* like a multi-column name grid to the
eye turned out, on inspection, to be single wide text blocks with internal
whitespace padding between names — not a real structural column split.
**Why it matters:** a naive "read text top-to-bottom regardless of x" pass
would interleave left- and right-column content into nonsense. Conversely,
assuming every visually grid-like page is column-split leads to wasted
effort — verify per page/section, don't generalize from appearance alone.
**Check:** for any suspicious page, dump block bounding boxes
(`page.get_text("dict")["blocks"]`) and look at the actual x0/x1 ranges. A
real column split shows two clusters of x-ranges; a padded single-column
list shows one wide range per line.
**Handling:** PyMuPDF's default block-level reading order handled the real
two-column case correctly here (validated against a known monograph) — the
tool most likely to get column order wrong was `pdfplumber`'s general
`extract_text()` (see item 8), not PyMuPDF.
**Generalizes:** yes — this exact check (dump bboxes, look at x-clusters)
works on any PDF to determine real column count before writing extraction
logic.
### 4. Full-width content breaking out of the column grid
**What it looks like:** some pages have a table (or could have a figure)
that spans nearly the entire page width (confirmed: a body-surface-area
lookup table's blocks span x≈35 to x≈553, i.e. across both normal columns),
overriding the page's usual two-column layout.
**Why it matters:** logic written to always split a page into "left column"
and "right column" text will misbehave on these pages — the content isn't
in either column, it's a single full-width unit.
**Check:** for any block, compare its x-width against the known
single-column width; if a block's x-range spans (or nearly spans) both
known column ranges, treat it as a full-width unit, not part of a column.
**Generalizes:** yes — any multi-column layout can have occasional
full-width breakout elements (tables, figures, pull-quotes); always check
for this rather than assuming rigid column adherence everywhere.
---
## Table-specific risks
### 5. Tables split across a page break lose their header on the continuation page
**What it looks like:** confirmed directly — "Bảng 4: Xử trí về điều trị ARV
theo mức độ phát ban" (a 3-column table) starts on one page with its header
row (`['Mức độ', 'Biểu hiện', 'Xử trí']`) and 3 data rows; its 4th data row
("Mức độ 4...") appears on the **next page**, extracted by `pdfplumber`
as a **separate table object with no header row at all**.
**Why it matters:** if a pipeline treats each `find_tables()`/
`extract_tables()` result as an independent, self-contained table, the
orphaned continuation row is meaningless on its own — you lose the column
semantics for that row entirely.
**Check:** for any table-like structure, check whether the page/column
immediately preceding it ends with a same-shaped table lacking a natural
final row (e.g. an incomplete-looking sequence) — a strong heuristic is
"table starts at the very top of a page/column, no header, same column
count as the table ending at the bottom of the previous page/column."
**Handling:** never treat page-extracted tables as independent; track
continuation explicitly and re-attach the original header to orphaned
continuation rows before using them.
**Generalizes:** yes — this is a generic multi-page-table risk in any
paginated PDF with tall tables; the detection heuristic (position at
page/column top + no header + matching column count to the previous
table) applies broadly.
### 6. Tables can also split across a column boundary on the *same* page
**What it looks like:** confirmed — "Bảng 6" (ARV drug toxicity table)
starts in the left column near the bottom of a page (header + first data
row) and its remaining data rows appear at the **top of the right column of
the very same page**, again with no header repeated.
**Why it matters:** this is easy to miss because there's no literal page
break — it's tempting to assume "if it's the same page, it's not split,"
but a table can still be taller than one column's usable height.
**Check:** same heuristic as item 5, but also check column position, not
just page number — a header-less table fragment starting at the top of a
column (regardless of page) is a suspect continuation.
**Generalizes:** yes, wherever content flows in columns at all — this risk
exists any time column height is shorter than table height.
### 7. Two-dimensional grid/nomogram tables are not linearly recoverable
**What it looks like:** confirmed — a body-surface-area lookup table
(height across the top, weight down the side, a BSA value at each
intersection) extracts as a scrambled sequence of numbers with no
recoverable row/column association from plain text alone (e.g. `"0,50
0,52 0,54 0,56"` followed by `"0,55 0,57 0,59 0,61"` — these are almost
certainly column-wise fragments, not the visual rows).
**Why it matters:** unlike a normal bordered table (rows of related
values), a 2D lookup grid's *meaning* depends entirely on 2D position — a
number is meaningless without knowing both its row header (weight) and
column header (height). Flattened text extraction destroys exactly the
information needed to interpret it.
**Check:** any table where extracted "cells" are bare numbers with no
inline label, laid out in a dense grid, is a candidate — cross-check
against the source's own stated formula/description (this table is
explicitly a lookup version of a stated formula, see item 8).
**Handling:** for RAG purposes, prefer **not** to chunk this table as
literal text at all; either (a) reconstruct it properly using per-number
bounding-box position matched against header row/column bboxes (real 2D
table reconstruction, non-trivial), or (b) rely on the accompanying formula
being available for the LLM to compute from directly, and explicitly flag
this table's raw text as unreliable/do-not-cite in metadata.
**Generalizes:** yes — any nomogram, nutrition-fact grid, or nCk-style
lookup table in any PDF has this exact problem; detect by the "bare number
grid" pattern, don't assume normal table extraction works.
---
## Formula / equation risks
### 8. Formula rendering is inconsistent — some survive as linear text, some don't
**What it looks like:** two real formulas found, two different outcomes.
The Du Bois body-surface-area formula (simple inline exponents,
`"S = W0,425 × H0,725 × 71,84"`) extracted **cleanly as readable text**. The
Cockcroft-Gault creatinine-clearance formula (a stacked fraction —
numerator over denominator, visually 2D) extracted as **scattered,
disordered fragments** with no linear reading order.
**Why it matters:** it's tempting to write one rule ("formulas are
unreliable, always flag them") or its opposite ("formulas extract fine, no
special handling needed") — neither is true here. The determining factor is
whether the formula's visual layout is fundamentally 1D (left-to-right,
like an inline exponent) or 2D (a fraction, a matrix, stacked terms).
**Check:** no cheap automatic detector was built for this distinction yet —
treat any equation/formula-like content as a manual-review candidate,
especially anything with a fraction bar, until a real detector exists
(e.g. checking for large vertical bbox gaps between adjacent glyphs that
should be visually stacked).
**Generalizes:** yes — any technical/medical/scientific PDF with inline
math will have this exact split; don't assume all formulas behave the same
way in extraction.
---
## Character/glyph-level risks
### 9. Rare reversed (right-to-left) glyph-order defect
**What it looks like:** confirmed exactly once across the entire
1668-page book (physical page 1373): one short text run's glyphs are
positioned in **descending x-order** rather than ascending, producing
scrambled output (e.g. `" = tịx 8 yàgn gnàh uềil gnổt(..."`) that reverses
character-by-character back to the correct Vietnamese sentence
(`"(4 xịt = 800 microgam) vào buổi chiều..."`).
**Why it matters:** this is a genuine, confirmed data-corruption risk, not
theoretical — but it's also extremely rare (1 occurrence in 1668 pages), so
it must be *detected*, not assumed to be either absent or common.
**Check:** group text fragments into visual rows by rounded y-coordinate,
then check whether x-coordinates are non-decreasing across the row; flag
(and optionally auto-correct by re-sorting on x) any row that isn't. This
full-book check runs in about 20 seconds.
**Generalizes:** yes, directly — this is a cheap, universal sanity check
worth running on any PDF text-extraction pipeline as a standing QA gate,
regardless of source document, since it catches a class of PDF-authoring
defects (RTL/BiDi overrides, corrupted content streams) that have nothing
to do with this book specifically.
---
## Section/heading detection risks
### 10. Font size is not a reliable heading signal — bold is
**What it looks like:** confirmed two genuine, equally top-level monograph
titles at different font sizes (10.0pt and 9.5pt). An early detector
gated on `size >= 9.8` and silently dropped ~15% of real monographs as a
result.
**Why it matters:** a threshold calibrated from one or two examples will
look correct until validated at scale — this is the single clearest
"don't generalize from a small sample" lesson from this whole
investigation.
**Check:** whole-document validation against an independent ground truth
(here, the back-of-book page-numbered index) is what caught this — a
sample of 2-3 pages would not have.
**Generalizes:** yes — for any PDF, prefer a binary style signal (bold/not
bold, a specific font name) over a numeric threshold (size, weight value)
wherever possible, and always validate any numeric threshold against the
whole document, not a handful of examples.
### 11. Multi-line wrapped titles/headings must be merged before matching
**What it looks like:** confirmed as the dominant cause of missed
detections in whole-document validation — long titles (e.g. "CÁC CHẤT ỨC
CHẾ HMG-CoA REDUCTASE", "THUỐC TƯƠNG TỰ HORMON GIẢI PHÓNG GONADOTROPIN")
wrap across 2+ physical lines; a per-line detector catches only fragments,
which then fail to match a name-based ground truth AND can produce false
name collisions with an unrelated single-line heading elsewhere in the
document (this happened: a wrapped title's second line, "GONADOTROPIN",
collided with a genuine, different, single-line "GONADOTROPIN" monograph
elsewhere).
**Check:** whole-document recall measurement against ground truth; misses
clustered around long/compound names are the signature of this bug.
**Handling:** merge consecutive bold+all-caps lines (with compatible
positioning) into one candidate title before matching/keying, rather than
treating each line independently.
**Generalizes:** yes — any document with long titles/headings that can wrap
will have this exact failure mode; always merge candidate multi-line
headings before using them as unique keys.
### 12a. Class-level monographs cover multiple active ingredients (multiple ATC codes) — this is NOT rare
**What it looks like:** first noticed via two incidental examples
("GONADOTROPIN", "VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ"), then actually measured
across the whole 680-monograph corpus (not assumed from the 2 examples —
this distinction matters, see below). **Real, whole-corpus number: 173 of
680 detected monographs (25.4%) have more than one distinct ATC code**,
ranging up to extreme cases — INSULIN alone lists **20** different ATC
codes, BETAMETHASON and DEXAMETHASON 11 each, PREDNISOLON 10,
HYDROCORTISON 9. This is a quarter of the entire corpus, not a couple of
edge cases — the 2 incidental examples badly understated how common this
is, and stating "found 2 examples, pattern confirmed" without the
whole-corpus count would have been exactly the kind of unverified claim
this project's CLAUDE.md now forbids.
**Even the 25.4% is a floor, not the true number** — see item 12c below:
ATC-code text-extraction noise (stray whitespace, O/0 confusion) caused
some genuinely multi-ATC monographs (e.g. "TRIAMCINOLON", 5 codes) to be
undercounted by a naive regex. The true proportion is measurably higher
than 25.4%; re-measure after fixing the regex, don't keep citing 25.4% as
final.
**Why it matters:** a data model that assumes "one monograph = one drug =
one ATC code" is wrong for roughly a quarter or more of the corpus.
**Handling:** store ATC code (and dosage-form sub-entries) as a **list**
per monograph, not a scalar; when chunking, consider whether a
class-level monograph's sections should be tagged with the whole class
name, the specific sub-compound, or both, depending on what the retrieval
use case needs.
**Generalizes:** yes — any reference work organized primarily by drug
class or by generic substance will have entries that don't map 1:1 to a
single identifier. More importantly, the *methodology* generalizes: when
you notice a pattern from 1-2 examples, measure its real prevalence across
the whole corpus before deciding how much engineering effort it deserves —
"found 2 examples" and "25.4% of everything" call for very different
levels of investment, and you can't tell which one you're dealing with
without the whole-corpus count.
### 12c. ATC codes (and likely other structured codes) have real text-extraction noise
**What it looks like:** while investigating why 22/680 (3.2%) monographs
appeared to have zero ATC codes, spot-checked 14 of them directly and found
**two distinct, confirmed causes**, both text-extraction noise rather than
missing content:
- **Stray internal whitespace** splitting one code into two tokens, e.g.
`"L01X X02"` (should be `L01XX02`), `"J04A C01"` (should be `J04AC01`),
`"N05B A06"` (should be `N05BA06`).
- **Digit/letter confusion**: a literal "0" rendered/typeset as the letter
"O", e.g. `"NO3AX12"` (should be `N03AX12`), `"JO1DC07"` (should be
`J01DC07`).
A relaxed regex tolerating both patterns resolved **9 of the 14** spot-checked
cases as real ATC codes hiding behind extraction noise. The **remaining
~5 of 14** were genuinely different: the source text explicitly states
`"Mã ATC: Chưa có."` or `"Mã ATC: Không có."` ("not yet available" / "none")
— a real, valid data state, not an error, and not something to paper over
as if a code exists.
**Why it matters:** a strict ATC-code regex silently undercounts real ATC
data; distinguishing "extraction noise hiding a real code" from "the book
says there is no code" requires checking the actual field text, not just
whether a regex matched.
**Handling:** normalize ATC-code-shaped text before matching (strip internal
whitespace between the letter/digit groups, treat a digit-position "O" as
"0") and explicitly check for the "Chưa có"/"Không có" literal strings as a
valid "no ATC" state rather than a parse failure.
**Generalizes:** yes — any structured code/identifier extracted from a PDF
(product codes, classification codes, reference numbers) can suffer this
same whitespace-injection and O/0 confusion; validate structured-looking
fields against their expected format and investigate exceptions rather than
assuming a strict pattern match is reliable.
### 12d. A section-title (part-divider) page can be falsely detected as a monograph
**What it looks like:** confirmed — the very first item in a whole-corpus
boundary scan was "CÁC CHUYÊN LUẬN THUỐC" (the literal title of Part 2 of
the book, "The Drug Monographs" — a part-divider heading, not a drug) at
physical page 98, picked up as a false-positive monograph boundary because
it happened to be bold, all-caps, short, and was followed (a few real
monograph-boundaries later) by some "Tên chung quốc tế" text from the
actual first real monograph.
**Why it matters:** without a whole-corpus scan this would have gone
unnoticed indefinitely — it doesn't look wrong from a single-page read of
Abacavir, and the discovery methodology this catalog is built on is
exhaustive scans, so this is a good example of a defect that only surfaces
at full scale.
**Handling:** exclude a small, known set of non-drug part/section-divider
strings ("CÁC CHUYÊN LUẬN THUỐC", "CÁC CHUYÊN LUẬN CHUNG", "CÁC PHỤ LỤC",
etc. — enumerable from the book's own table of contents) from the
monograph-boundary detector, or require the anchor phrase ("Tên chung quốc
tế") within a tighter line-distance so an unrelated real monograph several
lines away doesn't false-confirm a divider title.
**Generalizes:** yes — any document with part/section-divider title pages
styled similarly to its content headings (bold, prominent, short) risks
this exact false positive; explicitly exclude known structural/navigational
titles from content-boundary detectors.
### 12b. Genuine spelling/capitalization typos exist in the source text
**What it looks like:** confirmed real example — the running header on the
Vitamin D monograph's continuation pages reads `"Vitamin d và các thuốc
tương tự"` (lowercase "d"), while the real ALL-CAPS heading correctly reads
`"VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ"`. This is a genuine typesetting mistake
in the 2018 print, confirmed via font/bbox inspection (same bold 10pt font
as the correct heading — not an extraction artifact, the source text itself
has the typo). The page's bottom running *footer* uses yet another variant,
the short form `"Vitamin D"` (correctly capitalized) — meaning the same
monograph has **three different boilerplate text variants** across one
page (top header with a typo, the real heading, bottom footer).
**Why it matters:** don't treat running headers/footers as a perfectly
clean, typo-free secondary signal (item 13 in this catalog already
recommends using them as a cross-check) — they can themselves contain
source-level errors. In this specific case, the detection heuristic
(strict ALL-CAPS requirement, item 10) happened to still work correctly,
because "Vitamin d và các thuốc tương tự" and "Vitamin D" are not fully
uppercase and so are correctly rejected as monograph-boundary candidates —
but this was not a designed defense against typos specifically, just a
side effect of the all-caps requirement. A future/different typo (e.g. an
accidentally all-caps running header) would not be caught the same way.
**Check:** no systematic typo-detection was built (out of scope — this is
about parsing robustness, not proofreading the source); the practical
takeaway is to keep relying on the strict structural signals (bold + all
caps + short + anchor phrase) as primary, and treat any single text-based
signal (including running headers) as fallible.
**Generalizes:** yes — any real-world print-to-PDF source will have some
rate of genuine typos/inconsistencies; parsing logic should be robust to
them by relying on multiple independent structural signals (font,
position, anchor phrases) rather than trusting any single text match to be
error-free.
### 12e. Monograph length and section coverage vary enormously — measured, not assumed
**What it looks like:** across all 680 detected monographs, length ranges
from **2,331 to 45,623 characters** (~20x spread) and the number of known
section labels found per monograph ranges from as few as **8** up to
**20** (out of a ~19-20 item known vocabulary) — most cluster around
16-19, but the tails are real: "ASPARAGINASE"-adjacent short entries around
2,300-4,300 chars vs. "AMOXICILIN VÀ KALI CLAVULANAT" at 45,623 chars.
**Why it matters:** don't design chunking limits (e.g. a fixed max tokens
per monograph, or an assumption that "a monograph roughly fits in N
chunks") around a single example — the real distribution has a long tail
on both ends.
**Check:** this came from the same whole-corpus survey used for items 12a
and 12c — computing length and detected-section-count per monograph is
cheap and worth keeping as a standing sanity metric (e.g. flag any
monograph outside some percentile range for manual review).
**Generalizes:** yes — any corpus of "similar" documents (monographs,
product entries, articles) will have a real length/completeness
distribution; measure it before assuming uniformity.
### 12. The documented taxonomy is not exhaustive — keep it open
**What it looks like:** the book explicitly documents a 19-field template
for every drug monograph (page 38), but real monographs contain at least
one undocumented extra field ("Tên thương mại" — brand/trade names) not in
that list.
**Why it matters:** treating a documented schema as a closed enum will
silently misclassify or drop real content that doesn't fit it.
**Generalizes:** yes — any document that describes its own structure in a
preface/README should still be validated against real instances; documented
schemas are frequently incomplete in practice.
---
## Noise / boilerplate risks
### 13. Header/footer boilerplate must be stripped, but can double as a signal
**What it looks like:** every page carries a page number and a repeating
string (`"DTQGVN 2"`), and body pages additionally carry a running header
naming the current monograph/section.
**Handling:** strip the fixed boilerplate before parsing content, but the
running monograph-name header is a **useful secondary cross-check** for
"which monograph is this page's body text currently part of" — don't
discard it as pure noise.
**Generalizes:** yes — running headers/footers are common in print-derived
PDFs and are usually worth extracting as metadata, not just filtering out.
### 14. Blank/near-empty separator pages at section transitions are expected
**What it looks like:** exactly 6 near-empty pages (<20 characters) found
across the whole 1668-page book, and every single one lands exactly on a
major section-transition boundary (before general chapters, before
individual monographs, before appendices, near the book's end).
**Why it matters:** a naive pipeline might treat a near-empty page as an
extraction failure and error out or flag it, when it's actually an
intentional print-layout convention (forcing a new part to start on a
fresh page).
**Check:** whole-document scan for pages under some small character
threshold; cross-reference their positions against known section
boundaries before treating them as errors.
**Generalizes:** yes — this print convention is extremely common in
formally typeset books; always expect and gracefully skip near-empty pages
rather than treating them as failures.
---
## Not yet investigated (flagged for future work, not silently ignored)
- **Footnote-style superscript reference markers** (seen as `a, b, c, d` in
one table) — not yet checked for whether the footnote text stays
correctly associated with its marker/row during extraction.
- **Formula detection heuristic** (item 8) — no automatic detector exists
yet to flag 2D-formula regions before they're trusted as chunk content.
- **2D grid table reconstruction** (item 7) — no implementation yet for
recovering row/column-correct values from a nomogram-style table.