51 KiB
PDF Parsing Outlier Catalog
A generalized checklist of structural risks found while parsing
duoc-thu-quoc-gia-viet-nam-2018.pdf (1668 pages). Every item here was
confirmed with real evidence (bounding-box inspection, cross-tool
comparison, or a whole-document scan) — not assumed. The goal of this
document is reuse: if this project (or a future one) needs to parse another
structured reference PDF — another national formulary, a different
government-published multi-part document, any dense print-layout book —
this is the checklist of "things that go wrong that a small page sample
won't reveal," and how to actually check for each one cheaply (most checks
here run over the whole 1668-page book in under a minute).
For the narrative investigation and drug-formulary-specific numbers, see
docs/adr/0003-pdf-parsing-strategy.md. This document is the distilled,
reusable checklist form of the same findings, plus items found afterward.
Structural discovery risks (before you even parse content)
1. No bookmarks/TOC
What it looks like: doc.get_toc() (PyMuPDF) returns an empty list.
Why it matters: the obvious, easiest structural signal for section
boundaries simply doesn't exist — don't design a pipeline that assumes it
will.
Check: one line, len(doc.get_toc()). Do this first, always, before
assuming a bookmark-based approach.
Generalizes: yes, directly — always check this before designing around
bookmarks, for any PDF.
2. Shallow/unusable tagged-PDF structure tree
What it looks like: the PDF has a /StructTreeRoot (looks promising —
"tagged PDF"), but it only covers a handful of generic /H1//P elements
for a fraction of the document (here: ~29 elements for 1668 pages).
Why it matters: easy to assume "tagged PDF = rich semantic structure
available"; in practice many tagging tools produce a minimal
compliance-only tree that covers almost nothing.
Check: walk the struct tree (doc.xref_object on /StructTreeRoot,
recurse into /K) and count real leaf elements vs. total page count. If the
ratio is tiny, it's not a usable data source.
Generalizes: yes — always verify depth/coverage before trusting a
struct tree, don't just check for its existence.
Page layout risks
3. Multi-column body layout
What it looks like: body pages are genuinely two-column (confirmed via
bounding boxes: left column x≈44-299, right column x≈308-562, page width
≈595). Front-matter pages that look like a multi-column name grid to the
eye turned out, on inspection, to be single wide text blocks with internal
whitespace padding between names — not a real structural column split.
Why it matters: a naive "read text top-to-bottom regardless of x" pass
would interleave left- and right-column content into nonsense. Conversely,
assuming every visually grid-like page is column-split leads to wasted
effort — verify per page/section, don't generalize from appearance alone.
Check: for any suspicious page, dump block bounding boxes
(page.get_text("dict")["blocks"]) and look at the actual x0/x1 ranges. A
real column split shows two clusters of x-ranges; a padded single-column
list shows one wide range per line.
Handling: PyMuPDF's default block-level reading order handled the real
two-column case correctly here (validated against a known monograph) — the
tool most likely to get column order wrong was pdfplumber's general
extract_text() (see item 8), not PyMuPDF.
Generalizes: yes — this exact check (dump bboxes, look at x-clusters)
works on any PDF to determine real column count before writing extraction
logic.
4. Full-width content breaking out of the column grid
What it looks like: some pages have a table (or could have a figure) that spans nearly the entire page width (confirmed: a body-surface-area lookup table's blocks span x≈35 to x≈553, i.e. across both normal columns), overriding the page's usual two-column layout. Why it matters: logic written to always split a page into "left column" and "right column" text will misbehave on these pages — the content isn't in either column, it's a single full-width unit. Check: for any block, compare its x-width against the known single-column width; if a block's x-range spans (or nearly spans) both known column ranges, treat it as a full-width unit, not part of a column. Generalizes: yes — any multi-column layout can have occasional full-width breakout elements (tables, figures, pull-quotes); always check for this rather than assuming rigid column adherence everywhere.
Table-specific risks
5. Tables split across a page break lose their header on the continuation page
What it looks like: confirmed directly — "Bảng 4: Xử trí về điều trị ARV
theo mức độ phát ban" (a 3-column table) starts on one page with its header
row (['Mức độ', 'Biểu hiện', 'Xử trí']) and 3 data rows; its 4th data row
("Mức độ 4...") appears on the next page, extracted by pdfplumber
as a separate table object with no header row at all.
Why it matters: if a pipeline treats each find_tables()/
extract_tables() result as an independent, self-contained table, the
orphaned continuation row is meaningless on its own — you lose the column
semantics for that row entirely.
Check: for any table-like structure, check whether the page/column
immediately preceding it ends with a same-shaped table lacking a natural
final row (e.g. an incomplete-looking sequence) — a strong heuristic is
"table starts at the very top of a page/column, no header, same column
count as the table ending at the bottom of the previous page/column."
Handling: never treat page-extracted tables as independent; track
continuation explicitly and re-attach the original header to orphaned
continuation rows before using them.
Generalizes: yes — this is a generic multi-page-table risk in any
paginated PDF with tall tables; the detection heuristic (position at
page/column top + no header + matching column count to the previous
table) applies broadly.
6. Tables can also split across a column boundary on the same page
What it looks like: confirmed — "Bảng 6" (ARV drug toxicity table) starts in the left column near the bottom of a page (header + first data row) and its remaining data rows appear at the top of the right column of the very same page, again with no header repeated. Why it matters: this is easy to miss because there's no literal page break — it's tempting to assume "if it's the same page, it's not split," but a table can still be taller than one column's usable height. Check: same heuristic as item 5, but also check column position, not just page number — a header-less table fragment starting at the top of a column (regardless of page) is a suspect continuation. Generalizes: yes, wherever content flows in columns at all — this risk exists any time column height is shorter than table height.
7. Two-dimensional grid/nomogram tables are not linearly recoverable
What it looks like: confirmed — a body-surface-area lookup table
(height across the top, weight down the side, a BSA value at each
intersection) extracts as a scrambled sequence of numbers with no
recoverable row/column association from plain text alone (e.g. "0,50 0,52 0,54 0,56" followed by "0,55 0,57 0,59 0,61" — these are almost
certainly column-wise fragments, not the visual rows).
Why it matters: unlike a normal bordered table (rows of related
values), a 2D lookup grid's meaning depends entirely on 2D position — a
number is meaningless without knowing both its row header (weight) and
column header (height). Flattened text extraction destroys exactly the
information needed to interpret it.
Check: any table where extracted "cells" are bare numbers with no
inline label, laid out in a dense grid, is a candidate — cross-check
against the source's own stated formula/description (this table is
explicitly a lookup version of a stated formula, see item 8).
Handling: for RAG purposes, prefer not to chunk this table as
literal text at all; either (a) reconstruct it properly using per-number
bounding-box position matched against header row/column bboxes (real 2D
table reconstruction, non-trivial), or (b) rely on the accompanying formula
being available for the LLM to compute from directly, and explicitly flag
this table's raw text as unreliable/do-not-cite in metadata.
Generalizes: yes — any nomogram, nutrition-fact grid, or nCk-style
lookup table in any PDF has this exact problem; detect by the "bare number
grid" pattern, don't assume normal table extraction works.
Formula / equation risks
8. Formula rendering is inconsistent — some survive as linear text, some don't
What it looks like: two real formulas found, two different outcomes.
The Du Bois body-surface-area formula (simple inline exponents,
"S = W0,425 × H0,725 × 71,84") extracted cleanly as readable text. The
Cockcroft-Gault creatinine-clearance formula (a stacked fraction —
numerator over denominator, visually 2D) extracted as scattered,
disordered fragments with no linear reading order.
Why it matters: it's tempting to write one rule ("formulas are
unreliable, always flag them") or its opposite ("formulas extract fine, no
special handling needed") — neither is true here. The determining factor is
whether the formula's visual layout is fundamentally 1D (left-to-right,
like an inline exponent) or 2D (a fraction, a matrix, stacked terms).
Check: a detector now exists — residual_ink.py's
fraction_bar_candidate, which finds the bar as ink no extracted span
accounts for. Measured on this book: precision 16/23 = 69.6% (the misses
are decorative underlines and table borders), recall unknown, and it is blind
by construction to a fraction printed without a bar (item 25). Its output is
therefore a review queue, not a verdict: all 23 candidates were rendered and
read one at a time before any was acted on, and only the confirmed ones went
into ingestion/data/verified/formula_regions_2d.json.
Generalizes: yes — any technical/medical/scientific PDF with inline
math will have this exact split; don't assume all formulas behave the same
way in extraction.
Character/glyph-level risks
9. Rare reversed/misordered glyph defects — corrected count: 2, not 1
What it looks like: re-implemented as real, tested production code
(ingestion/ingestion/extract/glyph_order.py) rather than trusted from the
earlier exploratory script's claim. Found two distinct shapes, not the
one originally reported:
- Within-span character reversal (physical page 1373, the originally
reported case): one span's glyphs are positioned in descending x-order,
producing
" = tịx 8 yàgn gnàh uềil gnổt(...", which reverses character-by-character back to"(4 xịt = 800 microgam) vào buổi chiều...". - Cross-fragment row misordering, newly found (physical page 714): a
single visual row is split by PyMuPDF into multiple
lineobjects within one block that are then emitted out of left-to-right order — each fragment's own characters are fine, but concatenating fragments in extraction order produces"...bảo quản ộđ tệihn "instead of the correct"...bảo quản nhiệt độ ". This is a different underlying shape from item 1 (multiple mis-ordered fragments, not one reversed span) and was missed by the original narrower (within-span-only) check — the ADR 0003 claim of "exactly 1 occurrence in the whole book" undercounted the real defect population; corrected here. Getting a trustworthy count took three detector iterations (documented in the module's own docstring) — the first naive whole-book implementation of the row-level check reported 1113 "issues," almost all false positives from two mechanisms: (a) ordinary font-kerning jitter (e.g. in "mefloquin," two adjacent glyphs differ by 0.095pt — normal kerning, not a defect) treated as a reversal with no decrease-tolerance, which then actively corrupted correct text into "mefolquin"; and (b) reconstructing "visual rows" from raw x/y coordinates using a hand-picked column-boundary threshold, which misclassifies a paragraph that happens to start near the natural column gap (confirmed real case: a right-column paragraph starting at x=299.4 got merged with an unrelated left-column paragraph at the same y). The fix that survived whole-book testing: group by PyMuPDF's ownblockindex (already validated in ADR 0003 to respect this document's column structure) instead of re-deriving columns from coordinates, plus a minimum-decrease threshold (1.0pt — safely between the ~0.3pt kerning noise floor and the >2pt real-defect magnitude). Final whole-book result: 11 row-level issues on 5 pages — 3 of those pages (92, 94, 805) are formula regions already flagged as unreliable in item 8 below (2D-layout formulas scramble on extraction; this check's "corrected" text for those rows should not be trusted or auto-applied, same as item 8's existing guidance), leaving exactly the 2 genuine prose defects above (pages 714, 1373). Why it matters: both genuine defects are confirmed real data-corruption risks, not theoretical — but both are also extremely rare (2 occurrences in 1668 pages of prose), so they must be detected, not assumed either absent or common. Equally important: a naive implementation of "the obvious check" can itself introduce false positives and even actively corrupt correct text — this detector's own false-positive history is as important a lesson as the defects it catches. Check:ingestion.extract.scan_glyph_order(within-span) andingestion.extract.scan_reading_order(cross-fragment, grouped by real PyMuPDF block index + row y, with a 1.0pt minimum-decrease threshold and header-band exclusion). Both run in seconds over the full book. Generalizes: yes, directly — this is a cheap, universal sanity check worth running on any PDF text-extraction pipeline as a standing QA gate, regardless of source document. The false-positive history also generalizes: any "reconstruct visual rows from raw coordinates" approach needs a decrease-tolerance (font kerning is universal) and should prefer the source tool's own layout-analysis groupings (blocks/lines) over hand-picked coordinate thresholds wherever available.
Section/heading detection risks
10. Font size is not a reliable heading signal — bold is
What it looks like: confirmed two genuine, equally top-level monograph
titles at different font sizes (10.0pt and 9.5pt). An early detector
gated on size >= 9.8 and silently dropped ~15% of real monographs as a
result.
Why it matters: a threshold calibrated from one or two examples will
look correct until validated at scale — this is the single clearest
"don't generalize from a small sample" lesson from this whole
investigation.
Check: whole-document validation against an independent ground truth
(here, the back-of-book page-numbered index) is what caught this — a
sample of 2-3 pages would not have.
Generalizes: yes — for any PDF, prefer a binary style signal (bold/not
bold, a specific font name) over a numeric threshold (size, weight value)
wherever possible, and always validate any numeric threshold against the
whole document, not a handful of examples.
11. Multi-line wrapped titles/headings must be merged before matching
What it looks like: confirmed as the dominant cause of missed detections in whole-document validation — long titles (e.g. "CÁC CHẤT ỨC CHẾ HMG-CoA REDUCTASE", "THUỐC TƯƠNG TỰ HORMON GIẢI PHÓNG GONADOTROPIN") wrap across 2+ physical lines; a per-line detector catches only fragments, which then fail to match a name-based ground truth AND can produce false name collisions with an unrelated single-line heading elsewhere in the document (this happened: a wrapped title's second line, "GONADOTROPIN", collided with a genuine, different, single-line "GONADOTROPIN" monograph elsewhere). Check: whole-document recall measurement against ground truth; misses clustered around long/compound names are the signature of this bug. Handling: merge consecutive bold+all-caps lines (with compatible positioning) into one candidate title before matching/keying, rather than treating each line independently. Generalizes: yes — any document with long titles/headings that can wrap will have this exact failure mode; always merge candidate multi-line headings before using them as unique keys.
12a. Class-level monographs cover multiple active ingredients (multiple ATC codes) — this is NOT rare
What it looks like: first noticed via two incidental examples ("GONADOTROPIN", "VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ"), then actually measured across the whole 680-monograph corpus (not assumed from the 2 examples — this distinction matters, see below). Real, whole-corpus number: 173 of 680 detected monographs (25.4%) have more than one distinct ATC code, ranging up to extreme cases — INSULIN alone lists 20 different ATC codes, BETAMETHASON and DEXAMETHASON 11 each, PREDNISOLON 10, HYDROCORTISON 9. This is a quarter of the entire corpus, not a couple of edge cases — the 2 incidental examples badly understated how common this is, and stating "found 2 examples, pattern confirmed" without the whole-corpus count would have been exactly the kind of unverified claim this project's CLAUDE.md now forbids. Even the 25.4% is a floor, not the true number — see item 12c below: ATC-code text-extraction noise (stray whitespace, O/0 confusion) caused some genuinely multi-ATC monographs (e.g. "TRIAMCINOLON", 5 codes) to be undercounted by a naive regex. The true proportion is measurably higher than 25.4%; re-measure after fixing the regex, don't keep citing 25.4% as final. Why it matters: a data model that assumes "one monograph = one drug = one ATC code" is wrong for roughly a quarter or more of the corpus. Handling: store ATC code (and dosage-form sub-entries) as a list per monograph, not a scalar; when chunking, consider whether a class-level monograph's sections should be tagged with the whole class name, the specific sub-compound, or both, depending on what the retrieval use case needs. Generalizes: yes — any reference work organized primarily by drug class or by generic substance will have entries that don't map 1:1 to a single identifier. More importantly, the methodology generalizes: when you notice a pattern from 1-2 examples, measure its real prevalence across the whole corpus before deciding how much engineering effort it deserves — "found 2 examples" and "25.4% of everything" call for very different levels of investment, and you can't tell which one you're dealing with without the whole-corpus count.
12c. ATC codes (and likely other structured codes) have real text-extraction noise
What it looks like: while investigating why 22/680 (3.2%) monographs appeared to have zero ATC codes, spot-checked 14 of them directly and found two distinct, confirmed causes, both text-extraction noise rather than missing content:
- Stray internal whitespace splitting one code into two tokens, e.g.
"L01X X02"(should beL01XX02),"J04A C01"(should beJ04AC01),"N05B A06"(should beN05BA06). - Digit/letter confusion: a literal "0" rendered/typeset as the letter
"O", e.g.
"NO3AX12"(should beN03AX12),"JO1DC07"(should beJ01DC07). A relaxed regex tolerating both patterns resolved 9 of the 14 spot-checked cases as real ATC codes hiding behind extraction noise. The remaining ~5 of 14 were genuinely different: the source text explicitly states"Mã ATC: Chưa có."or"Mã ATC: Không có."("not yet available" / "none") — a real, valid data state, not an error, and not something to paper over as if a code exists. Why it matters: a strict ATC-code regex silently undercounts real ATC data; distinguishing "extraction noise hiding a real code" from "the book says there is no code" requires checking the actual field text, not just whether a regex matched. Handling: normalize ATC-code-shaped text before matching (strip internal whitespace between the letter/digit groups, treat a digit-position "O" as "0") and explicitly check for the "Chưa có"/"Không có" literal strings as a valid "no ATC" state rather than a parse failure. Generalizes: yes — any structured code/identifier extracted from a PDF (product codes, classification codes, reference numbers) can suffer this same whitespace-injection and O/0 confusion; validate structured-looking fields against their expected format and investigate exceptions rather than assuming a strict pattern match is reliable.
12d. A section-title (part-divider) page can be falsely detected as a monograph
What it looks like: confirmed — the very first item in a whole-corpus boundary scan was "CÁC CHUYÊN LUẬN THUỐC" (the literal title of Part 2 of the book, "The Drug Monographs" — a part-divider heading, not a drug) at physical page 98, picked up as a false-positive monograph boundary because it happened to be bold, all-caps, short, and was followed (a few real monograph-boundaries later) by some "Tên chung quốc tế" text from the actual first real monograph. Why it matters: without a whole-corpus scan this would have gone unnoticed indefinitely — it doesn't look wrong from a single-page read of Abacavir, and the discovery methodology this catalog is built on is exhaustive scans, so this is a good example of a defect that only surfaces at full scale. Handling: exclude a small, known set of non-drug part/section-divider strings ("CÁC CHUYÊN LUẬN THUỐC", "CÁC CHUYÊN LUẬN CHUNG", "CÁC PHỤ LỤC", etc. — enumerable from the book's own table of contents) from the monograph-boundary detector, or require the anchor phrase ("Tên chung quốc tế") within a tighter line-distance so an unrelated real monograph several lines away doesn't false-confirm a divider title. Generalizes: yes — any document with part/section-divider title pages styled similarly to its content headings (bold, prominent, short) risks this exact false positive; explicitly exclude known structural/navigational titles from content-boundary detectors.
12b. Genuine spelling/capitalization typos exist in the source text
What it looks like: confirmed real example — the running header on the
Vitamin D monograph's continuation pages reads "Vitamin d và các thuốc tương tự" (lowercase "d"), while the real ALL-CAPS heading correctly reads
"VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ". This is a genuine typesetting mistake
in the 2018 print, confirmed via font/bbox inspection (same bold 10pt font
as the correct heading — not an extraction artifact, the source text itself
has the typo). The page's bottom running footer uses yet another variant,
the short form "Vitamin D" (correctly capitalized) — meaning the same
monograph has three different boilerplate text variants across one
page (top header with a typo, the real heading, bottom footer).
Why it matters: don't treat running headers/footers as a perfectly
clean, typo-free secondary signal (item 13 in this catalog already
recommends using them as a cross-check) — they can themselves contain
source-level errors. In this specific case, the detection heuristic
(strict ALL-CAPS requirement, item 10) happened to still work correctly,
because "Vitamin d và các thuốc tương tự" and "Vitamin D" are not fully
uppercase and so are correctly rejected as monograph-boundary candidates —
but this was not a designed defense against typos specifically, just a
side effect of the all-caps requirement. A future/different typo (e.g. an
accidentally all-caps running header) would not be caught the same way.
Check: no systematic typo-detection was built (out of scope — this is
about parsing robustness, not proofreading the source); the practical
takeaway is to keep relying on the strict structural signals (bold + all
caps + short + anchor phrase) as primary, and treat any single text-based
signal (including running headers) as fallible.
Generalizes: yes — any real-world print-to-PDF source will have some
rate of genuine typos/inconsistencies; parsing logic should be robust to
them by relying on multiple independent structural signals (font,
position, anchor phrases) rather than trusting any single text match to be
error-free.
12e. Monograph length and section coverage vary enormously — measured, not assumed
What it looks like: across all 680 detected monographs, length ranges from 2,331 to 45,623 characters (~20x spread) and the number of known section labels found per monograph ranges from as few as 8 up to 20 (out of a ~19-20 item known vocabulary) — most cluster around 16-19, but the tails are real: "ASPARAGINASE"-adjacent short entries around 2,300-4,300 chars vs. "AMOXICILIN VÀ KALI CLAVULANAT" at 45,623 chars. Why it matters: don't design chunking limits (e.g. a fixed max tokens per monograph, or an assumption that "a monograph roughly fits in N chunks") around a single example — the real distribution has a long tail on both ends. Check: this came from the same whole-corpus survey used for items 12a and 12c — computing length and detected-section-count per monograph is cheap and worth keeping as a standing sanity metric (e.g. flag any monograph outside some percentile range for manual review). Generalizes: yes — any corpus of "similar" documents (monographs, product entries, articles) will have a real length/completeness distribution; measure it before assuming uniformity.
12. The documented taxonomy is not exhaustive — keep it open
What it looks like: the book explicitly documents a 19-field template for every drug monograph (page 38), but real monographs contain at least one undocumented extra field ("Tên thương mại" — brand/trade names) not in that list. Why it matters: treating a documented schema as a closed enum will silently misclassify or drop real content that doesn't fit it. Generalizes: yes — any document that describes its own structure in a preface/README should still be validated against real instances; documented schemas are frequently incomplete in practice.
Noise / boilerplate risks
13. Header/footer boilerplate must be stripped, but can double as a signal
What it looks like: every page carries a page number and a repeating
string ("DTQGVN 2"), and body pages additionally carry a running header
naming the current monograph/section.
Handling: strip the fixed boilerplate before parsing content, but the
running monograph-name header is a useful secondary cross-check for
"which monograph is this page's body text currently part of" — don't
discard it as pure noise.
Generalizes: yes — running headers/footers are common in print-derived
PDFs and are usually worth extracting as metadata, not just filtering out.
14. Blank/near-empty separator pages at section transitions are expected
What it looks like: exactly 6 near-empty pages (<20 characters) found across the whole 1668-page book, and every single one lands exactly on a major section-transition boundary (before general chapters, before individual monographs, before appendices, near the book's end). Why it matters: a naive pipeline might treat a near-empty page as an extraction failure and error out or flag it, when it's actually an intentional print-layout convention (forcing a new part to start on a fresh page). Check: whole-document scan for pages under some small character threshold; cross-reference their positions against known section boundaries before treating them as errors. Generalizes: yes — this print convention is extremely common in formally typeset books; always expect and gracefully skip near-empty pages rather than treating them as failures.
15. No embedded images anywhere in the book — measured, not assumed
What it looks like: a whole-book scan of page.get_images(full=True) across
all 1668 pages returns zero embedded raster/vector images, confirmed via
PyMuPDF's own image extraction API (not just "the text doesn't mention an
image").
Why it matters: avoids over-investing in image/caption validation tooling
for a corpus that has no images to validate — but this must be a measured
fact, not an assumption from the book's general description as "text-heavy."
Generalizes: the check (get_images(full=True) summed over every page)
is a cheap one-line whole-document verification worth running on any PDF
before deciding whether image-handling code is needed at all.
16. Chemical reaction arrows render as Private-Use-Area glyphs, not Unicode arrows
What it looks like: confirmed real example — physical page 1033 contains a
genuine chemical reaction equation (Na2S2O3 + CN⁻ → SCN⁻ + Na2SO3, part of
the cyanide-antidote/rhodanese mechanism description). The reaction arrow
extracts as a Private-Use-Area codepoint (), not a standard Unicode
arrow (→) — the source PDF's font maps a custom symbol glyph (likely from a
symbol/wingdings-style embedded font) into a PUA slot, and raw text extraction
faithfully returns that codepoint rather than a human-readable arrow.
Why it matters: any pipeline that treats extracted text as directly
human-readable/citable will surface a mangled or invisible character where a
reaction arrow should be; a naive keyword/embedding step over raw text would
either silently drop it (if PUA codepoints get filtered as junk) or leave a
confusing tofu/box character in a chunk shown to a doctor or pharmacist.
Check: scan extracted text for codepoints in the Unicode Private Use Area
ranges (U+E000–U+F8FF) — cheap and generalizes to any custom-glyph symbol
substitution, not just arrows.
Handling: for now, flag any monograph/section containing a PUA codepoint
for manual review or map known PUA codepoints (e.g. this book's →
→) via an explicit substitution table; do not pass raw PUA codepoints
through to chunking/embedding untranslated.
Generalizes: yes — any PDF built from print-authoring software that uses a
symbol font for arrows/special glyphs (common in scientific/medical/chemistry
documents) can exhibit this; always check for PUA codepoints in extracted
text as a standing sanity check, not just assume standard Unicode symbols.
Confirmed real chemical formula in the corpus, but rare: a regex scan for
molecular-formula-shaped tokens ([A-Z][a-z]?\d{1,3} repeated) across the
monograph page range found 9 raw hits; manual inspection found most are
false positives (H5N1 = flu strain name, P2Y12 = a receptor name, not
molecular formulas) and only one confirmed genuine chemical formula/equation
(the Na2S2O3 case above) — real chemical notation exists in this corpus but
is genuinely rare, not a systemic pattern requiring a general chemistry
parser.
17. Adult/child dosing-population splits are the norm, not an edge case
What it looks like: measured via a whole-monograph-range text scan for "Người lớn"/"Trẻ em"/"Trẻ sơ sinh" (adult/child/newborn) — these terms appear on 1121 of ~1400 monograph-range pages, i.e. the large majority of drug monographs split dosing by patient population. Why it matters: this is exactly the kind of structural content where a segmentation/chunking bug that interleaves or merges adjacent subsections (e.g. a table/list continuation bug, see items 5-6) would be a genuine patient-safety risk, not just a data-quality nicety — mixing an adult dose into a child-dose chunk (or vice versa) is a plausible, concrete failure mode given how common this structure is. Handling: treat "does this monograph's dosing section correctly keep adult/child/newborn subsections un-interleaved" as a standing validation check (not a rare-case afterthought), given the measured prevalence. Generalizes: yes — any clinical/pharmacological reference document organized with population-specific subsections has this same risk profile; measure real prevalence before deciding how much validation effort a structural risk deserves (same methodology lesson as item 12a).
18. A monograph title can legitimately repeat — disambiguated by a bold, non-caps qualifier line
What it looks like: confirmed real example, found while smoke-testing
the real segment/detector.py against the full book: "SALBUTAMOL" is
detected as a monograph title twice (physical pages 1261 and 1263).
Rendering both pages to images and reading them directly (not inferred from
coordinates) confirmed these are two genuinely different, complete
monographs — "SALBUTAMOL (Dùng trong hô hấp)" (respiratory use) and
"SALBUTAMOL (Dùng trong sản khoa)" (obstetric/tocolytic use) — each with
its own full 18-section template. The qualifier ("(Dùng trong hô hấp)" /
"(Dùng trong sản khoa)") is a bold line immediately below the all-caps
title, but is not itself all-caps (mixed case inside the parens), so it
is correctly excluded from detect_monograph_titles's all-caps candidate
filter — it must instead be captured as a separate signal and folded into
the monograph's disambiguating identity downstream.
Why it matters: an assembler that derives drug_id from the title text
alone (e.g. a simple slug of "SALBUTAMOL") will produce a real collision
between two legitimately different monographs — this is not the same
failure mode as the already-fixed GONADOTROPIN false-collision (that one
was a detector artifact from unmerged multi-line wrapping; this one is a
genuine same-name-different-monograph case that must be preserved, not
merged away).
Handling (for Phase 1.3's assembler): after detecting a monograph title,
check for an immediately-following bold, parenthesized, non-all-caps line
directly below it (same page, small y-gap) and include it in drug_id
generation when present, so "salbutamol_ho_hap" and "salbutamol_san_khoa"
remain distinct rather than colliding as "salbutamol" twice. The
assembler.py duplicate-drug_id check (outlier-catalog reasoning already
established: raise on a genuine duplicate rather than silently overwriting)
must be designed with this real case in mind, or it will incorrectly reject
a legitimate second "SALBUTAMOL" entry.
Generalizes: yes — any drug/entity reference work that documents the
same base substance under multiple distinct use-contexts (formulation,
indication, route) can have this exact pattern; never assume a title string
alone is a unique key without checking for a disambiguating qualifier line.
19. Table column headers can be bold + all-caps + short — identical shape to a real title
What it looks like: confirmed real example, found via a whole-book
assemble() run raising a duplicate-drug_id error: "HSV" and "CMV" each
appear twice as bold, all-caps, short (3-char) spans on physical page 698 —
not drug names at all, but column headers in a dosing-by-renal-function
table inside the "Foscarnet natri" monograph ("Liều đối với HSV / HSV /
CMV / CMV"). Bold+all-caps+short is exactly the monograph-title signal
(item 10/12d), so this is a genuine detector ambiguity, not a coding bug.
Why it matters: unlike item 12d's part-divider titles (a small,
enumerable, fixed set of known strings), a table's column headers are
unbounded and content-dependent (any future table could use "HSV", "CMV",
or something else entirely as a header) — an exclusion list approach
doesn't generalize here the way it did for part-dividers.
Handling: require a structural anchor rather than a text exclusion
list: a real monograph title is always followed shortly by at least one
recognized section heading from the vocabulary (in practice, always "Tên
chung quốc tế" first) before the next title-shaped candidate. A
table-header false positive is not — the table's own cells are numbers/
plain text, matching no vocabulary entry. Implemented as
assembler._filter_false_positive_titles (lookahead of 6 events, checked
against the same coalesced event stream already built for assembly — no
separate detection pass, no duplicated logic).
Generalizes: yes, more broadly than item 12d — any document where
section/entity boundaries are marked by a shape (bold+caps+short) that a
table, list, or figure caption could coincidentally also match should
verify a structural follow-on anchor, not just a shape match or a
denylist of known bad strings, since the space of possible false-shaped
content (table headers, figure labels, pull-quotes) is unbounded while the
space of "what a real boundary is followed by" is small and known.
20. Section headings are not consistently bold across monographs — some combine label+value in one plain span
What it looks like: confirmed real example, found by investigating why
a whole-book assemble() run showed 48 monographs with zero ATC codes and
not stated-absent (far more than the ~13-14 the original spot-check
extrapolated). AMITRIPTYLIN's real "Mã ATC:" field is a single, plain
(non-bold) span containing the label AND value together: "Mã ATC: N06AA09." — unlike Abacavir's equivalent, which is a bold "Mã ATC: "
label span followed by a separate plain "J05AF06." value span. Both
render visually similar but have completely different span/style
structure. Given the book's own foreword states it was "biên soạn bởi
nhiều tác giả" (written by many authors), this kind of per-author styling
inconsistency across ~700 individually-authored monographs is plausible
and, once checked, confirmed real — not a one-off.
Why it matters: a detector that requires span.bold to recognize a
section heading (reasonable-looking given every title is confirmed bold)
silently drops entire sections for a meaningful fraction of the corpus —
this directly caused undercounted ATC codes (and, structurally, would
equally affect any other section) for monographs using this looser style.
Handling: match section headings by vocabulary text, not by
boldness — the same "don't gate on a styling attribute, only content is
reliable" lesson as item 10 (font size), now applied to boldness. Also
handle the "label + value combined in one span" shape explicitly (a prefix
match: does the span start with a known label followed by ":", with the
remainder treated as the section's inline value) rather than assuming
label and value are always separate spans.
Generalizes: yes — any print-authored reference work assembled from
many individual authors/editors over a long production process should
expect inconsistent low-level styling of nominally-identical structural
elements; verify a structural signal (styling) against the content it's
supposed to correlate with, across a large real sample, before trusting it
as a universal discriminator — the same methodology lesson as item 10,
found again independently here.
21. "All-caps" is not 100% reliable either — and a class-level monograph's own internal sub-headings can masquerade as new monographs
What it looks like: two distinct confirmed real findings from the same investigation:
- The class-level monograph "CÁC CHẤT ỨC CHẾ HMG-CoA REDUCTASE" embeds the
mixed-case abbreviation "CoA" (Coenzyme A) inside an otherwise all-caps
title. A strict
text.isupper()check requires zero lowercase letters, so this single embedded abbreviation caused the entire monograph to be silently dropped from the corpus — found only by directly checking whether this specific, previously-known (outlier item 12a) class-level monograph was present in a real whole-bookassemble()run, and discovering it was not. - Within that same class-level monograph, individual statin names
("SIMVASTATIN", "LOVASTATIN", "PRAVASTATIN", "FLUVASTATIN") appear as
their own bold+all-caps+short sub-headings, each introducing its own
"Liều lượng và cách dùng" sub-section — shape-identical to a real
monograph title, and (after fix 1 above made the loosened "any known
section" anchor check pass) briefly became a second false-positive
category alongside item 19's table headers, since these sub-headings
are followed by a recognized section, just never by "Tên chung quốc
tế" specifically (that section belongs only to the parent).
Why it matters: together these show that neither "all-caps" nor "loosen
the anchor to any section" is safe in isolation — the fix for one false
positive (item 19, HSV/CMV) reopened a different one (SIMVASTATIN) until
the anchor check was tightened back to the specific section the book's
own template guarantees is always first for a genuine top-level monograph.
Handling:
detector._is_mostly_upperuses a lowercase-letter ratio (≤10%), not an absolute count — an earlier absolute-count version (≤2 lowercase letters) let a real regression through: "Mã ATC:" has only 1 lowercase letter (a normal Vietnamese diacritic, 'ã') but that's 20% of its 5 letters, correctly rejected by the ratio while HMG-CoA's 1/27 ≈ 3.7% correctly passes.assembler._has_anchor_aheadrequires specifically the "ten_chung_quoc_te" section key, not just any recognized section, since that is the one invariant the book's documented template actually guarantees is unique to real top-level monographs. Generalizes: yes — (1) don't assume a styling/casing convention holds with zero exceptions across an entire corpus, even one confirmed exception matters at whole-corpus scale; (2) when a document has nested substructure that mimics top-level structure (a class monograph containing per-item sub-entries), the anchor used to confirm a real boundary must be the most specific invariant available, not just "some known follow-on content" — a looser check that fixes one false positive can silently reopen another.
22. Running-header boilerplate was never actually stripped, despite item 13's warning — measured whole-corpus at 98.4% of monographs affected
What it looks like: the running header at the top of every physical page
("DTQGVN 2" + printed page number + the current monograph's name, e.g.
physical page 1008's "DTQGVN 2" / "1009" / "Morphin sulfat", tagged
column="full_width" by extract/spans.py) matches no section heading and
isn't a real all-caps title, so it fell through every classification branch
in assembler._classify into plain body text — splicing itself into the
middle of whatever section is open when a physical page turns. Real
example, MORPHIN SULFAT's liều lượng và cách dùng: "...Nếu\nDTQGVN 2\n 1009\nMorphin sulfat\nuống viên thuốc..." — the header text lands inside a
real dosing sentence.
Why it matters: item 13 (above) already warned "strip the fixed
boilerplate before parsing content" back when the extraction layer was
first built, but that step was never actually implemented in assembler.py
— the warning existed in the catalog without a corresponding code path or
test enforcing it, and nothing caught the gap until a whole-corpus
measurement was actually run. Measured: 1,374 of 11,409 sections (12.0%)
contained a literal "DTQGVN" string mid-text; 671 of 682 monographs (98.4%)
had at least one affected section — this is not a rare edge case, it's
the default outcome for any section whose text happens to cross a physical
page boundary (i.e. most sections longer than about half a page). Left
unfixed, boilerplate gets baked into chunks and embeddings and can surface
mid-sentence in a citation shown to a doctor/pharmacist.
Handling: assembler._is_page_boilerplate drops any span with
column == "full_width" and y0 < HEADER_BAND_Y (the same header-band
threshold page_map.py already uses to read the folio) before it reaches
any other classification branch. Whole-corpus re-measurement after the fix:
0 of 11,409 sections contain "DTQGVN". Regression test uses the exact real
MORPHIN SULFAT span shape.
Generalizes: a documented risk in this catalog is not the same as a
verified-fixed risk — "we know this could happen" needs a whole-corpus
measurement (not just a warning paragraph) before it can be crossed off,
and ideally a regression test that would fail if the fix were ever reverted.
23. PyMuPDF's raw block order doesn't reliably sequence left-column-before-right-column — confirmed wrong on 12 of 1398 pages
What it looks like: extract/spans.py originally trusted PyMuPDF's own
block iteration order to already emit left-column content before
right-column content, validated only against one example page during ADR
0003. On physical page 1100 (the OXYBUTYNIN/OXYMETAZOLIN monograph
boundary) and 11 other pages, PyMuPDF's raw block order emits the right
column first. Since assembler.assemble appends section content to
whichever monograph is currently open, this silently attributed
OXYMETAZOLIN's right-column sections (Chống chỉ định, Thận trọng, Thời kỳ
mang thai, Thời kỳ cho con bú, ADR, Hướng dẫn xử trí ADR, Liều lượng và
cách dùng) to the still-open OXYBUTYNIN monograph — overwriting
OXYBUTYNIN's real sections and leaving OXYMETAZOLIN missing all 7.
Why it matters: medically relevant — wrong contraindication/ADR content
silently attached to the wrong drug. Found via a whole-document
(1668-page) character-similarity diff against an independent parser
(opendataloader-pdf), not from a sample; confirmed by rendering the page
to an image and reading it directly, then confirmed again in the actual
assemble() output.
Handling: extract.spans._sort_blocks_reading_order explicitly sorts
each page's blocks by (full_width header band first, then left column,
then right column) and then by y-position, instead of trusting raw PyMuPDF
order. Whole-range (99-1496) re-scan after the fix: 0 pages with the
reversed-order signature (was 12). Directly verified OXYBUTYNIN's and
OXYMETAZOLIN's assemble()-produced sections are now distinct and
drug-appropriate.
Generalizes: don't trust an upstream library's element ordering just
because it happened to be correct on the one page checked during initial
validation — for a whole-corpus pipeline, explicitly sort by the actual
signal you care about (here: visual column position) rather than an
implicit "the library probably does this right" assumption.
24. Some text exists only as vector outlines — no text extractor can read it, and single dropped glyphs corrupt otherwise-clean sentences
What it looks like: physical page 714 prints 17 full lines of ordinary
GATIFLOXACIN prose that page.get_text() does not return, page.search_for()
cannot find, and neither pdfplumber nor opendataloader-pdf returns either.
page.get_drawings() shows why: each line is a filled path of 1,126-1,831
items, shaped exactly like one line of type and filled with the body-text
colour. The same defect occurs at glyph granularity (39-45 path items), and
that form is far more dangerous — a single Vietnamese diacritic character
drops out of a line that otherwise extracts perfectly: Độ ổn định extracts
as Độ n định, ≥ 1 tuổi as ≥ 1 tu i, tại chỗ as tại ch . The result
reads as ordinary text, so no structural check, no count and no cross-tool
comparison notices it.
Why it matters: this is silent loss of clinical prose in a drug
formulary, and it is invisible to every check that asks a text layer a
question. It survived a whole-document span-coverage ledger reporting
unassigned = 0, because the spans that existed were all routed correctly —
the missing content was never a span at all.
Check: render the page, white out every extracted span's bbox, and look
at the ink that survives (ingestion/validation/residual_ink.py, ~0.06
s/page). Confirm with page.get_drawings(): a filled path with ≥30 items
whose box is 3-20pt tall is type, not decoration (real decoration on this
book carries 1-2 items).
Handling: ingestion/extract/outlined_text.py detects the runs;
recovery cannot be automatic because the paths carry no character codes, so
each run was rendered and transcribed by reading it, into
ingestion/data/verified/outlined_text_transcriptions.json with page, bbox,
and the extracted line it belongs to. Whole-document scope: 51 runs on 5
pages (714 ×31, 736 ×16, 1373, 1444, 1445 ×2), 1,116 characters.
Generalizes: yes — any PDF produced by a layout tool that converts
selected text to outlines (common when a font cannot be embedded) has this.
Never treat "the text layer returned something for this page" as evidence
the page was fully extracted; compare against the rendered pixels.
25. A fraction can be printed with no fraction bar at all, so no geometric detector can find it
What it looks like: ADENOSIN (physical page 147) prints its infusion-rate
formula as three plain lines — Tốc độ truyền dịch (ml/phút) = 0,140 (mg/kg/phút) × trọng lượng cơ thể (kg) / Nồng độ adenosin (3 mg/ml). —
with no rule drawn between numerator and denominator, confirmed by
rendering the region and reading it. Extracted linearly it reads as a
multiplication chain, i.e. the division silently disappears.
Why it matters: it defeats the detector that catches every other 2D
formula in this book. The fraction-bar signal (item 8, and
residual_ink.py's fraction_bar_candidate) finds ink; there is no ink to
find here. It was caught only because a prose-leak gate matched its text.
Check: there is no cheap automatic check. Treat any line ending in a
unit-bearing quantity immediately followed by a line that is itself a
unit-bearing quantity as a division candidate for human review.
Handling: quarantined via the verified region list with
source_prints_no_bar: true. The count of bar-less formulas in this book is
unmeasured — recorded as recall_limit in
ingestion/data/verified/formula_regions_2d.json so the bar scan is never
mistaken for complete formula coverage.
Generalizes: yes — measured precision of the fraction-bar rule on this
book is 16/23 = 69.6%, and its recall is unknown. A geometric heuristic
finds candidates; it never proves absence.
Not yet investigated (flagged for future work, not silently ignored)
- Footnote-style superscript reference markers (seen as
a, b, c, din one table) — not yet checked for whether the footnote text stays correctly associated with its marker/row during extraction. - How many bar-less formulas exist (item 25) — one confirmed, total unmeasured; no geometric signal can bound it.
- Merging the 51 transcribed outlined runs back into monograph text
(item 24) — transcribed and stored, but the corpus still contains
Độ n định. - 2D grid table reconstruction (item 7) — no implementation yet for recovering row/column-correct values from a nomogram-style table.
- Exact shortest monograph name+page — a quick unmerged crude scan (no multi-line title merge) gave a different longest-monograph ranking than the already-documented authoritative one (item 12e: "AMOXICILIN VÀ KALI CLAVULANAT", 45,623 chars), meaning the crude scan's numbers are not reliable enough to name an exact shortest monograph — deferred to the real Phase 1.2 detector (with proper multi-line merge and back-index-validated boundaries), which will produce a trustworthy number as a side effect of its own validation run, rather than trusting today's quick, differently- scoped script.