7.6 KiB
05 — Document parsing
How 1,668 PDF pages become 684 structured monographs. The empirical background
is in the pre-existing docs/adr/0003-pdf-parsing-strategy.md,
docs/document-profile.md and docs/pdf-parsing-outlier-catalog.md; this page
describes the code that resulted.
Parsing pipeline
flowchart TD
PDF[/PDF page/]
SP["extract_spans<br/>text + bold flag + bbox + page"]
PM["build_page_map<br/>printed folio per physical page"]
OT["merge_outlined_runs<br/>put vector-path-only text back"]
NG["normalize/glyphs.py<br/>PUA + known-corruption substitution"]
NF["normalize/text_flow.py<br/>visual-line joining"]
CL["assembler._classify<br/>span → Span | _SectionEvent | _TextEvent"]
MT["detect_monograph_titles<br/>bold + mostly-upper + 3..60 chars + page range"]
SH["detect_section_headings<br/>bold + match_section(vocab)"]
CO["_coalesce_titles<br/>merge multi-line headings"]
FP["_filter_false_positive_titles<br/>needs an anchor section ahead"]
AS["assemble<br/>emit Monograph"]
PDF --> SP --> OT --> NG --> NF --> CL
PDF --> PM --> SP
CL --> MT --> CO --> FP --> AS
CL --> SH --> AS
Monograph title detection — segment/detector.py
Rule (validated, ADR 0003): bold + mostly-upper + short line + inside the
monograph page range. Font size is explicitly not part of the rule — a
size >= 9.8 threshold was measured dropping ~15% of real monographs.
MONOGRAPH_PRINTED_PAGE_START = 99 # both printed AND physical bounds
MONOGRAPH_PRINTED_PAGE_END = 1496 # are checked; either alone has
MONOGRAPH_PHYSICAL_PAGE_START = 99 # known failure modes
MONOGRAPH_PHYSICAL_PAGE_END = 1496
_MIN_TITLE_LEN = 3
_MAX_TITLE_LEN = 60
_MAX_LOWERCASE_RATIO = 0.10
_is_mostly_upper tolerates up to 10% lowercase letters rather than requiring
str.isupper(). The reason is a real regression: the class-level monograph
CÁC CHẤT ỨC CHẾ HMG-CoA REDUCTASE embeds the mixed-case CoA, and a strict
check silently dropped the whole monograph. The threshold is a ratio because
an earlier absolute-count version let the short label Mã ATC: through as a
false title.
Known false positive, excluded by name rather than tuned around: part-divider
titles like CÁC CHUYÊN LUẬN THUỐC sit exactly at the printed-page-99 boundary
and are bold + all-caps + short — vocab.is_part_divider rejects them.
A second guard, _filter_false_positive_titles + _has_anchor_ahead, requires a
plausible section heading to follow a candidate title before it is accepted.
Section heading detection — segment/detector.py + vocab.py
Bold spans within the page range are matched against an open vocabulary
(segment/vocab.py::match_section). There is no all-caps requirement here,
because most section headings (Chỉ định, Liều lượng và cách dùng) are not
all-caps. The vocabulary is data, so adding a phrasing is an entry, not a code
change.
The 19 canonical section keys are listed in
06-document-model-and-chunking.md and
duplicated (deliberately, as a closed vocabulary for the LLM) in
apps/ai-service/rag/understanding.py::SECTION_KEYS.
Line-level heuristics — segment/assembler.py
The classifier is where most of the accumulated PDF-specific knowledge lives:
| Helper | Purpose |
|---|---|
_is_page_boilerplate |
Drop running headers/footers |
_starts_its_visual_line / _continues_previous_visual_line |
Rebuild visual lines from spans |
_is_body_line_that_reads_like_a_label |
Stop body prose being read as a heading |
_is_mid_line_label |
A label appearing mid-line, not at line start |
_is_italic_cross_reference |
Italic "see also" runs |
_is_qualifier_line |
Parenthetical qualifiers under a title |
_slugify |
Drug name → drug_id |
Text between a monograph title and its first section heading is captured as
Monograph.preamble rather than dropped — the code names the case: ARTEMETHER
(physical page 210) opens with the regulatory notice that single-agent
artemisinin products were withdrawn.
Table and formula handling
Detection — tables/detect.py, tables/classify.py
Regions are located and classified into shapes:
| Shape | Meaning |
|---|---|
simple_table |
Regular rows/columns |
multi_level_or_merged_header |
Merged/multi-level header |
cross_page_continuation |
Continues onto the next page |
grid_2d_numeric |
2-D numeric lookup grid |
formula_2d |
A 2-D formula (from data/verified/formula_regions_2d.json) |
not_a_table_full_page |
False positive, full-page region |
single_column_boxed_list |
Boxed list, not a table |
QUARANTINE_SHAPES is the subset whose flattened text would be actively
misleading. assembler.py marks spans inside those regions
SPAN_STATE_QUARANTINED; everything else inside a region is SPAN_STATE_TABLE.
The quarantine contract
A quarantined block:
- is lifted out of the section's prose (
SectionSpan.prose_textfiltersquarantinedparts); - becomes a
TableBlockon the monograph with its owntable_id,bbox,physical_pageandshape; - produces a
block_descriptorchunk whose text is built only from metadata — drug name, section display name, "bảng"/"công thức", printed page, and the sentence "Nội dung chỉ tra cứu được trên ảnh trang gốc, không trích dẫn được dưới dạng văn bản." No cell value ever appears; - sets
has_quarantined_content=Trueon every prose chunk of that section, which the retrieval layer reads asrequires_visual_checkand turns into aVERIFY_PDFdecision.
Header rows are deliberately not embedded either
(chunker.py::_attachment forces header_row=[]). The measured reason: 42 of
124 simple-table headers contain a digit, and AMIODARON's (physical page 183)
"header" was a dose — Thời gian liệu pháp tĩnh mạch Liều 720 mg/ngày (0,5 mg/phút).
Normalization
| Concern | Module |
|---|---|
| Private-use-area and known-corruption glyph substitution | normalize/glyphs.py |
| Joining spans into flowing text, hyphenation, line breaks | normalize/text_flow.py |
| Diacritic-stripped casefolding for matching (never for storage) | apps/ai-service/rag/text.py::normalize_name |
Gates pua_char and replacement_char_ufffd both target zero, so a surviving
U+FFFD or PUA codepoint fails the readiness check rather than being embedded.
Verification instruments (no ground truth required)
Three independent instruments, each answering a different question:
| Command | Question | Output |
|---|---|---|
validate |
Did we find the monographs the book's own back index lists? | recall / precision, plus unmatched entries both ways (validation/back_index.py) |
coverage |
Where did every extracted span end up? | span + character counts per state, with the unassigned bucket broken out by page |
residual-ink |
What ink is on the page that no span accounts for? | region census by kind; gate: unclassified must be 0 (validation/residual_ink.py) |
residual-ink is the one that needs no extraction at all to be trusted — it
rasterises the page and asks what the text layer failed to emit.
Known parsing limits, stated by the code itself
scan_glyph_order/scan_reading_orderreport glyph and reading-order defects; they do not correct them. Formula-region issues are expected and left alone.- Table row/column reconstruction is not verified — the
chunk-readyoutput says so. - Recall for borderless tables and bar-less formulas is unquantified.
pdfplumberis used only for its table API; its body-text order is unreliable for this layout.