diff --git a/.gitignore b/.gitignore index 65b11b1..e732155 100644 --- a/.gitignore +++ b/.gitignore @@ -12,6 +12,7 @@ pnpm-debug.log* __pycache__/ *.pyc .venv/ +.venv*/ venv/ .pytest_cache/ .ruff_cache/ diff --git a/README.md b/README.md index 2eabf9a..dadc8ef 100644 --- a/README.md +++ b/README.md @@ -5,7 +5,10 @@ Medical chatbot grounded in the Vietnamese National Drug Formulary See [docs/architecture.md](docs/architecture.md) for the full design (service responsibilities, data stores, RAG ingestion strategy, safety -guardrails), [docs/adr](docs/adr) for architecture decision records, and +guardrails), [docs/adr](docs/adr) for architecture decision records, +[docs/pdf-parsing-outlier-catalog.md](docs/pdf-parsing-outlier-catalog.md) +for a reusable checklist of confirmed PDF-parsing risks (useful for this +book and any similarly-structured PDF), and [docs/progress-log.md](docs/progress-log.md) for a running log of what's been done and what's next. diff --git a/docs/adr/0003-pdf-parsing-strategy.md b/docs/adr/0003-pdf-parsing-strategy.md new file mode 100644 index 0000000..fded011 --- /dev/null +++ b/docs/adr/0003-pdf-parsing-strategy.md @@ -0,0 +1,208 @@ +# ADR 0003: PDF parsing strategy for the drug formulary — validated empirically + +## Status + +Accepted (validated against the real 1668-page source PDF, not assumptions) + +## Context + +The original scaffold's ingestion design (see `docs/architecture.md` history) +assumed generic best practices for structured-PDF parsing: prefer the PDF's +bookmark/outline (`doc.get_toc()`) for section boundaries, fall back to +font-size heuristics. Before writing real ingestion code, this assumption was +tested against the actual `duoc-thu-quoc-gia-viet-nam-2018.pdf` (1668 pages), +because a 1668-page book has enough real-world irregularity that guessing +from a handful of sample pages is not sufficient grounds to trust a parsing +strategy — every claim below was checked against the whole document or a +independently-sourced ground truth, not a small sample. + +## What was actually tested + +1. **`doc.get_toc()`**: returns **0 entries**. No usable bookmark/outline. +2. **Tagged-PDF structure tree** (`/StructTreeRoot`): exists, but is shallow + — ~29 generic `/H1`/`/P` elements, evidently covering only a small + fraction of the document. Not usable as a structural signal at scale. + Confirmed dead end. +3. **Cross-tool text-extraction comparison** on the same real pages + (a known drug monograph, "Abacavir"): + - **PyMuPDF (`fitz`)**: correct reading order, matches the visual source. + - **pdfplumber** (`extract_text()`): **incorrect** — scrambles paragraph + order on this layout and surfaces a stray marked-content artifact + (`"PB
"`) as if it were visible text. Decision: + pdfplumber is kept **only** for its `extract_tables()` API (a genuinely + different, table-specific algorithm), never for general body text. + - **opendataloader-pdf** (Java-based, benchmarks #1 in public leaderboards + for reading order/tables): correct reading order, and its own computed + font metadata (per-span `font`/`font size` in its JSON output) + **independently agreed** with PyMuPDF's raw span data — two unrelated + tools agreeing on the same font facts is real cross-validation, not + opinion. However, its higher-level paragraph/heading classifier is + **inconsistent**: identical bold section-heading text (e.g. "Dược lý và + cơ chế tác dụng", "Liều lượng và cách dùng") is sometimes promoted to a + markdown `##` heading and sometimes silently merged into the following + body paragraph, for no discernible content-based reason. Conclusion: its + Markdown/heading output is not reliable enough to be the sole + structural signal, but it's a useful independent check and its + header/footer-stripping was notably better than raw PyMuPDF text. + - **docling**: attempted, blocked by a `numpy`/`pyarrow` ABI conflict in + the environment (numpy 2.x vs a pyarrow build expecting numpy 1.x, + pulled in transitively via `torch`/`transformers`). Tested inside an + isolated venv rather than fixed globally, to avoid destabilizing other + tools on the machine. See progress log for current status. +4. **The definitive structural signal — bold font spans**: at the raw + PyMuPDF span level, every section heading and every monograph title is + rendered in a **bold** font (`"...-BoldMT"`), while body text is not. + Italic spans exist too (foreign/Latin species names inline) but are + never confused with headings since they're not bold and appear mid + sentence. This was cross-confirmed by opendataloader's independently + computed font metadata for the same spans (see above) — not a + single-tool guess. + - **Font size is NOT a reliable discriminator on its own**: a monograph + title was observed at both 10.0pt ("ABACAVIR") and 9.5pt ("ACARBOSE") + for equally genuine, equally top-level monograph headings. An earlier + draft of the detector required `size >= 9.8` based on the first + example seen and it silently dropped ~15% of real monographs as a + result — a concrete instance of exactly the "don't generalize from one + example" risk this investigation was meant to guard against. The fix: + drop the size floor; use **bold + all-caps + short line length** for + monograph titles, and **bold** alone (cross-checked against the known + section-name vocabulary) for section headings. +5. **Ground truth for validation**: the book has **two** indexes: + - The front-matter "Danh mục các chuyên luận thuốc" (pages 12-31, + 0-indexed): an alphabetical name list with **no page numbers** — useful + only for a name-overlap sanity check, not page-level validation. + - The back-of-book "Mục lục tra cứu" (from page ~1529 printed / ~1528 + 0-indexed onward): a proper index with **exact page numbers** per + generic-name entry (e.g. `"Abacavir, 101"`), plus brand-name + cross-references (`"Ziagen - Abacavir, 101"`, skipped for ground truth). + This is the real, page-verifiable ground truth and should be used for + any future re-validation, not the front-matter list. + - The front matter's own "NỘI DUNG" (table of contents, page 7 0-indexed) + also gives exact page ranges for the book's 3 parts: general topic + chapters (37-98 printed), individual drug monographs (**99-1496 + printed**), appendices (1497-1528), back index (1529+). Any monograph- + boundary detector should be scoped to the 99-1496 printed page range — + scanning the whole book without this scope produces false positives + from front-matter/general-chapter bold-caps lines (org names, decree + headers, chapter titles) that are not drug monographs. + +## Decision + +- **PyMuPDF is the primary and only general-text extractor.** No TOC + dependency, no reliance on the structure tree. +- **Section/monograph boundary detection uses bold-font spans** (not font + size, not font size + vocabulary alone), scoped to the printed page range + of the actual monograph section (99-1496), with all-caps + short length as + the additional signal narrowing bold spans down to monograph titles + specifically. Multi-line wrapped titles must be merged before matching. +- **pdfplumber is retained only for table extraction** (`extract_tables()`), + never general reading order, per the confirmed scrambling issue. +- **The back-of-book "Mục lục tra cứu" is the ground truth for validation**, + not the front-matter drug list. +- **Validation is a repeatable, whole-document, automated check**, not a + one-time manual read of a handful of pages: a full 1668-page scan runs in + under a minute, so re-running it after every heuristic change is cheap and + should be standard practice before trusting a change. + +## Validation results (most recent full-document run) + +- Page-verified recall against the back-of-book index: **91.7%** (665/725 + primary entries had a detected boundary within ±2 physical pages of the + expected page). +- Remaining misses are overwhelmingly one identified, fixable cause: + **multi-line wrapped ALL-CAPS titles** (long Vietnamese drug/vaccine names + spanning 2+ physical lines) being matched as fragments rather than merged + — not a failure of the bold-span signal itself. A handful of misses are + ground-truth extraction noise (the back-index parser occasionally picks up + a non-drug appendix/table-of-contents line that happens to match the + `"Name, ###"` pattern) rather than real detector failures. +- Expected recall after fixing multi-line merging and cleaning non-drug + entries out of the ground truth: materially higher than 91.7%, to be + re-measured once that fix lands (Phase 1 implementation, not this ADR). + +## Follow-up validation: duplicates and cross-page/column data-loss risk + +Two further questions were raised and empirically tested against the full +1405-page monograph range (99-1496 printed): + +1. **Are any drugs detected twice (real content duplication)?** Scanned for + normalized-name collisions at physically distant pages. Found exactly + **one** candidate: `"GONADOTROPIN"` at physical pages 755 and 1371. On + inspection, this is **not** a real duplicate — page 755 is the genuine + "GONADOTROPIN" monograph (hCG/menotropin/follitropin), while page 1371 is + a different monograph, "THUỐC TƯƠNG TỰ HORMON GIẢI PHÓNG GONADOTROPIN" + (GnRH-analog drugs), whose title wraps across two lines — the detector + matched only the second line ("GONADOTROPIN"), colliding with the + unrelated monograph's normalized key. This is the **same multi-line + title-wrapping bug** already identified above, now confirmed with a + second concrete example, not a new failure mode. **Conclusion: no real + duplicate monographs found in the corpus**; the multi-line merge fix + (already required for Phase 1) also resolves this collision. + +2. **Is the PDF two-column, and can content be lost/corrupted across a page + or column boundary during chunking?** Confirmed via bounding-box + inspection: this document **is** genuinely two-column (left column + x≈44-299, right column x≈308-562, same page). PyMuPDF's block-level + reading order correctly sequences left-column-then-right-column content + (already implicitly validated by the correct Abacavir sample earlier). + However, a **separate, real defect** was found and confirmed: on physical + page 1373, one short run of text has its glyphs in **reversed + (right-to-left) x-order**, producing scrambled output — e.g. `" = tịx 8 + yàgn gnàh uềil gnổt( uềihc iổub oàv )magorcim 008 = tịx 4( "`, which + reverses character-by-character back to the correct + `"(4 xịt = 800 microgam) vào buổi chiều (tổng liều hàng ngày 8 xịt = ..."`. + This looks like an isolated PDF-authoring artifact (e.g. an accidental + RTL/BiDi override on one small span during editing), not a systemic + extraction bug. **Initially this scan was scoped to the monograph range + only (1405 of 1668 pages) — an oversight, caught and corrected**: + re-run across all 1668 pages (front matter, general chapters, + monographs, appendices, back index — the entire book, page 0 to the + last page), it still found **exactly 1 affected row, on the same page + 1373, and no others** — confirming the defect is genuinely isolated, not + hiding somewhere in the ~260 pages outside the original scan scope. + - The same full-book pass also checked for near-empty pages (<20 chars + extracted): found exactly **6** — physical pages 3, 37, 99, 1495, 1497, + 1666 — every one lands exactly at a major section boundary (before + "Các chuyên luận chung" at 37, before "Các chuyên luận thuốc" at 99, + before "Các phụ lục" at 1497, near the book's end at 1666). These are + intentional print-layout blank/separator pages, not lost content — + standard practice to force a new part to start on a fresh page. + +**Implications for Phase 1 implementation:** +- Build the pipeline as one **continuous cross-page stream** (text + page + number + bbox per fragment, in reading order), not per-page-isolated + chunks — this is required both for correctly merging multi-line + monograph/section titles (see above) and for never truncating a + paragraph/sentence that spans a page or column break. +- Add an automated **glyph-order sanity check** as a mandatory pass over + 100% of pages (not sampled): group text fragments into visual rows by + y-coordinate, verify x-coordinates are non-decreasing, and either + auto-correct (re-sort by x — the fix is deterministic since raw glyph + positions are known) or flag for manual QA. This check is cheap + (~16 seconds over the full monograph range) and should run before every + real ingestion, not just once. +- The book's content must ultimately be captured **from page 0 to the last + page** — but not all of it as drug-monograph chunks: front matter (pages + 0-36) is mostly low-value organizational/decree content and can be + largely skipped for RAG purposes; general topic chapters (37-98) and + appendices (1497-1528) are real, valuable content that must be ingested + too, using their own heading-hierarchy-based chunking (not the drug + template) — this was already noted in `docs/architecture.md`'s original + design and is reaffirmed here, not changed. The back-of-book index + (1529+) does not need its own chunks (it's a page-locator, not content) + but remains the validation ground truth. + +## Consequences + +- The real ingestion pipeline (Phase 1) should implement the bold-span + detector directly (reusing the validated logic, not the exploratory + scratch scripts), scoped to the correct page range, with multi-line + heading merging as a required fix before first real ingestion run. +- Every future change to the segmentation heuristic should be re-validated + with the same whole-document + back-index cross-reference script (or its + Phase 1 equivalent) before being trusted — this is now the project's + standard rigor bar for this pipeline, not an optional nice-to-have. +- `opendataloader-pdf` (Java-based) and `pdfplumber`'s table extraction + remain candidate tools for the table/formula-handling fallback path + described in `docs/architecture.md`; docling's viability is still + unresolved pending the environment fix. diff --git a/docs/architecture.md b/docs/architecture.md index 51b2b3c..d802f72 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -46,21 +46,40 @@ needed now. ## RAG ingestion pipeline (PDF-specific) The formulary is a structured per-drug reference, not free prose — the -pipeline exploits that structure instead of naive fixed-size chunking: +pipeline exploits that structure instead of naive fixed-size chunking. This +section reflects an actual empirical investigation of the real PDF (not +assumptions) — see `docs/adr/0003-pdf-parsing-strategy.md` for the full +methodology, cross-tool comparison, and validation numbers. -1. **Extraction**: PyMuPDF (`fitz`) as primary extractor (font size/style/ - position metadata enables heading detection); pdfplumber as a fallback - specifically for tabular content (dosing/interaction tables). Raw - per-page extraction is persisted to `ingestion/data/interim/` so +1. **Extraction**: PyMuPDF (`fitz`) as primary extractor. This document has + **no bookmark/outline** (`doc.get_toc()` returns 0 entries — confirmed, + do not rely on it) and is a **tagged PDF with only a shallow, unusable + structure tree** (~29 generic H1/P elements covering a fraction of 1668 + pages — also confirmed dead-end, not a data source). PyMuPDF's reading + order was cross-validated against `pdfplumber` and `opendataloader-pdf` on + real sample pages: pdfplumber's default text order is **unreliable** for + this layout (scrambles paragraph order, leaks marked-content artifacts) — + use it only for its dedicated table-extraction API, never for body text. + Raw per-page extraction is persisted to `ingestion/data/interim/` so re-segmentation doesn't require re-running the expensive extraction step. -2. **Segmentation**: detect drug-entry boundaries (prefer the PDF's - bookmark/outline via `doc.get_toc()` when present, else font-size/style - heuristics), then classify each heading against a canonical section - taxonomy (`chi_dinh`, `chong_chi_dinh`, `lieu_dung`, `tac_dung_phu`, - `tuong_tac_thuoc`, etc., Vietnamese diacritic-insensitive matching). - Output: `{drug_id, drug_name, source_page_range, sections: {...}}` per - drug, persisted to `ingestion/data/processed/monographs.jsonl` and - manually spot-checked via `ingestion/notebooks/`. +2. **Segmentation**: drug-entry boundaries are detected via **bold-font + spans** (PyMuPDF span `font` containing `"Bold"`), not font-size alone — + font size for title/heading spans varies between monographs (confirmed: + 10.0pt and 9.5pt both occur for genuine drug-title headings), so bold is + the reliable signal, all-caps + short length narrows it to monograph + titles specifically. Section headings inside a monograph are also bold + spans, cross-checked against a canonical taxonomy (`chi_dinh`, + `chong_chi_dinh`, `lieu_dung`, `tac_dung_phu`, `tuong_tac_thuoc`, plus + real observed extras like `ten_thuong_mai` "Tên thương mại" not in the + book's own documented 19-field list — treat the taxonomy as open/ + extensible, not a fixed enum). Multi-line wrapped titles/headings (long + Vietnamese names/vaccine names) must be merged across consecutive + bold+all-caps lines before matching — this was the single largest source + of missed detections in validation. Output: `{drug_id, drug_name, + source_page_range, sections: {...}}` per drug, persisted to + `ingestion/data/processed/monographs.jsonl` and validated both + automatically (see ADR 0003) and via manual spot-check in + `ingestion/notebooks/`. 3. **Chunking**: each `(drug, section)` pair is the natural chunk unit; never split a section unless it exceeds a token budget (~500-800 tokens), in which case sub-chunk with a sliding window (400 tokens, 50 overlap), diff --git a/docs/pdf-parsing-outlier-catalog.md b/docs/pdf-parsing-outlier-catalog.md new file mode 100644 index 0000000..2fae0f5 --- /dev/null +++ b/docs/pdf-parsing-outlier-catalog.md @@ -0,0 +1,294 @@ +# PDF Parsing Outlier Catalog + +A generalized checklist of structural risks found while parsing +`duoc-thu-quoc-gia-viet-nam-2018.pdf` (1668 pages). Every item here was +**confirmed with real evidence** (bounding-box inspection, cross-tool +comparison, or a whole-document scan) — not assumed. The goal of this +document is reuse: if this project (or a future one) needs to parse another +structured reference PDF — another national formulary, a different +government-published multi-part document, any dense print-layout book — +this is the checklist of "things that go wrong that a small page sample +won't reveal," and how to actually check for each one cheaply (most checks +here run over the whole 1668-page book in under a minute). + +For the narrative investigation and drug-formulary-specific numbers, see +`docs/adr/0003-pdf-parsing-strategy.md`. This document is the distilled, +reusable checklist form of the same findings, plus items found afterward. + +--- + +## Structural discovery risks (before you even parse content) + +### 1. No bookmarks/TOC +**What it looks like:** `doc.get_toc()` (PyMuPDF) returns an empty list. +**Why it matters:** the obvious, easiest structural signal for section +boundaries simply doesn't exist — don't design a pipeline that assumes it +will. +**Check:** one line, `len(doc.get_toc())`. Do this first, always, before +assuming a bookmark-based approach. +**Generalizes:** yes, directly — always check this before designing around +bookmarks, for any PDF. + +### 2. Shallow/unusable tagged-PDF structure tree +**What it looks like:** the PDF has a `/StructTreeRoot` (looks promising — +"tagged PDF"), but it only covers a handful of generic `/H1`/`/P` elements +for a fraction of the document (here: ~29 elements for 1668 pages). +**Why it matters:** easy to assume "tagged PDF = rich semantic structure +available"; in practice many tagging tools produce a minimal +compliance-only tree that covers almost nothing. +**Check:** walk the struct tree (`doc.xref_object` on `/StructTreeRoot`, +recurse into `/K`) and count real leaf elements vs. total page count. If the +ratio is tiny, it's not a usable data source. +**Generalizes:** yes — always verify depth/coverage before trusting a +struct tree, don't just check for its existence. + +--- + +## Page layout risks + +### 3. Multi-column body layout +**What it looks like:** body pages are genuinely two-column (confirmed via +bounding boxes: left column x≈44-299, right column x≈308-562, page width +≈595). Front-matter pages that *look* like a multi-column name grid to the +eye turned out, on inspection, to be single wide text blocks with internal +whitespace padding between names — not a real structural column split. +**Why it matters:** a naive "read text top-to-bottom regardless of x" pass +would interleave left- and right-column content into nonsense. Conversely, +assuming every visually grid-like page is column-split leads to wasted +effort — verify per page/section, don't generalize from appearance alone. +**Check:** for any suspicious page, dump block bounding boxes +(`page.get_text("dict")["blocks"]`) and look at the actual x0/x1 ranges. A +real column split shows two clusters of x-ranges; a padded single-column +list shows one wide range per line. +**Handling:** PyMuPDF's default block-level reading order handled the real +two-column case correctly here (validated against a known monograph) — the +tool most likely to get column order wrong was `pdfplumber`'s general +`extract_text()` (see item 8), not PyMuPDF. +**Generalizes:** yes — this exact check (dump bboxes, look at x-clusters) +works on any PDF to determine real column count before writing extraction +logic. + +### 4. Full-width content breaking out of the column grid +**What it looks like:** some pages have a table (or could have a figure) +that spans nearly the entire page width (confirmed: a body-surface-area +lookup table's blocks span x≈35 to x≈553, i.e. across both normal columns), +overriding the page's usual two-column layout. +**Why it matters:** logic written to always split a page into "left column" +and "right column" text will misbehave on these pages — the content isn't +in either column, it's a single full-width unit. +**Check:** for any block, compare its x-width against the known +single-column width; if a block's x-range spans (or nearly spans) both +known column ranges, treat it as a full-width unit, not part of a column. +**Generalizes:** yes — any multi-column layout can have occasional +full-width breakout elements (tables, figures, pull-quotes); always check +for this rather than assuming rigid column adherence everywhere. + +--- + +## Table-specific risks + +### 5. Tables split across a page break lose their header on the continuation page +**What it looks like:** confirmed directly — "Bảng 4: Xử trí về điều trị ARV +theo mức độ phát ban" (a 3-column table) starts on one page with its header +row (`['Mức độ', 'Biểu hiện', 'Xử trí']`) and 3 data rows; its 4th data row +("Mức độ 4...") appears on the **next page**, extracted by `pdfplumber` +as a **separate table object with no header row at all**. +**Why it matters:** if a pipeline treats each `find_tables()`/ +`extract_tables()` result as an independent, self-contained table, the +orphaned continuation row is meaningless on its own — you lose the column +semantics for that row entirely. +**Check:** for any table-like structure, check whether the page/column +immediately preceding it ends with a same-shaped table lacking a natural +final row (e.g. an incomplete-looking sequence) — a strong heuristic is +"table starts at the very top of a page/column, no header, same column +count as the table ending at the bottom of the previous page/column." +**Handling:** never treat page-extracted tables as independent; track +continuation explicitly and re-attach the original header to orphaned +continuation rows before using them. +**Generalizes:** yes — this is a generic multi-page-table risk in any +paginated PDF with tall tables; the detection heuristic (position at +page/column top + no header + matching column count to the previous +table) applies broadly. + +### 6. Tables can also split across a column boundary on the *same* page +**What it looks like:** confirmed — "Bảng 6" (ARV drug toxicity table) +starts in the left column near the bottom of a page (header + first data +row) and its remaining data rows appear at the **top of the right column of +the very same page**, again with no header repeated. +**Why it matters:** this is easy to miss because there's no literal page +break — it's tempting to assume "if it's the same page, it's not split," +but a table can still be taller than one column's usable height. +**Check:** same heuristic as item 5, but also check column position, not +just page number — a header-less table fragment starting at the top of a +column (regardless of page) is a suspect continuation. +**Generalizes:** yes, wherever content flows in columns at all — this risk +exists any time column height is shorter than table height. + +### 7. Two-dimensional grid/nomogram tables are not linearly recoverable +**What it looks like:** confirmed — a body-surface-area lookup table +(height across the top, weight down the side, a BSA value at each +intersection) extracts as a scrambled sequence of numbers with no +recoverable row/column association from plain text alone (e.g. `"0,50 +0,52 0,54 0,56"` followed by `"0,55 0,57 0,59 0,61"` — these are almost +certainly column-wise fragments, not the visual rows). +**Why it matters:** unlike a normal bordered table (rows of related +values), a 2D lookup grid's *meaning* depends entirely on 2D position — a +number is meaningless without knowing both its row header (weight) and +column header (height). Flattened text extraction destroys exactly the +information needed to interpret it. +**Check:** any table where extracted "cells" are bare numbers with no +inline label, laid out in a dense grid, is a candidate — cross-check +against the source's own stated formula/description (this table is +explicitly a lookup version of a stated formula, see item 8). +**Handling:** for RAG purposes, prefer **not** to chunk this table as +literal text at all; either (a) reconstruct it properly using per-number +bounding-box position matched against header row/column bboxes (real 2D +table reconstruction, non-trivial), or (b) rely on the accompanying formula +being available for the LLM to compute from directly, and explicitly flag +this table's raw text as unreliable/do-not-cite in metadata. +**Generalizes:** yes — any nomogram, nutrition-fact grid, or nCk-style +lookup table in any PDF has this exact problem; detect by the "bare number +grid" pattern, don't assume normal table extraction works. + +--- + +## Formula / equation risks + +### 8. Formula rendering is inconsistent — some survive as linear text, some don't +**What it looks like:** two real formulas found, two different outcomes. +The Du Bois body-surface-area formula (simple inline exponents, +`"S = W0,425 × H0,725 × 71,84"`) extracted **cleanly as readable text**. The +Cockcroft-Gault creatinine-clearance formula (a stacked fraction — +numerator over denominator, visually 2D) extracted as **scattered, +disordered fragments** with no linear reading order. +**Why it matters:** it's tempting to write one rule ("formulas are +unreliable, always flag them") or its opposite ("formulas extract fine, no +special handling needed") — neither is true here. The determining factor is +whether the formula's visual layout is fundamentally 1D (left-to-right, +like an inline exponent) or 2D (a fraction, a matrix, stacked terms). +**Check:** no cheap automatic detector was built for this distinction yet — +treat any equation/formula-like content as a manual-review candidate, +especially anything with a fraction bar, until a real detector exists +(e.g. checking for large vertical bbox gaps between adjacent glyphs that +should be visually stacked). +**Generalizes:** yes — any technical/medical/scientific PDF with inline +math will have this exact split; don't assume all formulas behave the same +way in extraction. + +--- + +## Character/glyph-level risks + +### 9. Rare reversed (right-to-left) glyph-order defect +**What it looks like:** confirmed exactly once across the entire +1668-page book (physical page 1373): one short text run's glyphs are +positioned in **descending x-order** rather than ascending, producing +scrambled output (e.g. `" = tịx 8 yàgn gnàh uềil gnổt(..."`) that reverses +character-by-character back to the correct Vietnamese sentence +(`"(4 xịt = 800 microgam) vào buổi chiều..."`). +**Why it matters:** this is a genuine, confirmed data-corruption risk, not +theoretical — but it's also extremely rare (1 occurrence in 1668 pages), so +it must be *detected*, not assumed to be either absent or common. +**Check:** group text fragments into visual rows by rounded y-coordinate, +then check whether x-coordinates are non-decreasing across the row; flag +(and optionally auto-correct by re-sorting on x) any row that isn't. This +full-book check runs in about 20 seconds. +**Generalizes:** yes, directly — this is a cheap, universal sanity check +worth running on any PDF text-extraction pipeline as a standing QA gate, +regardless of source document, since it catches a class of PDF-authoring +defects (RTL/BiDi overrides, corrupted content streams) that have nothing +to do with this book specifically. + +--- + +## Section/heading detection risks + +### 10. Font size is not a reliable heading signal — bold is +**What it looks like:** confirmed two genuine, equally top-level monograph +titles at different font sizes (10.0pt and 9.5pt). An early detector +gated on `size >= 9.8` and silently dropped ~15% of real monographs as a +result. +**Why it matters:** a threshold calibrated from one or two examples will +look correct until validated at scale — this is the single clearest +"don't generalize from a small sample" lesson from this whole +investigation. +**Check:** whole-document validation against an independent ground truth +(here, the back-of-book page-numbered index) is what caught this — a +sample of 2-3 pages would not have. +**Generalizes:** yes — for any PDF, prefer a binary style signal (bold/not +bold, a specific font name) over a numeric threshold (size, weight value) +wherever possible, and always validate any numeric threshold against the +whole document, not a handful of examples. + +### 11. Multi-line wrapped titles/headings must be merged before matching +**What it looks like:** confirmed as the dominant cause of missed +detections in whole-document validation — long titles (e.g. "CÁC CHẤT ỨC +CHẾ HMG-CoA REDUCTASE", "THUỐC TƯƠNG TỰ HORMON GIẢI PHÓNG GONADOTROPIN") +wrap across 2+ physical lines; a per-line detector catches only fragments, +which then fail to match a name-based ground truth AND can produce false +name collisions with an unrelated single-line heading elsewhere in the +document (this happened: a wrapped title's second line, "GONADOTROPIN", +collided with a genuine, different, single-line "GONADOTROPIN" monograph +elsewhere). +**Check:** whole-document recall measurement against ground truth; misses +clustered around long/compound names are the signature of this bug. +**Handling:** merge consecutive bold+all-caps lines (with compatible +positioning) into one candidate title before matching/keying, rather than +treating each line independently. +**Generalizes:** yes — any document with long titles/headings that can wrap +will have this exact failure mode; always merge candidate multi-line +headings before using them as unique keys. + +### 12. The documented taxonomy is not exhaustive — keep it open +**What it looks like:** the book explicitly documents a 19-field template +for every drug monograph (page 38), but real monographs contain at least +one undocumented extra field ("Tên thương mại" — brand/trade names) not in +that list. +**Why it matters:** treating a documented schema as a closed enum will +silently misclassify or drop real content that doesn't fit it. +**Generalizes:** yes — any document that describes its own structure in a +preface/README should still be validated against real instances; documented +schemas are frequently incomplete in practice. + +--- + +## Noise / boilerplate risks + +### 13. Header/footer boilerplate must be stripped, but can double as a signal +**What it looks like:** every page carries a page number and a repeating +string (`"DTQGVN 2"`), and body pages additionally carry a running header +naming the current monograph/section. +**Handling:** strip the fixed boilerplate before parsing content, but the +running monograph-name header is a **useful secondary cross-check** for +"which monograph is this page's body text currently part of" — don't +discard it as pure noise. +**Generalizes:** yes — running headers/footers are common in print-derived +PDFs and are usually worth extracting as metadata, not just filtering out. + +### 14. Blank/near-empty separator pages at section transitions are expected +**What it looks like:** exactly 6 near-empty pages (<20 characters) found +across the whole 1668-page book, and every single one lands exactly on a +major section-transition boundary (before general chapters, before +individual monographs, before appendices, near the book's end). +**Why it matters:** a naive pipeline might treat a near-empty page as an +extraction failure and error out or flag it, when it's actually an +intentional print-layout convention (forcing a new part to start on a +fresh page). +**Check:** whole-document scan for pages under some small character +threshold; cross-reference their positions against known section +boundaries before treating them as errors. +**Generalizes:** yes — this print convention is extremely common in +formally typeset books; always expect and gracefully skip near-empty pages +rather than treating them as failures. + +--- + +## Not yet investigated (flagged for future work, not silently ignored) + +- **Footnote-style superscript reference markers** (seen as `a, b, c, d` in + one table) — not yet checked for whether the footnote text stays + correctly associated with its marker/row during extraction. +- **Formula detection heuristic** (item 8) — no automatic detector exists + yet to flag 2D-formula regions before they're trusted as chunk content. +- **2D grid table reconstruction** (item 7) — no implementation yet for + recovering row/column-correct values from a nomogram-style table. diff --git a/docs/progress-log.md b/docs/progress-log.md index 47860e8..fd89693 100644 --- a/docs/progress-log.md +++ b/docs/progress-log.md @@ -13,6 +13,130 @@ end if that risk is showing. --- +## 2026-07-30 — Comprehensive PDF outlier catalog (tables, formulas, columns) + +**Done (in response to direct follow-up questions about table/formula +handling and full-book coverage):** +- Found and confirmed a **table split across a page break loses its header + on the continuation page** — real example: "Bảng 4" (ARV rash management + table) ends with an orphaned, header-less data row on the next page when + extracted with `pdfplumber`. +- Found the **same header-loss risk also happens across a column boundary + within a single page** (no page break needed) — real example: "Bảng 6". +- Found and confirmed **2D grid/nomogram tables are not linearly + recoverable** — the body-surface-area lookup table (appendix) extracts as + scrambled bare numbers with no row/column association. +- Found **two different formula-rendering outcomes**: a simple inline- + exponent formula (Du Bois BSA) extracts cleanly as text; a stacked- + fraction formula (Cockcroft-Gault) extracts as disordered fragments — + confirmed the determining factor is 1D vs 2D visual layout, not "formulas + are always broken." +- Found and confirmed a **full-width table that breaks out of the normal + two-column page grid** (bbox spans nearly the full page width). +- Checked whether front-matter "committee list" pages are genuinely + multi-column (the user suspected 3 columns) — confirmed via bbox + inspection they are **not** true structural columns, just single wide + text blocks with internal whitespace padding between names. +- Consolidated **all** outlier findings from this investigation (this entry + and the previous one) into a single, reusable, generalized reference: + `docs/pdf-parsing-outlier-catalog.md` — written so it can guide parsing of + other similarly-structured PDFs, not just this book. + +**Not done yet / next up:** +- No automatic detector exists yet for (a) 2D-formula regions, or (b) 2D + grid-table reconstruction — both flagged as open items in the catalog, + not silently skipped. +- Table-continuation re-attachment (page-break and column-break cases) has + no implementation yet — needed before Phase 1 can trust any multi-row + table content. +- Phase 1 real implementation still pending overall (see previous entry). + +--- + +## 2026-07-30 — PDF parsing strategy validated empirically (pre-Phase-1) + +**Done:** +- Investigated the real PDF structure before writing any ingestion code + (previous scaffold's assumptions about `doc.get_toc()` turned out wrong). +- Confirmed: 1668 pages, no bookmark/outline (0 TOC entries), tagged-PDF + structure tree exists but is too shallow to use (~29 elements only). +- Cross-tested 3 extraction tools on real sample pages: PyMuPDF (correct + reading order — kept as primary), pdfplumber (scrambled reading order on + this layout — demoted to table-extraction-only use), opendataloader-pdf + (correct reading order, useful independent font-metadata cross-check, but + inconsistent heading classification — not trusted as sole signal). Docling + install hit a numpy/pyarrow ABI conflict in the global Python env; tested + in an isolated `.venv_docling_test/` (gitignored) instead of risking the + global environment — see whether that resolved before relying on it. +- Found the real structural ground truth: every section/monograph heading is + a **bold font span** in the PDF (confirmed at the PyMuPDF span level AND + independently by opendataloader's own font metadata — two tools agreeing). + Font **size** is not reliable (10.0pt and 9.5pt both occur for genuine + monograph titles) — an early size-based threshold silently dropped ~15% of + real monographs; caught and fixed via whole-document validation, not + spot-checking. +- Found the real ground truth for validation: the back-of-book "Mục lục tra + cứu" (page ~1528 onward) has exact page numbers per drug — much stronger + than the front-matter drug list (which has no page numbers). Also found + the book's own contents page states individual monographs run printed + pages 99-1496 exactly. +- Ran automated whole-document (1668-page, ~20-50s per run) validation + against that page-verified ground truth: **91.7% recall** (665/725), with + the remaining gap traced to one concrete, fixable cause (multi-line + wrapped ALL-CAPS titles not yet merged across lines) rather than a flaw in + the bold-span signal itself. +- Documented the full methodology and results in + `docs/adr/0003-pdf-parsing-strategy.md` and updated the ingestion section + of `docs/architecture.md` to match reality (removed the incorrect + TOC-preference assumption). + +**Also validated (in response to direct user questions about correctness):** +- **No real duplicate drug monographs** found across the full 1405-page + monograph range. The one apparent collision ("GONADOTROPIN" at 2 pages) + is a detector artifact from the known multi-line-title bug (a different + monograph's wrapped title fragment collided with it), not real content + duplication. +- **Confirmed the PDF is genuinely two-column** (bounding-box verified: left + column x≈44-299, right column x≈308-562). PyMuPDF's reading order across + columns is correct (already implied by earlier validation). +- **Found and precisely characterized one real data-corruption defect**: + a single text run on physical page 1373 has reversed (right-to-left) + glyph order, producing scrambled text — confirmed by reversing the + string, which recovers the correct Vietnamese sentence. A full scan of + all 1405 monograph pages (grouping fragments into visual rows, checking + for descending x-order) found this exact **1 occurrence and no others** — + rare, isolated, but real, and now has a cheap (~16s) automated detector. +- Full details, methodology, and exact numbers added to + `docs/adr/0003-pdf-parsing-strategy.md` under "Follow-up validation." +- **Caught a real scope gap**: the glyph-reversal scan above was initially + run on the monograph range only (1405 of 1668 pages), leaving ~260 pages + (front matter, appendices, back index) unchecked. Re-ran across the full + 1668 pages: still exactly 1 defect (same page, 1373) — confirmed isolated, + not hiding elsewhere. Also found 6 near-empty pages (3, 37, 99, 1495, 1497, + 1666), all of which land exactly on major section-transition boundaries — + intentional print blank pages, not lost content. + +**Not done yet / next up:** +- Resolve/confirm docling status in the isolated venv (numpy/pyarrow + conflict was fixed by using a separate venv; install completed — actual + parsing comparison against the sample pages still pending). +- Phase 1 real implementation: build `ingestion/` for real using the + validated bold-span detector (not the exploratory scratch scripts) as one + continuous cross-page stream (not per-page silos), fix the multi-line + heading-merge gap, add the glyph-order sanity check as a mandatory + pre-ingestion pass, re-run the validation script to confirm improved + recall, then proceed to chunking + embedding + Qdrant upsert. +- Decide and implement chunking strategy for the non-monograph parts of the + book (general chapters pages 37-98, appendices 1497-1528) — needed so the + full book (page 0 to last) ends up captured in the RAG corpus in some + appropriate form, per the user's explicit requirement that no content be + silently dropped. +- Clean up exploratory `scratch_*` files from the repo root as they + accumulate during investigation (routinely deleted after findings are + persisted to docs — not left in git history). + +--- + ## 2026-07-30 — Initial monorepo scaffold **Done:** @@ -30,7 +154,9 @@ end if that risk is showing. `infra/argocd/`). CI's job is build/test/push image + bump the Helm values image tag; ArgoCD does the actual sync. - `git init` + initial commit (this scaffold). -- Created a private GitHub repo and pushed the initial commit. +- Created a private GitHub repo (`BaoVu2k4/vsf-duocthu`, default branch + `master`) and pushed the initial commit; fixed `targetRevision` in the + ArgoCD Application manifests to `master` to match. **Not done yet / next up (Phase 1 of the build roadmap in `docs/architecture.md`):** - No business logic exists yet anywhere — this was scaffold only. diff --git a/infra/argocd/README.md b/infra/argocd/README.md index 1f578fc..35e1576 100644 --- a/infra/argocd/README.md +++ b/infra/argocd/README.md @@ -7,7 +7,7 @@ this project) rather than a custom push-based CD pipeline. See ## Flow 1. CI (`infra/ci/github-actions/*-ci.yml`) builds and pushes a container image - per app on merge to main, then bumps that app's image tag in + per app on merge to master, then bumps that app's image tag in `infra/helm/medical-chatbot/values-.yaml` (or a per-app values file) and pushes that commit back to the repo. CI never runs `kubectl apply` or `helm upgrade` directly. diff --git a/infra/argocd/applications/dev/app.yaml b/infra/argocd/applications/dev/app.yaml index b5284b9..e45e0de 100644 --- a/infra/argocd/applications/dev/app.yaml +++ b/infra/argocd/applications/dev/app.yaml @@ -7,7 +7,7 @@ spec: project: default # TODO: confirm the team's ArgoCD project/RBAC scope for this app source: repoURL: https://github.com/BaoVu2k4/vsf-duocthu.git # TODO: confirm once repo is created - targetRevision: main + targetRevision: master path: infra/helm/medical-chatbot helm: valueFiles: diff --git a/infra/argocd/applications/prod/app.yaml b/infra/argocd/applications/prod/app.yaml index ab56cdf..944ae28 100644 --- a/infra/argocd/applications/prod/app.yaml +++ b/infra/argocd/applications/prod/app.yaml @@ -7,7 +7,7 @@ spec: project: default # TODO: confirm the team's ArgoCD project/RBAC scope for this app source: repoURL: https://github.com/BaoVu2k4/vsf-duocthu.git # TODO: confirm once repo is created - targetRevision: main + targetRevision: master path: infra/helm/medical-chatbot helm: valueFiles: diff --git a/infra/argocd/applications/staging/app.yaml b/infra/argocd/applications/staging/app.yaml index f55eb6d..5eac9da 100644 --- a/infra/argocd/applications/staging/app.yaml +++ b/infra/argocd/applications/staging/app.yaml @@ -7,7 +7,7 @@ spec: project: default # TODO: confirm the team's ArgoCD project/RBAC scope for this app source: repoURL: https://github.com/BaoVu2k4/vsf-duocthu.git # TODO: confirm once repo is created - targetRevision: main + targetRevision: master path: infra/helm/medical-chatbot helm: valueFiles: