diff --git a/docs-legacy/README.md b/docs-legacy/README.md deleted file mode 100644 index bb9706b..0000000 --- a/docs-legacy/README.md +++ /dev/null @@ -1,31 +0,0 @@ -# docs-legacy — lịch sử dự án - -`docs/` là bộ tài liệu chuẩn. Thư mục này **chỉ còn giữ lịch sử**: những gì -không tái tạo được từ code. - -| Mục | Là gì | Vì sao giữ | -|---|---|---| -| `adr/` | 11 Architecture Decision Record | Lịch sử quyết định. `apps/ai-service/routers/rag.py:436` tham chiếu trực tiếp `adr/0006` | -| `pdf-parsing-outlier-catalog.md` | Danh mục ca lỗi khi bóc PDF | **Code đang dùng**: `ingestion/cli.py`, `extract/glyph_order.py`, `extract/models.py` và một test đều trỏ tới file này | - -Nhật ký phát triển chi tiết (`progress-log.md`, ~326 KB, đo lường/ngõ cụt/quyết định -theo từng phiên làm việc) không nằm trong bản mirror này — chỉ có trong repo gốc. - -## Đã xoá 2026-08-24 - -Bộ `00-29`, `architecture.md`, các thư mục diataxis (`explanation/`, `how-to/`, -`reference/`, `runbooks/`, `tutorials/`) và các tài liệu kế hoạch/kiểm kê -(`DOCUMENTATION_PLAN.md`, `diataxis-audit.md`, `document-profile.md`, -`bao-cao-kiem-ke-...-2026-08-13.md`, `ke-hoach-showcase-...`, -`pipeline-tu-pdf-den-chatbot-production.md`). - -Lý do: `docs/` đã thay thế chúng và được viết lại từ code, còn bộ này mô tả trạng -thái cũ nên đọc vào dễ hiểu sai. Đã kiểm không file nào trong số đó được code hay -`docs/` tham chiếu. - -Cần đọc lại thì lấy từ lịch sử Git — chúng được track, không mất: - -```bash -git log --oneline -- docs-legacy/00-project-overview.md -git show ^:docs-legacy/00-project-overview.md -``` diff --git a/docs-legacy/adr/0001-vector-db-qdrant.md b/docs-legacy/adr/0001-vector-db-qdrant.md deleted file mode 100644 index a471958..0000000 --- a/docs-legacy/adr/0001-vector-db-qdrant.md +++ /dev/null @@ -1,45 +0,0 @@ -# ADR 0001: Use Qdrant as the vector database - -## Status - -Accepted - -## Context - -The RAG pipeline needs a vector store for drug-monograph chunks. The main -alternative considered was **pgvector** (a Postgres extension), which would -let us reuse the Postgres instance already needed for users/chat history — -one fewer moving part to operate. - -The corpus is not free-flowing prose: it's a structured per-drug reference -with rich per-chunk metadata (drug name, section type, page range). The -common retrieval pattern this domain calls for is "vector similarity search, -filtered by metadata" — e.g. "search only within chỉ định sections" or -"filter to a specific drug the user named" combined with the semantic query. - -## Decision - -Use **Qdrant** as a dedicated vector database, separate from Postgres. - -## Rationale - -- Qdrant gives first-class combined payload-filtering + ANN search in a - single query, which is exactly the retrieval pattern this structured - corpus needs — pgvector supports filtering too, but it's a less natural - fit layered on top of a general-purpose relational engine. -- Vector search becomes its own independent scaling axis, separate from the - transactional Postgres workload (users/chat) — re-indexing or re-ingesting - the formulary doesn't contend with transactional traffic. -- Mature standalone Docker image for local dev, a well-supported Python - client, and a Helm chart for the production Kubernetes deployment target. -- Corpus size (tens of thousands of chunks) is trivial for Qdrant's HNSW - indexing. - -## Consequences - -- One additional service to operate/deploy/monitor compared to pgvector - (which would ride on the existing Postgres). -- Revisit if operational overhead becomes a real burden at our actual scale, - or if we want tighter transactional consistency between chat data and - retrieval — pgvector remains a viable fallback documented here for that - case. diff --git a/docs-legacy/adr/0002-argocd-gitops.md b/docs-legacy/adr/0002-argocd-gitops.md deleted file mode 100644 index b721393..0000000 --- a/docs-legacy/adr/0002-argocd-gitops.md +++ /dev/null @@ -1,68 +0,0 @@ -# ADR 0002: Use the team's existing ArgoCD instance for deployment (GitOps) - -## Status - -**Accepted — still the target, not yet implemented.** Not superseded by the -current production setup. - -Since 2026-08-10 the project has a *different*, interim deployment: a single -EC2 box running `infra/docker/docker-compose.prod.yml` behind Caddy, deployed -by `.github/workflows/deploy.yml` over SSH. That was built to get a working -demo online, not to replace this decision. Migrating to the team's Kubernetes -+ ArgoCD remains planned work, and the expensive prerequisite — containerising -both apps — is already done, so the Dockerfiles and compose services port over. - -Two things must still happen and neither has been started: - -1. **Repository move to the team's self-hosted Gitea** (company domain), which - is where the GitOps repo is meant to live. The project stays on private - GitHub until that move is deliberately made. Note the hard boundary already - in force: the team's existing `git.vinmec.tech/ai-team/gitops` repository is - **reference-only** — never push this project into it. -2. **Filling in the scaffolds this ADR assumes exist.** `infra/helm/medical-chatbot/templates/` - and `infra/k8s/**` are empty (`.gitkeep` only), the chart is version `0.0.0`, - and every `infra/argocd/applications/*/app.yaml` still carries unresolved - TODOs for project, repo URL and destination cluster. - -## Context - -Phase 6 of the build roadmap needs a way to actually deploy the Helm chart to -Kubernetes across dev/staging/prod. The original scaffold (`infra/ci/github-actions/deploy-cd.yml`) -assumed a push-based CI deploy step (CI runs `helm upgrade`/`kubectl apply` -directly against the cluster). The team already runs an ArgoCD instance used -by other projects. - -## Decision - -Deploy via **GitOps through the team's existing ArgoCD instance** instead of -building a custom push-based CD pipeline. ArgoCD Applications -(`infra/argocd/applications/{dev,staging,prod}/app.yaml`) point at -`infra/helm/medical-chatbot` in this repo; ArgoCD watches the repo and -reconciles the cluster to match. - -## Rationale - -- Reuses infrastructure the team already operates and trusts, instead of - standing up a parallel deploy mechanism. -- GitOps gives an auditable history of every deploy (it's just git commits - changing values files/image tags) and a built-in rollback path (revert the - commit). -- Removes the need for CI to hold cluster credentials — CI's job shrinks to - "build, test, push image, bump tag," which is a smaller security surface - than "CI can directly mutate the production cluster." -- Prod uses a non-automated `syncPolicy` (manual approval in ArgoCD) while - dev/staging auto-sync, matching normal caution around production changes. - -## Consequences - -- CI workflows (`infra/ci/github-actions/*.yml`) build/test/push images and - bump the relevant `values-.yaml` image tag + push that commit; they do - **not** call `kubectl`/`helm` against any cluster directly. -- Actual deploy execution and health/sync status live in the team's ArgoCD - UI/CLI, outside this repo — runbooks in `docs/runbooks/` should document how - to check sync status and roll back once the team's ArgoCD instance details - (cluster/server, project, repo URL) are confirmed (see TODOs in - `infra/argocd/README.md`). -- If the team's ArgoCD instance becomes unavailable or this project needs to - fully own its own deploy tooling later, the push-based `deploy-cd.yml` - approach remains a documented fallback. diff --git a/docs-legacy/adr/0003-pdf-parsing-strategy.md b/docs-legacy/adr/0003-pdf-parsing-strategy.md deleted file mode 100644 index fded011..0000000 --- a/docs-legacy/adr/0003-pdf-parsing-strategy.md +++ /dev/null @@ -1,208 +0,0 @@ -# ADR 0003: PDF parsing strategy for the drug formulary — validated empirically - -## Status - -Accepted (validated against the real 1668-page source PDF, not assumptions) - -## Context - -The original scaffold's ingestion design (see `docs/architecture.md` history) -assumed generic best practices for structured-PDF parsing: prefer the PDF's -bookmark/outline (`doc.get_toc()`) for section boundaries, fall back to -font-size heuristics. Before writing real ingestion code, this assumption was -tested against the actual `duoc-thu-quoc-gia-viet-nam-2018.pdf` (1668 pages), -because a 1668-page book has enough real-world irregularity that guessing -from a handful of sample pages is not sufficient grounds to trust a parsing -strategy — every claim below was checked against the whole document or a -independently-sourced ground truth, not a small sample. - -## What was actually tested - -1. **`doc.get_toc()`**: returns **0 entries**. No usable bookmark/outline. -2. **Tagged-PDF structure tree** (`/StructTreeRoot`): exists, but is shallow - — ~29 generic `/H1`/`/P` elements, evidently covering only a small - fraction of the document. Not usable as a structural signal at scale. - Confirmed dead end. -3. **Cross-tool text-extraction comparison** on the same real pages - (a known drug monograph, "Abacavir"): - - **PyMuPDF (`fitz`)**: correct reading order, matches the visual source. - - **pdfplumber** (`extract_text()`): **incorrect** — scrambles paragraph - order on this layout and surfaces a stray marked-content artifact - (`"PB
"`) as if it were visible text. Decision: - pdfplumber is kept **only** for its `extract_tables()` API (a genuinely - different, table-specific algorithm), never for general body text. - - **opendataloader-pdf** (Java-based, benchmarks #1 in public leaderboards - for reading order/tables): correct reading order, and its own computed - font metadata (per-span `font`/`font size` in its JSON output) - **independently agreed** with PyMuPDF's raw span data — two unrelated - tools agreeing on the same font facts is real cross-validation, not - opinion. However, its higher-level paragraph/heading classifier is - **inconsistent**: identical bold section-heading text (e.g. "Dược lý và - cơ chế tác dụng", "Liều lượng và cách dùng") is sometimes promoted to a - markdown `##` heading and sometimes silently merged into the following - body paragraph, for no discernible content-based reason. Conclusion: its - Markdown/heading output is not reliable enough to be the sole - structural signal, but it's a useful independent check and its - header/footer-stripping was notably better than raw PyMuPDF text. - - **docling**: attempted, blocked by a `numpy`/`pyarrow` ABI conflict in - the environment (numpy 2.x vs a pyarrow build expecting numpy 1.x, - pulled in transitively via `torch`/`transformers`). Tested inside an - isolated venv rather than fixed globally, to avoid destabilizing other - tools on the machine. See progress log for current status. -4. **The definitive structural signal — bold font spans**: at the raw - PyMuPDF span level, every section heading and every monograph title is - rendered in a **bold** font (`"...-BoldMT"`), while body text is not. - Italic spans exist too (foreign/Latin species names inline) but are - never confused with headings since they're not bold and appear mid - sentence. This was cross-confirmed by opendataloader's independently - computed font metadata for the same spans (see above) — not a - single-tool guess. - - **Font size is NOT a reliable discriminator on its own**: a monograph - title was observed at both 10.0pt ("ABACAVIR") and 9.5pt ("ACARBOSE") - for equally genuine, equally top-level monograph headings. An earlier - draft of the detector required `size >= 9.8` based on the first - example seen and it silently dropped ~15% of real monographs as a - result — a concrete instance of exactly the "don't generalize from one - example" risk this investigation was meant to guard against. The fix: - drop the size floor; use **bold + all-caps + short line length** for - monograph titles, and **bold** alone (cross-checked against the known - section-name vocabulary) for section headings. -5. **Ground truth for validation**: the book has **two** indexes: - - The front-matter "Danh mục các chuyên luận thuốc" (pages 12-31, - 0-indexed): an alphabetical name list with **no page numbers** — useful - only for a name-overlap sanity check, not page-level validation. - - The back-of-book "Mục lục tra cứu" (from page ~1529 printed / ~1528 - 0-indexed onward): a proper index with **exact page numbers** per - generic-name entry (e.g. `"Abacavir, 101"`), plus brand-name - cross-references (`"Ziagen - Abacavir, 101"`, skipped for ground truth). - This is the real, page-verifiable ground truth and should be used for - any future re-validation, not the front-matter list. - - The front matter's own "NỘI DUNG" (table of contents, page 7 0-indexed) - also gives exact page ranges for the book's 3 parts: general topic - chapters (37-98 printed), individual drug monographs (**99-1496 - printed**), appendices (1497-1528), back index (1529+). Any monograph- - boundary detector should be scoped to the 99-1496 printed page range — - scanning the whole book without this scope produces false positives - from front-matter/general-chapter bold-caps lines (org names, decree - headers, chapter titles) that are not drug monographs. - -## Decision - -- **PyMuPDF is the primary and only general-text extractor.** No TOC - dependency, no reliance on the structure tree. -- **Section/monograph boundary detection uses bold-font spans** (not font - size, not font size + vocabulary alone), scoped to the printed page range - of the actual monograph section (99-1496), with all-caps + short length as - the additional signal narrowing bold spans down to monograph titles - specifically. Multi-line wrapped titles must be merged before matching. -- **pdfplumber is retained only for table extraction** (`extract_tables()`), - never general reading order, per the confirmed scrambling issue. -- **The back-of-book "Mục lục tra cứu" is the ground truth for validation**, - not the front-matter drug list. -- **Validation is a repeatable, whole-document, automated check**, not a - one-time manual read of a handful of pages: a full 1668-page scan runs in - under a minute, so re-running it after every heuristic change is cheap and - should be standard practice before trusting a change. - -## Validation results (most recent full-document run) - -- Page-verified recall against the back-of-book index: **91.7%** (665/725 - primary entries had a detected boundary within ±2 physical pages of the - expected page). -- Remaining misses are overwhelmingly one identified, fixable cause: - **multi-line wrapped ALL-CAPS titles** (long Vietnamese drug/vaccine names - spanning 2+ physical lines) being matched as fragments rather than merged - — not a failure of the bold-span signal itself. A handful of misses are - ground-truth extraction noise (the back-index parser occasionally picks up - a non-drug appendix/table-of-contents line that happens to match the - `"Name, ###"` pattern) rather than real detector failures. -- Expected recall after fixing multi-line merging and cleaning non-drug - entries out of the ground truth: materially higher than 91.7%, to be - re-measured once that fix lands (Phase 1 implementation, not this ADR). - -## Follow-up validation: duplicates and cross-page/column data-loss risk - -Two further questions were raised and empirically tested against the full -1405-page monograph range (99-1496 printed): - -1. **Are any drugs detected twice (real content duplication)?** Scanned for - normalized-name collisions at physically distant pages. Found exactly - **one** candidate: `"GONADOTROPIN"` at physical pages 755 and 1371. On - inspection, this is **not** a real duplicate — page 755 is the genuine - "GONADOTROPIN" monograph (hCG/menotropin/follitropin), while page 1371 is - a different monograph, "THUỐC TƯƠNG TỰ HORMON GIẢI PHÓNG GONADOTROPIN" - (GnRH-analog drugs), whose title wraps across two lines — the detector - matched only the second line ("GONADOTROPIN"), colliding with the - unrelated monograph's normalized key. This is the **same multi-line - title-wrapping bug** already identified above, now confirmed with a - second concrete example, not a new failure mode. **Conclusion: no real - duplicate monographs found in the corpus**; the multi-line merge fix - (already required for Phase 1) also resolves this collision. - -2. **Is the PDF two-column, and can content be lost/corrupted across a page - or column boundary during chunking?** Confirmed via bounding-box - inspection: this document **is** genuinely two-column (left column - x≈44-299, right column x≈308-562, same page). PyMuPDF's block-level - reading order correctly sequences left-column-then-right-column content - (already implicitly validated by the correct Abacavir sample earlier). - However, a **separate, real defect** was found and confirmed: on physical - page 1373, one short run of text has its glyphs in **reversed - (right-to-left) x-order**, producing scrambled output — e.g. `" = tịx 8 - yàgn gnàh uềil gnổt( uềihc iổub oàv )magorcim 008 = tịx 4( "`, which - reverses character-by-character back to the correct - `"(4 xịt = 800 microgam) vào buổi chiều (tổng liều hàng ngày 8 xịt = ..."`. - This looks like an isolated PDF-authoring artifact (e.g. an accidental - RTL/BiDi override on one small span during editing), not a systemic - extraction bug. **Initially this scan was scoped to the monograph range - only (1405 of 1668 pages) — an oversight, caught and corrected**: - re-run across all 1668 pages (front matter, general chapters, - monographs, appendices, back index — the entire book, page 0 to the - last page), it still found **exactly 1 affected row, on the same page - 1373, and no others** — confirming the defect is genuinely isolated, not - hiding somewhere in the ~260 pages outside the original scan scope. - - The same full-book pass also checked for near-empty pages (<20 chars - extracted): found exactly **6** — physical pages 3, 37, 99, 1495, 1497, - 1666 — every one lands exactly at a major section boundary (before - "Các chuyên luận chung" at 37, before "Các chuyên luận thuốc" at 99, - before "Các phụ lục" at 1497, near the book's end at 1666). These are - intentional print-layout blank/separator pages, not lost content — - standard practice to force a new part to start on a fresh page. - -**Implications for Phase 1 implementation:** -- Build the pipeline as one **continuous cross-page stream** (text + page - number + bbox per fragment, in reading order), not per-page-isolated - chunks — this is required both for correctly merging multi-line - monograph/section titles (see above) and for never truncating a - paragraph/sentence that spans a page or column break. -- Add an automated **glyph-order sanity check** as a mandatory pass over - 100% of pages (not sampled): group text fragments into visual rows by - y-coordinate, verify x-coordinates are non-decreasing, and either - auto-correct (re-sort by x — the fix is deterministic since raw glyph - positions are known) or flag for manual QA. This check is cheap - (~16 seconds over the full monograph range) and should run before every - real ingestion, not just once. -- The book's content must ultimately be captured **from page 0 to the last - page** — but not all of it as drug-monograph chunks: front matter (pages - 0-36) is mostly low-value organizational/decree content and can be - largely skipped for RAG purposes; general topic chapters (37-98) and - appendices (1497-1528) are real, valuable content that must be ingested - too, using their own heading-hierarchy-based chunking (not the drug - template) — this was already noted in `docs/architecture.md`'s original - design and is reaffirmed here, not changed. The back-of-book index - (1529+) does not need its own chunks (it's a page-locator, not content) - but remains the validation ground truth. - -## Consequences - -- The real ingestion pipeline (Phase 1) should implement the bold-span - detector directly (reusing the validated logic, not the exploratory - scratch scripts), scoped to the correct page range, with multi-line - heading merging as a required fix before first real ingestion run. -- Every future change to the segmentation heuristic should be re-validated - with the same whole-document + back-index cross-reference script (or its - Phase 1 equivalent) before being trusted — this is now the project's - standard rigor bar for this pipeline, not an optional nice-to-have. -- `opendataloader-pdf` (Java-based) and `pdfplumber`'s table extraction - remain candidate tools for the table/formula-handling fallback path - described in `docs/architecture.md`; docling's viability is still - unresolved pending the environment fix. diff --git a/docs-legacy/adr/0004-chunking-strategy.md b/docs-legacy/adr/0004-chunking-strategy.md deleted file mode 100644 index ebc0771..0000000 --- a/docs-legacy/adr/0004-chunking-strategy.md +++ /dev/null @@ -1,131 +0,0 @@ -# ADR 0004: Chunking strategy for drug monographs — validated against real per-section measurements - -## Status - -Accepted for the monograph range (printed pp. 99-1496) only. General -chapters (pp. 37-98) and appendices (pp. 1497-1528) are explicitly out of -scope — see Consequences. - -## Context - -`docs/architecture.md`'s original "Chunking" paragraph specified `(drug, -section)` as the chunk unit, a ~500-800 token budget, and a 400-token/ -50-overlap sliding window for oversized sections. Those numbers were written -before segmentation existed — a plausible guess, never checked against real -per-section text length. - -Phase 1 (extract → segment → validate) is now real, tested code producing -682 real monographs from the full 1668-page source PDF. This session ran -`python -m ingestion.cli run` for real and measured actual per-section -length across the whole corpus with a temporary investigation script -(`ingestion/scratch/chunking_stats_survey.py`, deleted after this ADR -captured its findings, per this project's investigation-script rule) — -something that had never been measured before this ADR. - -## What was actually measured (whole corpus, 682 monographs) - -- Sections per monograph: min 11, median 17, max 19 (of ~18-19 known - section keys in `segment/vocab.py`'s open taxonomy). -- Whole-monograph length: median 11,480 chars, p90 19,068 chars, max 38,786 - chars. -- Per-section length, converted to a **chars/4 token estimate — an - estimate, not a real tokenizer count**: - - Most of the ~18 section types sit comfortably under 800 estimated - tokens even at their p90 (e.g. `chi_dinh` p90≈268 tok, `dang_thuoc_va_ - ham_luong` p90≈115 tok, `tac_dung_khong_mong_muon` p90≈481 tok). - - **Two sections routinely exceed 800 tokens**: - `duoc_ly_va_co_che_tac_dung` (242 of 678 monographs that have this - section, 35.7%, max ≈3542 tok) and `lieu_luong_va_cach_dung` (200 of - 675, 29.6%, max ≈3631 tok). - - A smaller tail also exceeds it: `than_trong` (25/680, 3.7%), - `tuong_tac_thuoc` (22/642, 3.4%). - - This means: the original 800-token ceiling is directionally correct - (it clears ~16 of 18 section types at their p90 with room to spare), - but "sub-chunk in that case" is not a rare hedge as originally implied - — it is the routine path for roughly a third of all monographs, on two - specific, named, high-clinical-importance sections (mechanism of - action and dosing). - -**A separate, blocking bug was found while gathering this data, not fixed -by this ADR** (out of scope — belongs to `extract`/`segment`, owned by a -parallel session at the time of writing): running header/footer -boilerplate ("DTQGVN 2" + page number + repeated drug name, tagged -`column="full_width"` in `extract/spans.py`) is never filtered out of -section body text before it reaches `SectionSpan.text`. Measured: -1,374 of 11,409 sections (12.0%) contain a literal "DTQGVN" string -mid-text; 671 of 682 monographs (98.4%) have at least one affected section -(e.g. MORPHIN SULFAT's `lieu_luong_va_cach_dung`: `"...Nếu\nDTQGVN 2\n1009\n -Morphin sulfat\nuống viên thuốc..."`). This is `docs/pdf-parsing-outlier- -catalog.md` item 13's known risk, measured whole-corpus for the first time -here. **Chunking must not run against real data until this is fixed** — -otherwise boilerplate is baked into embeddings and can surface mid-sentence -in a chunk shown to a doctor or pharmacist. - -## Decision - -1. **Chunk unit stays `(drug_id, section_key)`** — matches - `segment/models.py`'s existing `Monograph.sections: Dict[str, - SectionSpan]`, matches how a doctor/pharmacist would query ("what does - it say about liều dùng"), and lets a citation point at one clinical - section rather than a whole 2,000-19,000-char monograph. -2. **Token budget: keep the 800-token ceiling** (chars/4 estimate) as the - split trigger. Below it, a section is one chunk, verbatim. This is now a - validated choice, not a guess. -3. **Sub-chunking only applies to the long-tail sections above** (~30-36% - of monographs for the two named sections, a few percent for the rest). - Method: **sentence-boundary-aware sliding window**, replacing the - originally-guessed fixed-character window. Target ~600-700 tokens per - sub-chunk (headroom under the 800 ceiling), ~1 sentence / 50-80 token - overlap between adjacent sub-chunks. Split only at a sentence boundary - (`.`/`;`/`:` followed by whitespace + capital letter), explicitly not - treating a Vietnamese decimal comma (e.g. "0,425") as a boundary. -4. **Why sentence-aware, not line- or character-based**: `assembler.py` - joins `body_lines` one line per PyMuPDF *span*, i.e. one PDF visual - line-wrap point — not a semantic paragraph or sentence boundary. A blind - character/line window can split a sentence mid-way. This is a real, - measured risk here, not theoretical: outlier-catalog item 17 found - adult/child dosing splits ("Người lớn"/"Trẻ em") appear on 1,121 of - ~1,400 monograph-range pages — a chunk boundary landing inside one of - those sentences would be a patient-safety-relevant defect, not a - cosmetic one. -5. **Chunk metadata / provenance** (extends the existing `drug_name, - section_type, source_page_range, chunk_id` list in `docs/architecture.md` - — per this project's provenance convention): `chunk_id` - (`{drug_id}__{section_key}__{part_index}`), `drug_id`, `drug_name`, - `section_key`, `section_display_name`, `atc_codes` (inherited from the - monograph — enables ATC-class-filtered retrieval), exact per-chunk - `source_page_range` and `printed_page_range`, `part_index`/`part_count` - (`0`/`1` for un-split sections, keeps the schema uniform across all chunks). -6. **Schema v4 separates source from retrieval context.** `source_text` is the - exact contiguous source span and is the basis for lossless reassembly and - page provenance. `text` may prefix repeated route/population labels so a - continuation chunk is independently safe to retrieve. Those retrieval-only - prefixes are recorded in `context_labels` and may not alter `source_text`. - Token counts use `cl100k_base`, not the earlier chars/4 estimate. - -## Consequences - -- **Scope**: this decision covers the monograph range only. General - chapters and appendices contain real tables and 2D stacked-fraction - formulas (`docs/document-profile.md`, investigation in progress as of - this ADR) that need their own structural survey before any chunking rule - can be designed for them — do not extend this ADR's rules to those ranges - without a fresh investigation. -- **Hard prerequisite**: the boilerplate-leakage bug described above must - be fixed in `extract`/`segment` before this chunking design is run - against real data for ingestion. This ADR does not fix it. -- **Known gap — sub-compound tagging inside class-level monographs**: 25.5% - of the corpus has more than one ATC code per monograph (outlier item - 12a), e.g. "VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ" documents dosing for 7 - different analogues inside one `lieu_luong_va_cach_dung` section. No - reliable structural signal was found in sampled text to split a section - by sub-compound — a chunk from this section is tagged with the class - name only, not the specific analogue a query might target. Deferred to - golden-dataset-driven eval rather than guessed at now. -- **Resolved — sub-chunk page precision**: schema v4 derives exact physical - support from the contiguous `source_text` span and maps it to verified - printed folios. Missing or ambiguous support fails readiness rather than - falling back to monograph-level provenance. -- **Implemented**: the sentence/label-aware splitter is in - `ingestion/ingestion/chunk/` with regression tests for dose continuations, - compound label boundaries, parent route context and lossless reassembly. diff --git a/docs-legacy/adr/0005-segment-output-contract-for-chunking.md b/docs-legacy/adr/0005-segment-output-contract-for-chunking.md deleted file mode 100644 index 9aa109f..0000000 --- a/docs-legacy/adr/0005-segment-output-contract-for-chunking.md +++ /dev/null @@ -1,215 +0,0 @@ -# ADR 0005: `segment/` output contract needed by `chunk/` — structure-preserving, not flattened - -## Status - -Proposed. **Contract/schema only — no implementation.** `segment/models.py`, -`segment/assembler.py`, and `segment/io.py` are actively owned by a parallel -session on the same checkout at the time of writing; this ADR specifies what -`chunk/` needs from `segment/`'s output precisely enough to implement and -test, but does not touch those files itself. Supersedes part of ADR 0004 -(see "Relationship to ADR 0004" below) — ADR 0004's `(drug_id, section_key)` -chunk-unit-as-leaf assumption is corrected here to `(drug_id, section_key)` -as a **parent**, with sentence-window splitting demoted from primary -strategy to fallback. - -## Context - -ADR 0004 designed chunking against `segment/models.py`'s current output: -`SectionSpan.text` is a single flattened string per section (`"\n".join( -body_line.strip() for ...)`), with all per-line style (`Span.bold`) and -per-line page position discarded once the string is built (confirmed by -reading `assembler.py`: `body_lines.append(span.text.strip())` keeps only -`span.text`, nothing else). Review of ADR 0004 surfaced four real problems -that trace back to this flattening, not to the chunking algorithm itself: - -1. **A section is not a single semantic unit.** `liều lượng và cách dùng` - and `tương tác thuốc` routinely contain multiple distinct facts (dosing - per patient population, dosing per organ-function impairment, multiple - separate drug interactions) that a doctor may want to retrieve - independently. Measured: an explicit population marker ("Người lớn"/ - "Trẻ em"/"Trẻ sơ sinh"/"Suy thận"/"Suy gan" immediately followed by `:` - or `.`) appears in **303 of 675 monographs (44.9%)** that have a `liều - lượng và cách dùng` section — this is common, not an edge case. -2. **A blind sentence-boundary sliding window (ADR 0004's original - sub-chunking method) can still split two different facts into the same - chunk, or split one fact across two chunks**, because it has no way to - know a population/interaction boundary exists — that information exists - in the source (as a bold or otherwise visually distinct sub-heading, per - direct reading of MORPHIN SULFAT/VITAMIN D section text: lines like - "Thuốc uống", "Cách dùng:" render as isolated bold short lines in the - PDF) but is discarded before `chunk/` ever sees it. -3. **Tables inside the monograph range are not addressed at all.** ADR 0004 - implicitly assumed monograph-range sections are prose. `docs/pdf-parsing- - outlier-catalog.md` item 19 already documents a real table (dosing by - renal function, HSV/CMV columns) inside a monograph body (Foscarnet - natri, physical page 698) — flattening a table's rows into - newline-joined body text destroys its row/column structure exactly the - way outlier item 7 already describes for the appendix's 2D nomogram - table. A whole-range survey to size this properly is in progress - alongside this ADR (see "Not yet resolved" below). -4. **Provenance is section-level, not chunk-level**, because per-line - `physical_page`/`y0` (which `Span` already carries — see - `extract/models.py`) is discarded at the same flattening point. For a - section spanning several physical pages, a sub-chunk built from its - final third currently has no way to know its own real page — it can - only inherit the whole monograph's `source_page_range`. For medical - citations this is not precise enough. - -**A fifth, independently-found data-quality bug makes precise provenance -even more necessary, not less**: the corpus's last-processed monograph -(ZOLPIDEM) is never closed until true end-of-stream, and `assembler._classify` -calls `match_section()`/`match_section_with_inline_value()` on every span -with **no `in_monograph_range` gate** (unlike `_TextEvent` handling, which -does check it). A spurious bold-text match on physical page 1655 — deep in -the back-of-book "Mục lục tra cứu" brand-name index, confirmed by reading -that page directly — overwrote ZOLPIDEM's real `tương tác thuốc` -`SectionSpan` with an empty one and corrupted its `source_page_range` to -`[1492, 1655]`. This is real content loss (measured: exactly 1 monograph -affected, the last one processed — every other monograph is closed on -schedule by the next monograph title, which *is* range-gated). Flagged for -the session that owns `extract`/`segment`, not fixed here. - -## Decision - -Extend `segment/models.py`'s `SectionSpan` with a structured, line-level -representation, additive to (not replacing) the existing flat `text` field -— `chunk/` becomes a real, structure-aware consumer instead of re-deriving -structure from a flattened string via ad hoc regex. - -### New/changed types (`segment/models.py`) - -```python -@dataclass(frozen=True) -class BodyLine: - text: str - physical_page: int - y0: float - bold: bool # Span.bold, preserved instead of discarded - -@dataclass -class SectionSpan: - key: str - display_name: str - heading: Heading - text: str # UNCHANGED meaning, kept for - # backward compat (see invariant below) - lines: List[BodyLine] = field(default_factory=list) # NEW -``` - -`lines` carries exactly the per-line signal `chunk/` needs to do its own -job (population/subheading detection, precise page provenance) without -`segment/` having to know anything about chunking — `segment/`'s -responsibility stays "detect boundaries and preserve source structure," not -"decide what a retrieval unit is" (Clean Architecture / SoC). Specifically, this is deliberately **not** a `is_subheading: -bool` field computed by `segment/` — classifying "is this line a -subheading a chunker should split on" is a chunking-time decision (what -counts as a good split point can vary by strategy/eval results), not a -segmentation-time one. `segment/` should stop discarding the raw signal -(`bold`, `y0`, `physical_page`) it already has per span; it should not also -start doing chunk-shaping judgment calls. - -### Invariants - -1. `text == "\n".join(l.text for l in lines).strip()` for every - `SectionSpan`, for the lifetime of this contract — `lines` is a strictly - additive refinement, never a divergent second source of truth. Any - change to how body text is assembled (e.g. the boilerplate-stripping fix - already applied by the other session) must update both fields from the - same filtered span list, not `text` alone. -2. `lines` is in reading order, matching the order `text`'s lines already - implicitly have. -3. Every `BodyLine.physical_page` satisfies `detector.in_monograph_range` - for a `Span` on that page — i.e., **no line in any `SectionSpan.lines` - may come from outside the monograph's real printed-page range**. This is - the ZOLPIDEM bug's exact failure mode stated as an invariant: it was - violated (a spurious section event was accepted from a fully - out-of-range page precisely because no such check existed for section - *events*, only for body *text* events). Enforcing this invariant closes - that bug as a side effect, but the invariant is stated here as a - contract requirement independent of any specific fix implementation. -4. Every currently-open monograph must be finalized exactly once, at either - (a) the next monograph title, or (b) true end-of-stream — with no third - path (e.g., a stray out-of-range section match) able to silently mutate - an already-"complete" monograph's sections after point (a) would - otherwise have applied. (This is a restatement of invariant 3 from the - monograph-lifecycle side, not a new requirement.) - -### Migration impact - -- **`segment/io.py`** (`_monograph_to_dict`/`_monograph_from_dict`, - `write_monographs_jsonl`/`read_monographs_jsonl`): additive — serialize - `lines` alongside the existing `text`/`heading` fields per section. - Existing consumers reading only `text` (e.g. `segment/atc.py`'s - `extract_atc_codes`, which regexes over `SectionSpan.text`) need no - change, per invariant 1. -- **`ingestion/data/processed/monographs.jsonl`**: schema grows a new - optional-shaped field (`sections[key].lines`). No `schema_version` field - currently exists in the serialized dict (checked `io.py` directly) — - worth adding as part of this change, both for this migration and because - `docs/architecture.md` already assumes "collection aliasing allows - re-ingesting with a changed chunking strategy," which implies the - ingestion output itself should be able to declare which schema shape it - is. -- **Existing 110 tests**: unaffected if invariant 1 holds — no assertion in - the current suite inspects `lines` (it doesn't exist yet), and `text`'s - value/semantics are unchanged. -- **New tests required** (this ADR specifies them; implementation and the - actual test code are not part of this ADR): - 1. Regression test reproducing the ZOLPIDEM failure shape: a synthetic - span stream — last monograph's title and real sections, followed by - spans whose `printed_page` is out of `in_monograph_range` but whose - text matches a `vocab.py` section label — asserting the monograph - closes with its real sections intact and the out-of-range spurious - match is ignored, not accepted. - 2. `SectionSpan.lines` fixture test: using the real MORPHIN SULFAT - boilerplate-fix fixture already in `tests/test_segment_assembler.py`, - assert `lines` preserves the correct `bold`/`physical_page`/`y0` per - retained line (and that stripped boilerplate lines are absent from - `lines` too, not just from `text`). - 3. Round-trip test: `write_monographs_jsonl` → `read_monographs_jsonl` - preserves `lines` exactly (dataclass equality per line). - 4. Whole-corpus invariant-1 check: for a real `cli run` output, assert - `text == "\n".join(l.text for l in lines).strip()` holds for every - section of every monograph, not a sample. - -## Relationship to ADR 0004 - -ADR 0004's chunk-unit decision (`(drug_id, section_key)`) is **not** -discarded — a section is still the natural *parent* grouping (matches how a -clinician thinks, matches `Monograph.sections`). What changes: ADR 0004 -described a section as directly *the* chunk when under the 800-token -ceiling, with sentence-window splitting as the fallback for oversized -sections. Per the review above, splitting must instead **first** attempt to -break at real structural boundaries available in `SectionSpan.lines` (a -bold, short, isolated line — the same "subheading" shape already visually -confirmed for route-of-administration/population sub-headers — or an -explicit population/organ-function marker), with the sentence-window method -demoted to a fallback for the remaining prose that has no such marker. The -exact splitting algorithm (how a "subheading-shaped line" is defined -precisely, in code) is a `chunk/`-side implementation detail *enabled* by -this contract, not decided by it. - -## Not yet resolved (explicitly out of scope for this ADR) - -- **Table/formula content blocks.** A separate whole-monograph-range survey - (pdfplumber `find_tables()` + PyMuPDF math-symbol scan, physical pages - 98-1494 excluding blank page 99 — the exact set `detector. - in_monograph_range` accepts, not an assumed offset) is in progress at the - time of writing, per explicit user instruction to measure before deciding - a table/formula chunk-unit strategy. This ADR's `BodyLine` - contract covers **text content only**; a table/formula region should - *not* currently be flattened into `BodyLine`s (doing so would repeat - exactly the "destroys row/column meaning" mistake outlier item 7 already - documents) — but the precise `ContentBlock`/table-row/formula-unit shape - is deferred to a follow-up revision of this ADR once the survey reports - real numbers (how many monographs/sections affected, page-break - continuation frequency, multi-tier headers, merged cells, footnotes). -- **Paragraph-boundary detection** (grouping consecutive `BodyLine`s into a - flowing paragraph vs. a new one) is left to `chunk/`, using the same - kind of y-gap heuristic `segment/merge.py` already validates for - multi-line title wraps (`_MAX_LINE_GAP_PT`) — `BodyLine.y0` is sufficient - raw signal for `chunk/` to compute this itself; `segment/` does not need - to pre-compute paragraph grouping. -- **The actual `chunk/` splitting implementation** (subheading detector, - population-marker regex, sentence-window fallback) is not part of this - ADR — this ADR defines the data contract that implementation will consume. diff --git a/docs-legacy/adr/0006-quarantined-block-references-in-chunks.md b/docs-legacy/adr/0006-quarantined-block-references-in-chunks.md deleted file mode 100644 index 46068ae..0000000 --- a/docs-legacy/adr/0006-quarantined-block-references-in-chunks.md +++ /dev/null @@ -1,167 +0,0 @@ -# ADR 0006: chunks must carry references to lifted table/formula blocks - -## Status - -Accepted and implemented in schema v4. Resolves the item ADR -0005 explicitly deferred ("Table/formula content blocks … the precise -`ContentBlock`/table-row/formula-unit shape is deferred to a follow-up -revision of this ADR once the survey reports real numbers"). The survey has -reported. - -## Context - -`segment/` now lifts table and formula regions out of section prose and -quarantines them (ADR 0003 lineage, outlier-catalog items 7, 8, 24, 25). -That was the right move — linearised, AMPICILIN VÀ SULBACTAM's -Cockcroft-Gault fraction read as `Clcr (ml/phút) = 72 x creatinin huyết -thanh`, i.e. a division presented as a multiplication, in a renal-dosing -section. - -But `chunk/models.py` has no field that refers to a lifted block. Measured on -the current whole-corpus output: - -| quantity | value | -|---|---| -| lifted blocks represented by descriptor chunks | 151, all quarantined | -| sections affected | 103 | -| **blocks in `lieu_luong_va_cach_dung`** | **125** | -| unverified header rows admitted to embedding text | **0** | - -So three quarters of everything removed from prose was removed from the -dosing section, in a drug formulary, for an audience of doctors and -pharmacists. - -**The failure this creates is silent, not visible.** A chunk of AMPICILIN VÀ -SULBACTAM's `lieu_luong_va_cach_dung` is grammatical, complete-looking prose -with the renal-dosing table absent and nothing marking the absence. Retrieval -ranks it, the model answers from it, and neither has any way to know a table -was taken out. A visible error would be safer than this. - -A second, quieter failure: a table is currently **unreachable**. Nothing in -the index represents it, so "bảng liều theo chức năng thận của ampicilin" -cannot retrieve it even in principle. - -## Decision - -Chunks reference blocks; blocks' content never becomes embedded text. - -### 1. `Chunk` gains typed attachments - -```python -@dataclass(frozen=True) -class ChunkAttachment: - block_id: str - kind: str # "table" | "formula" - shape: str # simple_table | multi_level_or_merged_header | - # cross_page_continuation | formula_2d - physical_page: int - printed_page: int - bbox: List[float] - quarantined: bool - header_row: List[str] = () # always empty until separately verified - source_crop: str | None = None - -@dataclass(frozen=True) -class Chunk: - ... - chunk_kind: str = "prose" # "prose" | "block_descriptor" - attachments: List[ChunkAttachment] = () - has_quarantined_content: bool = False -``` - -`has_quarantined_content` is derivable from `attachments`, and is serialized -anyway. A consumer that never looks at `attachments` must still be unable to -miss the fact — the whole defect being fixed here is a consumer not knowing -what it was not told. - -### 2. One descriptor chunk per block, built from metadata only - -A block also gets its own chunk so it is retrievable at all: - -``` -chunk_id = "{drug_id}:{section_key}:block:{block_id}" -chunk_kind = "block_descriptor" -text = "AMPICILIN VÀ SULBACTAM — Liều lượng và cách dùng — bảng, - trang in 204." -``` - -The text is assembled only from verified metadata: drug name, section display -name, block kind and printed page. **No cell value or inferred header appears.** -The earlier proposal to use `pdfplumber.find_tables()`'s first row was rejected -after corpus audit: a guessed first row can be a body row or can merge numeric -relationships. Until a separate human-verified header dataset exists, -`header_row` is embargoed for every shape and serialized as empty. - -### 3. The answer layer's obligations (binding on `ai-service`) - -These obligations are implemented across `ingestion/` and `ai-service` and are -enforced by tests/readiness gates. - -1. A retrieved chunk with `has_quarantined_content: true` **must** cause the - answer to state that a table or formula exists at the cited page, and to - surface its rendered crop. The answer may not present itself as complete. -2. A `block_descriptor` chunk may be answered **only** with the crop. It must - never be paraphrased, and its `header_row` must never be presented as the - table's content. -3. No chunk carrying a quarantined attachment may be used to state a numeric - dose. If the dose is in the table, the answer is the crop plus the page. - -### 4. `schema_version` - -`monographs.jsonl` and the chunk output both gain `schema_version`. ADR 0005 -flagged its absence; a schema that now has two chunk kinds and an attachment -list cannot be safely consumed without one. - -## Alternatives rejected - -- **Flatten the block into the chunk text.** This is the defect, not the fix - — it reproduces `Clcr = 72 x creatinin` exactly. -- **Chunk the block's linearised text as an ordinary chunk.** Worse than - flattening: it makes unsafe text independently retrievable *as prose*, with - its quarantine flag one dereference away from being ignored. -- **Drop the blocks.** Silent loss, and contrary to the standing rule that - unreconstructable content is quarantined with full provenance, never - deleted. -- **Rely on the prose saying "xem bảng".** The prose often does not, and a - retrieval layer cannot act on an unstructured hint. -- **Wait for row/column reconstruction and do this once.** Reconstruction is - days of work and would leave the corpus unchunkable meanwhile; worse, it - would make the schema question look answered when the *silent-incompleteness* - problem is independent of whether the rows are recovered. Reconstruction - later populates `rows` on the same attachment without touching consumers. - -## Why a crop is a legitimate answer, not a placeholder - -For doctors and pharmacists a rendered crop of the source page is the -highest-fidelity response available: it *is* the book, and it is verifiable at -a glance. Reconstruction earns its keep for a different job — comparing or -combining values across drugs, which is the synthesis use case this product -exists for — not for single-table lookup. - -## Invariants and gates - -Added to `cli chunk-ready` and to the chunk stage's own tests: - -1. `section_with_lifted_block_but_no_chunk_reference = 0` -2. `attachment_block_id_unknown = 0` — every referenced id exists on the - monograph -3. `attachment_without_page_or_bbox = 0` -4. `block_text_leaked_into_chunk_text = 0` — no chunk's embedded text - contains a quarantined block's text -5. `descriptor_chunk_count == block_count` -6. `descriptor_chunk_without_attachment = 0` -7. `attachment_header_row_present = 0` -8. `descriptor_with_unverified_columns = 0` -9. `descriptor_range_not_attachment_page = 0` -10. `attachment_without_printed_page = 0` - -## Consequences - -- Prose chunks shrink slightly in trustworthiness terms but grow in honesty: - the ones missing a table now say so. -- The current candidate index gains 151 descriptor chunks, - each cheap and none carrying unsafe text. -- `ai-service` cannot answer a dosing question from prose alone for the 103 - affected sections without violating a stated contract. -- The 14 `formula_2d` attachments make the two Cockcroft-Gault formulas - answerable as crops today, which they are not now. diff --git a/docs-legacy/adr/0007-conversational-reasoning-rag.md b/docs-legacy/adr/0007-conversational-reasoning-rag.md deleted file mode 100644 index 7726a2a..0000000 --- a/docs-legacy/adr/0007-conversational-reasoning-rag.md +++ /dev/null @@ -1,210 +0,0 @@ -# ADR 0007: Conversational reasoning RAG — state, bounded loop, and how it is measured - -**Status:** superseded by ADR 0008 (2026-08-07). See the note below before -reading this as a description of anything currently running. -**Supersedes:** nothing. Extends ADR 0005 (segment output contract) and ADR 0006 -(quarantined block references) rather than replacing them. - -> **2026-08-07 — why this was superseded, not deleted.** An independent -> 7-agent audit on 2026-08-06 found `bootstrap.py` never constructs any of -> `rag/conversation.py` / `rag/reasoning.py` / `rag/conversational.py` — the -> live agent (`rag/agent.py::RagAgent`, wired in since the F-03 rebuild on -> 2026-08-06) is a fixed one-shot pipeline (understand → route → retrieve -> once → generate → ≤2 same-claim entailment retries), not the PLAN/RETRIEVE/ -> ASSESS/REFINE/VERIFY loop or the `Focus`/`ConversationState`/TTL state -> design below. This was a real, deliberate pivot mid-implementation, not an -> abandoned-but-still-intended plan: `rag/agent.py`'s own module docstring -> says outright that `ConversationalLoopService` + `conversation.py` were -> replaced because "the LLM reads a plain turn history and resolves -> ['thuốc đó' / 'còn liều thì sao'] itself" — simpler than maintaining -> `Focus`/TTL/turn-budget state by hand, and proven live across many -> multi-turn conversations since. Section 6 below ("Refused: an LLM -> confidence score as the loop's uncertainty signal") is the clearest -> evidence this is a genuine architecture change, not a gap: the live system -> now uses exactly that — an LLM sufficiency/clarify judgment — as its -> ask-or-answer signal, the opposite of what this ADR chose. -> -> The three modules this ADR specified (1,314 lines) and their five dedicated -> test files (42 tests) were deleted on 2026-08-07 rather than left as dead -> code, once confirmed to have zero live importers anywhere -> (`bootstrap.py`/`main.py`/`agent.py`/`answer.py`/`routers/rag.py`). This -> document is kept, unedited below this notice, as the historical record of -> why that design was chosen and what it traded off — see ADR 0008 for what -> actually runs today, including what this ADR got right that ADR 0008 -> still owes (a real request-scoped time/call budget — F-08, still open; a -> durable, cross-worker conversation store — currently an in-process dict). - -## Context - -The service answers one question at a time. `POST /v1/rag/query` carries no -conversation id, `apps/chat-service` holds zero source files, and every request -re-resolves the drug from scratch. Three consequences, all observed in the UI on -2026-08-05: - -- `paracetamol` alone is refused rather than asked about. -- `liều dùng paracetamol cho người lớn` returns the identical answer to - `liều dùng paracetamol` — the qualifier is not used at any stage. -- A follow-up such as *"còn trẻ em thì sao?"* cannot work at all, because - nothing carries the drug forward. - -The owner's requirement is a **conversational reasoning RAG**: history, an -internal reasoning stage, and a bounded self-improvement loop. - -The binding constraint is that this is a drug formulary for clinicians. Every -capability below is designed so that adding it cannot widen what the system is -allowed to assert. - -## Decision - -### 1. Conversation state - -Two stores with different jobs, deliberately not merged. - -**`Focus` — structured, drives routing.** This is what makes *"còn trẻ em thì -sao?"* resolvable without an LLM. - -| Field | Purpose | -|---|---| -| `drug_id`, `drug_name` | The drug under discussion | -| `section_key` | The attribute last answered | -| `population` | `nguoi_lon` / `tre_em` / `phu_nu_co_thai` / … | -| `verbosity` | `concise` \| `detailed`, set when the user asks | -| `set_at_turn` | Turn index each field was last set | - -**`ConversationState` — the whole record.** - -``` -conversation_id -recent: tuple[Turn, ...] # last K turns, verbatim -summary: str # rolling prose summary of everything older -focus: Focus -turn_count: int -``` - -A `Turn` carries `role`, `text`, `at`, and — for assistant turns — the -`drug_id`, `section_key` and `evidence_ids` that produced it. Storing the -evidence ids is what lets the planner answer a follow-up **from evidence -already retrieved** instead of retrieving again. - -**Carry-over is never silent.** An inherited `drug_id` that is wrong is a -wrong-drug answer, so any answer built on inherited focus must name what it -inherited: *"Về Metformin, ở trẻ em: …"*. This is a hard rule, not a -presentation preference. - -**Focus expires.** A field older than `FOCUS_TTL_TURNS` (6) is dropped rather -than inherited. Conversations drift, and a drug from ten turns ago is not -context, it is a hazard. - -### 2. Recent history and summary - -- `recent` holds the last **K = 6** turns verbatim (three exchanges). -- When a turn falls out of `recent`, it is folded into `summary`. -- `summary` is regenerated at most every **S = 4** turns, capped at **400 - tokens**; `recent` is capped at **2000 tokens**, oldest dropped first. -- **The summary records what was discussed, never clinical content.** It may - say *"đã hỏi liều dùng của Metformin cho người lớn"*; it may not carry a dose. - A dose restated from a summary would have no citation and could not be - grounding-verified — the check compares against retrieved evidence, and a - summary is not evidence. - -### 3. Reasoning loop - -```mermaid -flowchart TD - A[User turn] --> B[UNDERSTAND
resolve against Focus] - B --> C{Clarify signal?} - C -->|ambiguous drug / no attribute /
multi-attribute| Z[ASK — 1 turn, no loop] - C -->|no| D{Simple?} - D -->|drug + section resolved,
no follow-up ambiguity| E[RETRIEVE] - D -->|complex / decomposable| P[PLAN
sub-questions + retrieval set] - P --> E - E --> F[ASSESS sufficiency] - F -->|insufficient AND rounds left| R[REFINE query] --> E - F -->|sufficient OR rounds exhausted| G[GENERATE] - G --> H[VERIFY
grounding + coverage] - H -->|ungrounded / off-target,
repairs left| G - H -->|grounded| Y[RESPOND] - H -->|repairs exhausted| X[FALL BACK
verbatim source] - F -->|exhausted AND still thin| Z -``` - -**Continue conditions** — a round is spent only when all hold: -1. `retrieval_rounds < MAX_RETRIEVAL_ROUNDS` (2) -2. the assessor named a *specific* missing thing (a section, a population, a - second drug) — "feels incomplete" is not a reason to spend a round -3. the refined query differs from every query already tried this turn - -**Stop conditions** — any one ends the loop: -- sufficiency satisfied -- budget exhausted (rounds, LLM calls, wall-clock, tokens) -- a clarify signal fires (these bypass the loop entirely — asking beats guessing) -- grounding verification fails after `MAX_REPAIRS` (1) → extractive fallback - -**Fast path.** When the drug resolves and `SectionResolver` returns a section -and no clarify signal fires, the loop is skipped: retrieve → generate → verify. -This is the majority path and it costs one LLM call. - -### 4. Budgets - -| Limit | Value | Enforced at | -|---|---|---| -| `MAX_RETRIEVAL_ROUNDS` | 2 | loop guard | -| `MAX_REPAIRS` | 1 | loop guard | -| `MAX_LLM_CALLS` per turn | 4 | budget object, checked before each call | -| `MAX_WALL_CLOCK_MS` | 20000 | checked between stages | -| `MAX_EVIDENCE_TOKENS` | 12000 | evidence assembly, oldest-dropped | -| `FOCUS_TTL_TURNS` | 6 | state update | - -The budget is a single object threaded through the loop and **decremented -before** each call, so exhaustion degrades to the best answer so far rather -than to an error. - -### 5. Integration - -New domain modules, no SDK imports: - -- `rag/conversation.py` — `Focus`, `Turn`, `ConversationState`, window and - focus-update rules. Pure; the follow-up resolution in it needs no LLM. -- `rag/reasoning.py` — the loop, its budget, and its stage protocols. -- `rag/ports.py` — `ConversationStore` (load/save), `Summariser`, `Planner`, - `SufficiencyAssessor`. Each has a deterministic no-LLM default so the whole - loop runs offline. - -New adapter: `adapters/postgres.py` gains `PostgresConversationStore`. - -Unchanged and still binding: `GroundedAnswerService` remains the single-turn -engine; `grounding.verify` gates every generated answer; `VERIFY_PDF` evidence -is never generated over. - -### 6. Measurement - -A capability that cannot be shown to help does not ship. Three modes are run -over the same cases — `single-shot`, `+history`, `+reasoning-loop`: - -| Metric | Answers | -|---|---| -| follow-up resolution accuracy | does *"còn trẻ em thì sao?"* reach the right drug+section+population | -| on-target rate | does the answer contain the population/attribute actually asked for | -| grounding rejection rate | does reasoning make fabrication more or less likely | -| clarify rate / clarify precision | does it ask when it should, and only then | -| median + p95 latency, LLM calls, tokens per answered turn | what the capability costs | - -The evaluation set is a **new multi-turn golden file** — the existing -`golden_e2e_v1.csv` is single-turn by construction and cannot measure any of -this. Counters land in `rag/metrics.py` and on the existing Grafana dashboard. - -## Consequences - -**Accepted.** More moving parts and more tokens per turn; a stateful service -where there was a stateless one; a summary that must be kept free of clinical -content by rule rather than by mechanism. - -**Refused.** An LLM confidence score as the loop's uncertainty signal. The -signals used are the resolver states that already exist — ambiguous drug, -unresolved section, multi-attribute question — because they are deterministic, -testable, and explainable to a reviewer. "The model felt 0.73 sure" is not a -defensible basis for asking or not asking a clinician a question. - -**Unchanged.** Nothing here lets the system assert a figure absent from the -retrieved source. Reasoning chooses *what to look up and how to say it*; it is -not a source of facts. diff --git a/docs-legacy/adr/0008-llm-understanding-one-shot-rag.md b/docs-legacy/adr/0008-llm-understanding-one-shot-rag.md deleted file mode 100644 index 1f82e33..0000000 --- a/docs-legacy/adr/0008-llm-understanding-one-shot-rag.md +++ /dev/null @@ -1,153 +0,0 @@ -# ADR 0008: LLM query understanding + one-shot grounded RAG (what is actually live) - -**Status:** accepted, live since 2026-08-06 (F-03), extended 2026-08-07 -**Supersedes:** ADR 0007 (conversational reasoning RAG — the `Focus`/ -`ConversationState`/TTL state design and the PLAN/RETRIEVE/ASSESS/REFINE/ -VERIFY bounded loop). ADR 0007's own `rag/conversation.py`/`rag/reasoning.py`/ -`rag/conversational.py` were deleted 2026-08-07 once confirmed unreachable -from `bootstrap.py` — see the notice at the top of ADR 0007 for the full -reasoning. -**Extends:** ADR 0006 (quarantined block references) — unchanged and still -binding: a chunk with `has_quarantined_content` still forces `VERIFY_PDF` -and is never generated over. - -## Context - -This ADR exists because `docs/architecture.md` and ADR 0007 described a -design that was never fully built, and the modules that partially -implemented it were never wired into `bootstrap.py`. A 2026-08-06 -independent 7-agent audit found this the hard way — it cost real time -establishing that `QdrantRetriever.search()` (dense vector search) and the -entire reasoning-loop module set were dead code, contradicting what the -docs claimed was live. The fix is not "finish building ADR 0007" — the -project deliberately moved to a simpler design that already works, proven -across many real multi-turn conversations (see `docs/progress-log.md`, -2026-08-05 through 2026-08-07 entries). This ADR documents that design so -the next reader doesn't have to re-discover it by audit. - -## Decision - -### 1. One LLM call understands the turn; no separate state object - -`rag/understanding.py::LlmQueryUnderstander.understand(turn, history)` reads -the raw current turn plus a **plain list of past turn strings** -(`"Người dùng: …"` / `"Trợ lý: …"`, kept by `RagAgent._history`, a -per-conversation-id in-process dict) and returns a `QueryFrame`: turn type, -resolved `drug_id`s (validated against a candidate set a deterministic -fuzzy/alias pass bounds *before* the model runs — F-04), section attribute, -population, weight, age, indication, route, and a `needs_clarify`/ -`clarify_reason`/`quick_replies` triple. - -There is no `Focus` struct, no TTL, no separate summariser. The model -re-reads the same history window (last `HISTORY_TURNS * 2` = 12 lines) every -turn and re-derives what's still relevant — cheaper to build and, so far, -more robust than hand-maintained state: it naturally handles "còn trẻ em thì -sao?" and short replies to its own clarify questions (population/route/etc. -— the latter only after a 2026-08-07 fix; see progress-log) without a -resolver state machine to keep in sync. - -**Known gap, inherited from ADR 0007 and still open:** this history is an -in-process dict — lost on restart, not shared across workers if the service -ever scales beyond one. ADR 0007's `PostgresConversationStore` was never -built either. - -### 2. Routing is a single dispatch, not a loop - -`RagAgent._route()` reads `frame.turn_type` and dispatches once: -`interaction` (2+ drugs) → gather each drug's evidence, combine, decide; -`drug_attribute`/`drug_overview`/`dosing_calc`/fallback → one drug, one -retrieval call; `smalltalk`/`out_of_scope` → canned reply, no retrieval; -`symptom_to_drug` with no drug named → an honest "not built yet" clarify. -There is no PLAN/REFINE step and no retrieval-round budget, because there is -only ever one retrieval call per turn. - -### 3. Retrieval is deterministic routing, not similarity ranking - -`RetrievalService.retrieve_framed(drug_id, section_key, query)`: -- `section_key` given (the dominant case, since `understand()` almost always - resolves it) → `find_by_section`, an **exact Qdrant payload filter** - (`drug_id` + `section_key`), returning the whole section as a scroll. - Score is a hardcoded 1.0 — this is a filter, not a ranked search, and nothing - here is "confidence" in the sense ADR 0007's retrieval-confidence gate meant. -- No section resolved → `find_by_drug` (whole monograph, book order), - trimmed to identity sections for a bare name or reranked (Cohere - cross-encoder over the ~29 sections of that one drug, not a corpus search) - for a free-form question. -- `QdrantRetriever.search()` — real dense vector similarity over the whole - corpus — exists and is unit-tested, but `RagAgent` never calls it. It is - reachable only through the legacy `RetrievalService.retrieve()` entry - point, itself only reachable when `ANSWER_PROVIDER=disabled` (no agent - configured at all — retrieval-only mode). `docs/architecture.md`'s - "Retrieval-confidence gate: below a similarity threshold, skip the LLM - call entirely" describes this legacy-only path, not the live one; that - section has been corrected to say so. -- Measured, and the reason this design was chosen over similarity ranking - for the live path: routing by exact `section_key` moved contraindication - hit@1 from 0.05 to 1.00 (`[[project-retrieval-quality-gap]]`, 2026-08-04). - A quarantined chunk anywhere in the retrieved set still forces the whole - result to `VERIFY_PDF` (`RetrievalService.decide`, a public wrapper added - 2026-08-07 so `RagAgent._interaction` applies the same policy to a - combined multi-drug evidence pool instead of hand-rolling it). - -### 4. Generation is one call, verified twice, with no confidence score - -`GroundedAnswerService.answer_from_result`: sufficiency-check (ask instead of -guessing when the evidence spans multiple populations/routes and the turn -hasn't disambiguated) → generate → `grounding.verify` (every number and -citation traces to the block it cites) → `_verify_entailment` (a second LLM -pass confirming each cited claim's *content*, not just its numbers, is -actually stated by that block; one same-claim retry on a lone reject, since -this call is measurably noisy — 2026-08-06 finding). A generation that fails -any check **abstains** — it does not fall back to a raw extractive quote -when a generator is configured (`[[feedback_no_extractive_fallback_when_llm_configured]]`). - -No `MAX_LLM_CALLS`/`MAX_WALL_CLOCK_MS` budget object exists. Each call is -bounded only by its own provider timeout. **This is ADR 0007's F-08 finding, -inherited unchanged and still open** — a real end-to-end request deadline -threaded through `RagAgent`'s sequence of up to 5 sequential Bedrock calls -(understand → sufficiency → generate → ≤2 entailment) is real remaining -work, not solved by this ADR. Measured live 2026-08-07: a single answerable -turn costs ~8-9s wall clock, ~75-80% of it the 4 sequential LLM calls -(understand ~2.6-3.3s dominates — an 80B model doing a classification task -that likely doesn't need one); a clarify chain compounds this linearly since -each round is a fresh request repeating the same call sequence from scratch. - -### 5. Context resolved across turns is folded into one self-contained string - -Added 2026-08-07, closing a P0 the 2026-08-06 audit named: `frame.population`/ -`weight_kg`/`age_text`/`route`/`indication` were extracted by `understand()` -but never reached `retrieve_framed`/`answer_from_result`, which took only -the bare current-turn text — so a reply like "Uống" three turns into a dose -conversation reached the sufficiency/generation LLM calls as literally just -"Uống", with no notion that population=adult was already established two -turns back. `RagAgent._synthesize_query` now folds every resolved field into -one string (`"Uống. Đối tượng: người lớn. Đường dùng: uống."`) before it -reaches retrieval's rerank signal and generation's `query` argument. No-op -for a fresh single-shot question that already states its own context. - -## Consequences - -**Accepted.** No confidence score, no retrieval-round budget, no PLAN/REFINE -step — the tradeoff ADR 0007 explicitly refused ("an LLM confidence score... -is not a defensible basis for asking or not asking a clinician a question") -is exactly what this design uses instead (an LLM sufficiency/clarify -judgment), because in practice it has been reliable enough and dramatically -simpler to build, extend (route/quick_replies were one schema field + one -prompt rule each, not a new state machine), and debug — every session this -month that touched the ADR 0007 modules found new bugs in the state-machine -edges (TTL boundaries, Focus inheritance correctness) rather than in the -domain logic itself. - -**Refused (again, restated from ADR 0007, still true):** an LLM confidence -score as a hard gate for retrieval — `RetrievalService.decide`'s -`VERIFY_PDF`/`ABSTAIN` decisions remain deterministic (quarantine flag, -missing provenance), never a model's self-reported certainty. - -**Still open, named rather than hidden:** -- No request-scoped time/call budget (F-08). -- Conversation history is in-process, not durable/shared (inherited from - ADR 0007, never built either way). -- No production-path adversarial regression suite beyond one live-verified - end-to-end case (F-10's remaining scope). -- `dosing_calc` (a real mg/kg calculator) and `symptom_to_drug` (reverse - indication lookup) remain honest "not ready" clarifies, not answers. diff --git a/docs-legacy/adr/0009-no-rag-framework.md b/docs-legacy/adr/0009-no-rag-framework.md deleted file mode 100644 index 60033c6..0000000 --- a/docs-legacy/adr/0009-no-rag-framework.md +++ /dev/null @@ -1,82 +0,0 @@ -# ADR 0009: No RAG framework — hand-written orchestration behind ports - -## Status - -Accepted. **Recorded retrospectively** during the 2026-08-12 documentation pass: -the decision is unambiguous in the implementation, but no ADR existed for it. - -## Context - -The system performs retrieval-augmented generation with query understanding, -multiple retrieval strategies, reranking, prompt construction, structured output -parsing, and post-generation verification — the exact feature set LangChain and -LlamaIndex exist to provide. - -## Decision - -Neither framework is used. There is no RAG or agent library of any kind. - -Verifiable from the repository: - -- `apps/ai-service/pyproject.toml` declares six runtime dependencies: - `fastapi`, `httpx`, `psycopg`, `pydantic-settings`, `qdrant-client`, - `uvicorn`. Optional extras add `prometheus-client`, `anthropic` and three - OpenTelemetry packages. -- `apps/ai-service/Dockerfile` installs that set plus `boto3`. -- No file imports `langchain`, `llama_index`, `haystack` or any equivalent. - -Instead: - -- Orchestration is a plain class with an explicit branch table - (`rag/agent.py::_route`). -- Prompts are module-level constants with JSON schemas (`rag/prompt.py`). -- Providers are injected through `typing.Protocol`s (`rag/ports.py`) and - implemented in `adapters/`, which is the only package importing an SDK — and - always lazily, inside a method. -- `bootstrap.py` is the single composition root. - -## Consequences - -**Enabled by this choice** - -- `rag/` imports no SDK, so the entire domain — including every safety check — - is unit-testable offline with stub objects. All 278 ai-service tests run in - 2.6 s with no network. -- Behaviour is inspectable: the retrieval route for a given turn is a readable - `if` chain, not framework dispatch. -- Failure semantics are chosen per call site. The fail-closed/fail-open - asymmetry in [02-system-architecture.md](../02-system-architecture.md#failure-boundaries) - is deliberate and would be hard to express through a framework's uniform - error handling. -- Prompt text is reviewable as domain policy in one file, and swapping providers - cannot silently change what the model was told. - -**Costs** - -- Retrieval strategies, rank fusion, context packing and evaluation harnesses - are all hand-written. Two of them (`fusion.py`, `expansion.py`) were written - and never wired ([27-technical-debt.md](../27-technical-debt.md#d-12--dead-code-three-tested-modules-with-no-runtime-caller)). -- Optional retriever capabilities are discovered with `getattr` rather than - declared, so the real interface is wider than `ports.py` documents (D-14). -- No community tooling for tracing, caching or evaluation applies; the - observability layer is bespoke. - -## Rationale - -Partially recoverable. The code does not state "we chose not to use a -framework", but the ports-and-adapters discipline is documented repeatedly in -module docstrings, and one of them makes the intent explicit — -`rag/understanding.py`: - -> `rag/` imports no SDK: the LLM is injected as a `JsonLlm` protocol … and a -> deterministic stub runs the whole path offline in tests. - -`rag/prompt.py` gives the parallel reason for prompts: - -> This is domain policy, not infrastructure … it lives here so it can be read, -> reviewed and tested without an SDK, and so swapping the provider cannot -> silently change what the model was told. - -The consistent theme is testability and reviewability of the safety layer. -Whether cost, lock-in or framework maturity also weighed in the decision is not -recoverable from the repository. diff --git a/docs-legacy/adr/0010-interim-single-host-compose-deployment.md b/docs-legacy/adr/0010-interim-single-host-compose-deployment.md deleted file mode 100644 index a9942c2..0000000 --- a/docs-legacy/adr/0010-interim-single-host-compose-deployment.md +++ /dev/null @@ -1,85 +0,0 @@ -# ADR 0010: Single-host Docker Compose as the interim deployment - -## Status - -Accepted. **Recorded retrospectively** during the 2026-08-12 documentation pass. -Does **not** supersede [ADR 0002](0002-argocd-gitops.md), whose own status line -says it remains the target: - -> **Accepted — still the target, not yet implemented.** Not superseded by the -> current production setup. - -## Context - -ADR 0002 chose GitOps on the team's ArgoCD instance. A complete Helm chart -(`infra/helm/medical-chatbot/`) and three ArgoCD `Application` manifests exist. -Neither has been applied: each `Application` carries three unresolved `TODO`s -(project/RBAC scope, repo URL, target cluster), `infra/k8s/base|overlays/` hold -only `.gitkeep`, and no image registry is configured anywhere. - -Meanwhile the product is live at `https://realvuxbaro.me`. - -## Decision - -Run production as Docker Compose on a single EC2 host, with Caddy terminating -TLS, and deploy by SSH from GitHub Actions. - -Verifiable from the repository: - -- `infra/docker/docker-compose.prod.yml` — postgres, qdrant, ai-service, web, - caddy, with named volumes. -- `infra/docker/docker-compose.observability.yml` — the OTel/Prometheus/Tempo/ - Grafana overlay, which also sets `OTEL_ENABLED=true`. -- `infra/docker/Caddyfile` — `realvuxbaro.me` → `web:3000`, `/grafana/*` → - `grafana:3000`. -- `.github/workflows/deploy.yml` — `appleboy/ssh-action`, `git reset --hard`, - `docker compose up -d --build`, `caddy reload`, `python -m migrate`, then ~18 - assertions. - -## Consequences - -**Accepted trade-offs** - -- Images are built on the production host and are untagged, so there is **no - artifact to roll back to**; recovery is a revert commit plus a rebuild. -- Deploys are in-place, with brief per-service downtime. -- No horizontal scaling. That happens to align with the in-process agent state - described in [02-system-architecture.md](../02-system-architecture.md#the-stateful-detail-that-constrains-scaling), - but the alignment is coincidental, not enforced. -- Configuration and secrets live in an uncommitted `.env.prod` on the host, so - production configuration cannot be reviewed in Git. -- `postgres` and `qdrant` are deliberately absent from the workflow's `up -d` - list, so a code deploy never restarts the stateful services — and changes to - their service definitions do not take effect until someone restarts them. - -**Preserved despite the simpler runtime** - -The deploy script asserts far more than a Compose deploy usually does: service -health, a **real grounded answer** from the real corpus (`decision=answerable` -with a `chi_dinh` citation), both Grafana datasources, the provisioned -dashboard, public reachability of `/grafana/login`, and end-to-end trace -propagation by asserting that a specific `X-Trace-ID` becomes retrievable from -Tempo. That verification block is what makes the simpler runtime defensible. - -**Migration path** - -The Helm chart already maps every setting in `config.py` to a ConfigMap, mounts -`POSTGRES_DSN` from a Secret, and configures readiness/liveness/startup probes -against the same `/ready` and `/health` endpoints Compose uses. Moving to -Kubernetes therefore needs: an image registry and tagging, a corpus-load or -snapshot-restore step (the chart provisions an **empty** Qdrant, against which -`ai-service`'s manifest check refuses to start), the three ArgoCD `TODO`s -resolved, and the `bump-image-tag` workflow that -`infra/ci/github-actions/README.md` describes but does not contain. - -## Rationale - -**Decision observed; rationale not fully recoverable from the repository.** The -Compose header comment records one constraint — - -> No GPU, no team k3s — Bedrock calls go out over the instance's IAM role … so -> no AWS access keys live in this file or its env files. - -— and ADR 0002 remaining un-superseded shows the Kubernetes target was not -abandoned. Beyond that, whether the driver was cost, cluster access, or time to -first deployment is not determinable from the code. diff --git a/docs-legacy/adr/README.md b/docs-legacy/adr/README.md deleted file mode 100644 index be0f74b..0000000 --- a/docs-legacy/adr/README.md +++ /dev/null @@ -1,20 +0,0 @@ -# Architecture decision records - -| ADR | Title | Status | Reflected in code? | -|---|---|---|---| -| [0001](0001-vector-db-qdrant.md) | Use Qdrant as the vector database | Accepted | **Yes** — `adapters/qdrant.py`, `ingestion/load/qdrant_repo.py` | -| [0002](0002-argocd-gitops.md) | Use the team's existing ArgoCD instance for deployment (GitOps) | Accepted — target, **not yet implemented** | **No** — production is Docker Compose on EC2 ([20](../20-deployment.md)) | -| [0003](0003-pdf-parsing-strategy.md) | PDF parsing strategy, validated empirically | Accepted | **Yes** — `ingestion/extract/`, `ingestion/segment/detector.py` | -| [0004](0004-chunking-strategy.md) | Chunking strategy for drug monographs | Accepted (monograph range only) | **Yes** — `ingestion/chunk/chunker.py` | -| [0005](0005-segment-output-contract-for-chunking.md) | `segment/` output contract needed by `chunk/` | Proposed; header says "contract only, no implementation" | **Yes, now implemented** — `segment/models.py` + `chunk/` follow it. The status line is stale | -| [0006](0006-quarantined-block-references-in-chunks.md) | Chunks must carry references to lifted table/formula blocks | Accepted, implemented in schema v4 | **Yes** — `ChunkAttachment`, `has_quarantined_content`, the ADR-0006 gate set | -| [0007](0007-conversational-reasoning-rag.md) | Conversational reasoning RAG (state + bounded loop) | **Superseded by 0008** | **No** — `rag/conversation.py` and `rag/reasoning.py` no longer exist | -| [0008](0008-llm-understanding-one-shot-rag.md) | LLM query understanding + one-shot grounded RAG | Accepted, live since 2026-08-06 | **Yes** — `rag/understanding.py`, `rag/agent.py`, `rag/answer.py` | -| [0009](0009-no-rag-framework.md) | No RAG framework — hand-written orchestration behind ports | Accepted (recorded retrospectively) | **Yes** | -| [0010](0010-interim-single-host-compose-deployment.md) | Single-host Docker Compose as the interim deployment | Accepted (recorded retrospectively) | **Yes** | - -ADRs 0009 and 0010 were written during the documentation pass described in -[DOCUMENTATION_PLAN.md](../DOCUMENTATION_PLAN.md). They record decisions that are -unambiguously visible in the implementation but had no ADR. Where the rationale -could not be recovered from the repository, they say so rather than inventing -one. diff --git a/docs-legacy/pdf-parsing-outlier-catalog.md b/docs-legacy/pdf-parsing-outlier-catalog.md deleted file mode 100644 index 49d3fcc..0000000 --- a/docs-legacy/pdf-parsing-outlier-catalog.md +++ /dev/null @@ -1,888 +0,0 @@ -# PDF Parsing Outlier Catalog - -A generalized checklist of structural risks found while parsing -`duoc-thu-quoc-gia-viet-nam-2018.pdf` (1668 pages). Every item here was -**confirmed with real evidence** (bounding-box inspection, cross-tool -comparison, or a whole-document scan) — not assumed. The goal of this -document is reuse: if this project (or a future one) needs to parse another -structured reference PDF — another national formulary, a different -government-published multi-part document, any dense print-layout book — -this is the checklist of "things that go wrong that a small page sample -won't reveal," and how to actually check for each one cheaply (most checks -here run over the whole 1668-page book in under a minute). - -For the narrative investigation and drug-formulary-specific numbers, see -`docs/adr/0003-pdf-parsing-strategy.md`. This document is the distilled, -reusable checklist form of the same findings, plus items found afterward. - ---- - -## Structural discovery risks (before you even parse content) - -### 1. No bookmarks/TOC -**What it looks like:** `doc.get_toc()` (PyMuPDF) returns an empty list. -**Why it matters:** the obvious, easiest structural signal for section -boundaries simply doesn't exist — don't design a pipeline that assumes it -will. -**Check:** one line, `len(doc.get_toc())`. Do this first, always, before -assuming a bookmark-based approach. -**Generalizes:** yes, directly — always check this before designing around -bookmarks, for any PDF. - -### 2. Shallow/unusable tagged-PDF structure tree -**What it looks like:** the PDF has a `/StructTreeRoot` (looks promising — -"tagged PDF"), but it only covers a handful of generic `/H1`/`/P` elements -for a fraction of the document (here: ~29 elements for 1668 pages). -**Why it matters:** easy to assume "tagged PDF = rich semantic structure -available"; in practice many tagging tools produce a minimal -compliance-only tree that covers almost nothing. -**Check:** walk the struct tree (`doc.xref_object` on `/StructTreeRoot`, -recurse into `/K`) and count real leaf elements vs. total page count. If the -ratio is tiny, it's not a usable data source. -**Generalizes:** yes — always verify depth/coverage before trusting a -struct tree, don't just check for its existence. - ---- - -## Page layout risks - -### 3. Multi-column body layout -**What it looks like:** body pages are genuinely two-column (confirmed via -bounding boxes: left column x≈44-299, right column x≈308-562, page width -≈595). Front-matter pages that *look* like a multi-column name grid to the -eye turned out, on inspection, to be single wide text blocks with internal -whitespace padding between names — not a real structural column split. -**Why it matters:** a naive "read text top-to-bottom regardless of x" pass -would interleave left- and right-column content into nonsense. Conversely, -assuming every visually grid-like page is column-split leads to wasted -effort — verify per page/section, don't generalize from appearance alone. -**Check:** for any suspicious page, dump block bounding boxes -(`page.get_text("dict")["blocks"]`) and look at the actual x0/x1 ranges. A -real column split shows two clusters of x-ranges; a padded single-column -list shows one wide range per line. -**Handling:** PyMuPDF's default block-level reading order handled the real -two-column case correctly here (validated against a known monograph) — the -tool most likely to get column order wrong was `pdfplumber`'s general -`extract_text()` (see item 8), not PyMuPDF. -**Generalizes:** yes — this exact check (dump bboxes, look at x-clusters) -works on any PDF to determine real column count before writing extraction -logic. - -### 4. Full-width content breaking out of the column grid -**What it looks like:** some pages have a table (or could have a figure) -that spans nearly the entire page width (confirmed: a body-surface-area -lookup table's blocks span x≈35 to x≈553, i.e. across both normal columns), -overriding the page's usual two-column layout. -**Why it matters:** logic written to always split a page into "left column" -and "right column" text will misbehave on these pages — the content isn't -in either column, it's a single full-width unit. -**Check:** for any block, compare its x-width against the known -single-column width; if a block's x-range spans (or nearly spans) both -known column ranges, treat it as a full-width unit, not part of a column. -**Generalizes:** yes — any multi-column layout can have occasional -full-width breakout elements (tables, figures, pull-quotes); always check -for this rather than assuming rigid column adherence everywhere. - ---- - -## Table-specific risks - -### 5. Tables split across a page break lose their header on the continuation page -**What it looks like:** confirmed directly — "Bảng 4: Xử trí về điều trị ARV -theo mức độ phát ban" (a 3-column table) starts on one page with its header -row (`['Mức độ', 'Biểu hiện', 'Xử trí']`) and 3 data rows; its 4th data row -("Mức độ 4...") appears on the **next page**, extracted by `pdfplumber` -as a **separate table object with no header row at all**. -**Why it matters:** if a pipeline treats each `find_tables()`/ -`extract_tables()` result as an independent, self-contained table, the -orphaned continuation row is meaningless on its own — you lose the column -semantics for that row entirely. -**Check:** for any table-like structure, check whether the page/column -immediately preceding it ends with a same-shaped table lacking a natural -final row (e.g. an incomplete-looking sequence) — a strong heuristic is -"table starts at the very top of a page/column, no header, same column -count as the table ending at the bottom of the previous page/column." -**Handling:** never treat page-extracted tables as independent; track -continuation explicitly and re-attach the original header to orphaned -continuation rows before using them. -**Generalizes:** yes — this is a generic multi-page-table risk in any -paginated PDF with tall tables; the detection heuristic (position at -page/column top + no header + matching column count to the previous -table) applies broadly. - -### 6. Tables can also split across a column boundary on the *same* page -**What it looks like:** confirmed — "Bảng 6" (ARV drug toxicity table) -starts in the left column near the bottom of a page (header + first data -row) and its remaining data rows appear at the **top of the right column of -the very same page**, again with no header repeated. -**Why it matters:** this is easy to miss because there's no literal page -break — it's tempting to assume "if it's the same page, it's not split," -but a table can still be taller than one column's usable height. -**Check:** same heuristic as item 5, but also check column position, not -just page number — a header-less table fragment starting at the top of a -column (regardless of page) is a suspect continuation. -**Generalizes:** yes, wherever content flows in columns at all — this risk -exists any time column height is shorter than table height. - -### 7. Two-dimensional grid/nomogram tables are not linearly recoverable -**What it looks like:** confirmed — a body-surface-area lookup table -(height across the top, weight down the side, a BSA value at each -intersection) extracts as a scrambled sequence of numbers with no -recoverable row/column association from plain text alone (e.g. `"0,50 -0,52 0,54 0,56"` followed by `"0,55 0,57 0,59 0,61"` — these are almost -certainly column-wise fragments, not the visual rows). -**Why it matters:** unlike a normal bordered table (rows of related -values), a 2D lookup grid's *meaning* depends entirely on 2D position — a -number is meaningless without knowing both its row header (weight) and -column header (height). Flattened text extraction destroys exactly the -information needed to interpret it. -**Check:** any table where extracted "cells" are bare numbers with no -inline label, laid out in a dense grid, is a candidate — cross-check -against the source's own stated formula/description (this table is -explicitly a lookup version of a stated formula, see item 8). -**Handling:** for RAG purposes, prefer **not** to chunk this table as -literal text at all; either (a) reconstruct it properly using per-number -bounding-box position matched against header row/column bboxes (real 2D -table reconstruction, non-trivial), or (b) rely on the accompanying formula -being available for the LLM to compute from directly, and explicitly flag -this table's raw text as unreliable/do-not-cite in metadata. -**Generalizes:** yes — any nomogram, nutrition-fact grid, or nCk-style -lookup table in any PDF has this exact problem; detect by the "bare number -grid" pattern, don't assume normal table extraction works. - ---- - -## Formula / equation risks - -### 8. Formula rendering is inconsistent — some survive as linear text, some don't -**What it looks like:** two real formulas found, two different outcomes. -The Du Bois body-surface-area formula (simple inline exponents, -`"S = W0,425 × H0,725 × 71,84"`) extracted **cleanly as readable text**. The -Cockcroft-Gault creatinine-clearance formula (a stacked fraction — -numerator over denominator, visually 2D) extracted as **scattered, -disordered fragments** with no linear reading order. -**Why it matters:** it's tempting to write one rule ("formulas are -unreliable, always flag them") or its opposite ("formulas extract fine, no -special handling needed") — neither is true here. The determining factor is -whether the formula's visual layout is fundamentally 1D (left-to-right, -like an inline exponent) or 2D (a fraction, a matrix, stacked terms). -**Check:** a detector now exists — `residual_ink.py`'s -`fraction_bar_candidate`, which finds the bar as ink no extracted span -accounts for. Measured on this book: **precision 16/23 = 69.6%** (the misses -are decorative underlines and table borders), recall unknown, and it is blind -by construction to a fraction printed without a bar (item 25). Its output is -therefore a review queue, not a verdict: all 23 candidates were rendered and -read one at a time before any was acted on, and only the confirmed ones went -into `ingestion/data/verified/formula_regions_2d.json`. -**Generalizes:** yes — any technical/medical/scientific PDF with inline -math will have this exact split; don't assume all formulas behave the same -way in extraction. - ---- - -## Character/glyph-level risks - -### 9. Rare reversed/misordered glyph defects — corrected count: 2, not 1 -**What it looks like:** re-implemented as real, tested production code -(`ingestion/ingestion/extract/glyph_order.py`) rather than trusted from the -earlier exploratory script's claim. Found **two distinct shapes**, not the -one originally reported: -1. **Within-span character reversal** (physical page 1373, the originally - reported case): one span's glyphs are positioned in descending x-order, - producing `" = tịx 8 yàgn gnàh uềil gnổt(..."`, which reverses - character-by-character back to `"(4 xịt = 800 microgam) vào buổi - chiều..."`. -2. **Cross-fragment row misordering, newly found** (physical page 714): a - single visual row is split by PyMuPDF into multiple `line` objects - *within one block* that are then emitted out of left-to-right order — - each fragment's own characters are fine, but concatenating fragments in - extraction order produces `"...bảo quản ộđ tệihn "` instead of the - correct `"...bảo quản nhiệt độ "`. This is a different underlying shape - from item 1 (multiple mis-ordered fragments, not one reversed span) and - was missed by the original narrower (within-span-only) check — the - ADR 0003 claim of "exactly 1 occurrence in the whole book" undercounted - the real defect population; corrected here. -**Getting a trustworthy count took three detector iterations** (documented -in the module's own docstring) — the first naive whole-book implementation -of the row-level check reported **1113** "issues," almost all false -positives from two mechanisms: (a) ordinary font-kerning jitter (e.g. in -"mefloquin," two adjacent glyphs differ by 0.095pt — normal kerning, not a -defect) treated as a reversal with no decrease-tolerance, which then -actively *corrupted* correct text into "mefolquin"; and (b) reconstructing -"visual rows" from raw x/y coordinates using a hand-picked column-boundary -threshold, which misclassifies a paragraph that happens to start near the -natural column gap (confirmed real case: a right-column paragraph starting -at x=299.4 got merged with an unrelated left-column paragraph at the same -y). The fix that survived whole-book testing: group by PyMuPDF's own -`block` index (already validated in ADR 0003 to respect this document's -column structure) instead of re-deriving columns from coordinates, plus a -minimum-decrease threshold (1.0pt — safely between the ~0.3pt kerning noise -floor and the >2pt real-defect magnitude). Final whole-book result: **11 -row-level issues on 5 pages** — 3 of those pages (92, 94, 805) are formula -regions already flagged as unreliable in item 8 below (2D-layout formulas -scramble on extraction; this check's "corrected" text for those rows should -**not** be trusted or auto-applied, same as item 8's existing guidance), -leaving exactly the 2 genuine prose defects above (pages 714, 1373). -**Why it matters:** both genuine defects are confirmed real data-corruption -risks, not theoretical — but both are also extremely rare (2 occurrences in -1668 pages of prose), so they must be *detected*, not assumed either absent -or common. Equally important: a naive implementation of "the obvious check" -can itself introduce false positives and even actively corrupt correct -text — this detector's own false-positive history is as important a lesson -as the defects it catches. -**Check:** `ingestion.extract.scan_glyph_order` (within-span) and -`ingestion.extract.scan_reading_order` (cross-fragment, grouped by real -PyMuPDF block index + row y, with a 1.0pt minimum-decrease threshold and -header-band exclusion). Both run in seconds over the full book. -**Generalizes:** yes, directly — this is a cheap, universal sanity check -worth running on any PDF text-extraction pipeline as a standing QA gate, -regardless of source document. The false-positive history also generalizes: -any "reconstruct visual rows from raw coordinates" approach needs a -decrease-tolerance (font kerning is universal) and should prefer the -source tool's own layout-analysis groupings (blocks/lines) over hand-picked -coordinate thresholds wherever available. - ---- - -## Section/heading detection risks - -### 10. Font size is not a reliable heading signal — bold is -**What it looks like:** confirmed two genuine, equally top-level monograph -titles at different font sizes (10.0pt and 9.5pt). An early detector -gated on `size >= 9.8` and silently dropped ~15% of real monographs as a -result. -**Why it matters:** a threshold calibrated from one or two examples will -look correct until validated at scale — this is the single clearest -"don't generalize from a small sample" lesson from this whole -investigation. -**Check:** whole-document validation against an independent ground truth -(here, the back-of-book page-numbered index) is what caught this — a -sample of 2-3 pages would not have. -**Generalizes:** yes — for any PDF, prefer a binary style signal (bold/not -bold, a specific font name) over a numeric threshold (size, weight value) -wherever possible, and always validate any numeric threshold against the -whole document, not a handful of examples. - -### 11. Multi-line wrapped titles/headings must be merged before matching -**What it looks like:** confirmed as the dominant cause of missed -detections in whole-document validation — long titles (e.g. "CÁC CHẤT ỨC -CHẾ HMG-CoA REDUCTASE", "THUỐC TƯƠNG TỰ HORMON GIẢI PHÓNG GONADOTROPIN") -wrap across 2+ physical lines; a per-line detector catches only fragments, -which then fail to match a name-based ground truth AND can produce false -name collisions with an unrelated single-line heading elsewhere in the -document (this happened: a wrapped title's second line, "GONADOTROPIN", -collided with a genuine, different, single-line "GONADOTROPIN" monograph -elsewhere). -**Check:** whole-document recall measurement against ground truth; misses -clustered around long/compound names are the signature of this bug. -**Handling:** merge consecutive bold+all-caps lines (with compatible -positioning) into one candidate title before matching/keying, rather than -treating each line independently. -**Generalizes:** yes — any document with long titles/headings that can wrap -will have this exact failure mode; always merge candidate multi-line -headings before using them as unique keys. - -### 12a. Class-level monographs cover multiple active ingredients (multiple ATC codes) — this is NOT rare -**What it looks like:** first noticed via two incidental examples -("GONADOTROPIN", "VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ"), then actually measured -across the whole 680-monograph corpus (not assumed from the 2 examples — -this distinction matters, see below). **Real, whole-corpus number: 173 of -680 detected monographs (25.4%) have more than one distinct ATC code**, -ranging up to extreme cases — INSULIN alone lists **20** different ATC -codes, BETAMETHASON and DEXAMETHASON 11 each, PREDNISOLON 10, -HYDROCORTISON 9. This is a quarter of the entire corpus, not a couple of -edge cases — the 2 incidental examples badly understated how common this -is, and stating "found 2 examples, pattern confirmed" without the -whole-corpus count would have been exactly the kind of unverified claim -this project's own validation standard now forbids. -**Even the 25.4% is a floor, not the true number** — see item 12c below: -ATC-code text-extraction noise (stray whitespace, O/0 confusion) caused -some genuinely multi-ATC monographs (e.g. "TRIAMCINOLON", 5 codes) to be -undercounted by a naive regex. The true proportion is measurably higher -than 25.4%; re-measure after fixing the regex, don't keep citing 25.4% as -final. -**Why it matters:** a data model that assumes "one monograph = one drug = -one ATC code" is wrong for roughly a quarter or more of the corpus. -**Handling:** store ATC code (and dosage-form sub-entries) as a **list** -per monograph, not a scalar; when chunking, consider whether a -class-level monograph's sections should be tagged with the whole class -name, the specific sub-compound, or both, depending on what the retrieval -use case needs. -**Generalizes:** yes — any reference work organized primarily by drug -class or by generic substance will have entries that don't map 1:1 to a -single identifier. More importantly, the *methodology* generalizes: when -you notice a pattern from 1-2 examples, measure its real prevalence across -the whole corpus before deciding how much engineering effort it deserves — -"found 2 examples" and "25.4% of everything" call for very different -levels of investment, and you can't tell which one you're dealing with -without the whole-corpus count. - -### 12c. ATC codes (and likely other structured codes) have real text-extraction noise -**What it looks like:** while investigating why 22/680 (3.2%) monographs -appeared to have zero ATC codes, spot-checked 14 of them directly and found -**two distinct, confirmed causes**, both text-extraction noise rather than -missing content: -- **Stray internal whitespace** splitting one code into two tokens, e.g. - `"L01X X02"` (should be `L01XX02`), `"J04A C01"` (should be `J04AC01`), - `"N05B A06"` (should be `N05BA06`). -- **Digit/letter confusion**: a literal "0" rendered/typeset as the letter - "O", e.g. `"NO3AX12"` (should be `N03AX12`), `"JO1DC07"` (should be - `J01DC07`). -A relaxed regex tolerating both patterns resolved **9 of the 14** spot-checked -cases as real ATC codes hiding behind extraction noise. The **remaining -~5 of 14** were genuinely different: the source text explicitly states -`"Mã ATC: Chưa có."` or `"Mã ATC: Không có."` ("not yet available" / "none") -— a real, valid data state, not an error, and not something to paper over -as if a code exists. -**Why it matters:** a strict ATC-code regex silently undercounts real ATC -data; distinguishing "extraction noise hiding a real code" from "the book -says there is no code" requires checking the actual field text, not just -whether a regex matched. -**Handling:** normalize ATC-code-shaped text before matching (strip internal -whitespace between the letter/digit groups, treat a digit-position "O" as -"0") and explicitly check for the "Chưa có"/"Không có" literal strings as a -valid "no ATC" state rather than a parse failure. -**Generalizes:** yes — any structured code/identifier extracted from a PDF -(product codes, classification codes, reference numbers) can suffer this -same whitespace-injection and O/0 confusion; validate structured-looking -fields against their expected format and investigate exceptions rather than -assuming a strict pattern match is reliable. - -### 12d. A section-title (part-divider) page can be falsely detected as a monograph -**What it looks like:** confirmed — the very first item in a whole-corpus -boundary scan was "CÁC CHUYÊN LUẬN THUỐC" (the literal title of Part 2 of -the book, "The Drug Monographs" — a part-divider heading, not a drug) at -physical page 98, picked up as a false-positive monograph boundary because -it happened to be bold, all-caps, short, and was followed (a few real -monograph-boundaries later) by some "Tên chung quốc tế" text from the -actual first real monograph. -**Why it matters:** without a whole-corpus scan this would have gone -unnoticed indefinitely — it doesn't look wrong from a single-page read of -Abacavir, and the discovery methodology this catalog is built on is -exhaustive scans, so this is a good example of a defect that only surfaces -at full scale. -**Handling:** exclude a small, known set of non-drug part/section-divider -strings ("CÁC CHUYÊN LUẬN THUỐC", "CÁC CHUYÊN LUẬN CHUNG", "CÁC PHỤ LỤC", -etc. — enumerable from the book's own table of contents) from the -monograph-boundary detector, or require the anchor phrase ("Tên chung quốc -tế") within a tighter line-distance so an unrelated real monograph several -lines away doesn't false-confirm a divider title. -**Generalizes:** yes — any document with part/section-divider title pages -styled similarly to its content headings (bold, prominent, short) risks -this exact false positive; explicitly exclude known structural/navigational -titles from content-boundary detectors. - -### 12b. Genuine spelling/capitalization typos exist in the source text -**What it looks like:** confirmed real example — the running header on the -Vitamin D monograph's continuation pages reads `"Vitamin d và các thuốc -tương tự"` (lowercase "d"), while the real ALL-CAPS heading correctly reads -`"VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ"`. This is a genuine typesetting mistake -in the 2018 print, confirmed via font/bbox inspection (same bold 10pt font -as the correct heading — not an extraction artifact, the source text itself -has the typo). The page's bottom running *footer* uses yet another variant, -the short form `"Vitamin D"` (correctly capitalized) — meaning the same -monograph has **three different boilerplate text variants** across one -page (top header with a typo, the real heading, bottom footer). -**Why it matters:** don't treat running headers/footers as a perfectly -clean, typo-free secondary signal (item 13 in this catalog already -recommends using them as a cross-check) — they can themselves contain -source-level errors. In this specific case, the detection heuristic -(strict ALL-CAPS requirement, item 10) happened to still work correctly, -because "Vitamin d và các thuốc tương tự" and "Vitamin D" are not fully -uppercase and so are correctly rejected as monograph-boundary candidates — -but this was not a designed defense against typos specifically, just a -side effect of the all-caps requirement. A future/different typo (e.g. an -accidentally all-caps running header) would not be caught the same way. -**Check:** no systematic typo-detection was built (out of scope — this is -about parsing robustness, not proofreading the source); the practical -takeaway is to keep relying on the strict structural signals (bold + all -caps + short + anchor phrase) as primary, and treat any single text-based -signal (including running headers) as fallible. -**Generalizes:** yes — any real-world print-to-PDF source will have some -rate of genuine typos/inconsistencies; parsing logic should be robust to -them by relying on multiple independent structural signals (font, -position, anchor phrases) rather than trusting any single text match to be -error-free. - -### 12e. Monograph length and section coverage vary enormously — measured, not assumed -**What it looks like:** across all 680 detected monographs, length ranges -from **2,331 to 45,623 characters** (~20x spread) and the number of known -section labels found per monograph ranges from as few as **8** up to -**20** (out of a ~19-20 item known vocabulary) — most cluster around -16-19, but the tails are real: "ASPARAGINASE"-adjacent short entries around -2,300-4,300 chars vs. "AMOXICILIN VÀ KALI CLAVULANAT" at 45,623 chars. -**Why it matters:** don't design chunking limits (e.g. a fixed max tokens -per monograph, or an assumption that "a monograph roughly fits in N -chunks") around a single example — the real distribution has a long tail -on both ends. -**Check:** this came from the same whole-corpus survey used for items 12a -and 12c — computing length and detected-section-count per monograph is -cheap and worth keeping as a standing sanity metric (e.g. flag any -monograph outside some percentile range for manual review). -**Generalizes:** yes — any corpus of "similar" documents (monographs, -product entries, articles) will have a real length/completeness -distribution; measure it before assuming uniformity. - -### 12. The documented taxonomy is not exhaustive — keep it open -**What it looks like:** the book explicitly documents a 19-field template -for every drug monograph (page 38), but real monographs contain at least -one undocumented extra field ("Tên thương mại" — brand/trade names) not in -that list. -**Why it matters:** treating a documented schema as a closed enum will -silently misclassify or drop real content that doesn't fit it. -**Generalizes:** yes — any document that describes its own structure in a -preface/README should still be validated against real instances; documented -schemas are frequently incomplete in practice. - ---- - -## Noise / boilerplate risks - -### 13. Header/footer boilerplate must be stripped, but can double as a signal -**What it looks like:** every page carries a page number and a repeating -string (`"DTQGVN 2"`), and body pages additionally carry a running header -naming the current monograph/section. -**Handling:** strip the fixed boilerplate before parsing content, but the -running monograph-name header is a **useful secondary cross-check** for -"which monograph is this page's body text currently part of" — don't -discard it as pure noise. -**Generalizes:** yes — running headers/footers are common in print-derived -PDFs and are usually worth extracting as metadata, not just filtering out. - -### 14. Blank/near-empty separator pages at section transitions are expected -**What it looks like:** exactly 6 near-empty pages (<20 characters) found -across the whole 1668-page book, and every single one lands exactly on a -major section-transition boundary (before general chapters, before -individual monographs, before appendices, near the book's end). -**Why it matters:** a naive pipeline might treat a near-empty page as an -extraction failure and error out or flag it, when it's actually an -intentional print-layout convention (forcing a new part to start on a -fresh page). -**Check:** whole-document scan for pages under some small character -threshold; cross-reference their positions against known section -boundaries before treating them as errors. -**Generalizes:** yes — this print convention is extremely common in -formally typeset books; always expect and gracefully skip near-empty pages -rather than treating them as failures. - ---- - -### 15. No embedded images anywhere in the book — measured, not assumed -**What it looks like:** a whole-book scan of `page.get_images(full=True)` across -all 1668 pages returns **zero** embedded raster/vector images, confirmed via -PyMuPDF's own image extraction API (not just "the text doesn't mention an -image"). -**Why it matters:** avoids over-investing in image/caption validation tooling -for a corpus that has no images to validate — but this must be a measured -fact, not an assumption from the book's general description as "text-heavy." -**Generalizes:** the check (`get_images(full=True)` summed over every page) -is a cheap one-line whole-document verification worth running on any PDF -before deciding whether image-handling code is needed at all. - -### 16. Chemical reaction arrows render as Private-Use-Area glyphs, not Unicode arrows -**What it looks like:** confirmed real example — physical page 1033 contains a -genuine chemical reaction equation (`Na2S2O3 + CN⁻ → SCN⁻ + Na2SO3`, part of -the cyanide-antidote/rhodanese mechanism description). The reaction arrow -extracts as a Private-Use-Area codepoint (``), not a standard Unicode -arrow (`→`) — the source PDF's font maps a custom symbol glyph (likely from a -symbol/wingdings-style embedded font) into a PUA slot, and raw text extraction -faithfully returns that codepoint rather than a human-readable arrow. -**Why it matters:** any pipeline that treats extracted text as directly -human-readable/citable will surface a mangled or invisible character where a -reaction arrow should be; a naive keyword/embedding step over raw text would -either silently drop it (if PUA codepoints get filtered as junk) or leave a -confusing tofu/box character in a chunk shown to a doctor or pharmacist. -**Check:** scan extracted text for codepoints in the Unicode Private Use Area -ranges (`U+E000–U+F8FF`) — cheap and generalizes to any custom-glyph symbol -substitution, not just arrows. -**Handling:** for now, flag any monograph/section containing a PUA codepoint -for manual review or map known PUA codepoints (e.g. this book's `` → -`→`) via an explicit substitution table; do not pass raw PUA codepoints -through to chunking/embedding untranslated. -**Generalizes:** yes — any PDF built from print-authoring software that uses a -symbol font for arrows/special glyphs (common in scientific/medical/chemistry -documents) can exhibit this; always check for PUA codepoints in extracted -text as a standing sanity check, not just assume standard Unicode symbols. -**Confirmed real chemical formula in the corpus, but rare:** a regex scan for -molecular-formula-shaped tokens (`[A-Z][a-z]?\d{1,3}` repeated) across the -monograph page range found 9 raw hits; manual inspection found most are -**false positives** (`H5N1` = flu strain name, `P2Y12` = a receptor name, not -molecular formulas) and only one confirmed genuine chemical formula/equation -(the Na2S2O3 case above) — real chemical notation exists in this corpus but -is genuinely rare, not a systemic pattern requiring a general chemistry -parser. - -### 17. Adult/child dosing-population splits are the norm, not an edge case -**What it looks like:** measured via a whole-monograph-range text scan for -"Người lớn"/"Trẻ em"/"Trẻ sơ sinh" (adult/child/newborn) — these terms appear -on **1121 of ~1400** monograph-range pages, i.e. the large majority of drug -monographs split dosing by patient population. -**Why it matters:** this is exactly the kind of structural content where a -segmentation/chunking bug that interleaves or merges adjacent subsections -(e.g. a table/list continuation bug, see items 5-6) would be a genuine -patient-safety risk, not just a data-quality nicety — mixing an adult dose -into a child-dose chunk (or vice versa) is a plausible, concrete failure -mode given how common this structure is. -**Handling:** treat "does this monograph's dosing section correctly keep -adult/child/newborn subsections un-interleaved" as a standing validation -check (not a rare-case afterthought), given the measured prevalence. -**Generalizes:** yes — any clinical/pharmacological reference document -organized with population-specific subsections has this same risk profile; -measure real prevalence before deciding how much validation effort a -structural risk deserves (same methodology lesson as item 12a). - -### 18. A monograph title can legitimately repeat — disambiguated by a bold, non-caps qualifier line -**What it looks like:** confirmed real example, found while smoke-testing -the real `segment/detector.py` against the full book: "SALBUTAMOL" is -detected as a monograph title **twice** (physical pages 1261 and 1263). -Rendering both pages to images and reading them directly (not inferred from -coordinates) confirmed these are two genuinely different, complete -monographs — "SALBUTAMOL (Dùng trong hô hấp)" (respiratory use) and -"SALBUTAMOL (Dùng trong sản khoa)" (obstetric/tocolytic use) — each with -its own full 18-section template. The qualifier ("(Dùng trong hô hấp)" / -"(Dùng trong sản khoa)") is a bold line immediately below the all-caps -title, but is **not itself all-caps** (mixed case inside the parens), so it -is correctly excluded from `detect_monograph_titles`'s all-caps candidate -filter — it must instead be captured as a *separate* signal and folded into -the monograph's disambiguating identity downstream. -**Why it matters:** an assembler that derives `drug_id` from the title text -alone (e.g. a simple slug of "SALBUTAMOL") will produce a real collision -between two legitimately different monographs — this is **not** the same -failure mode as the already-fixed GONADOTROPIN false-collision (that one -was a detector artifact from unmerged multi-line wrapping; this one is a -genuine same-name-different-monograph case that must be preserved, not -merged away). -**Handling (for Phase 1.3's assembler):** after detecting a monograph title, -check for an immediately-following bold, parenthesized, non-all-caps line -directly below it (same page, small y-gap) and include it in `drug_id` -generation when present, so "salbutamol_ho_hap" and "salbutamol_san_khoa" -remain distinct rather than colliding as "salbutamol" twice. The -`assembler.py` duplicate-drug_id check (outlier-catalog reasoning already -established: raise on a genuine duplicate rather than silently overwriting) -must be designed with this real case in mind, or it will incorrectly reject -a legitimate second "SALBUTAMOL" entry. -**Generalizes:** yes — any drug/entity reference work that documents the -same base substance under multiple distinct use-contexts (formulation, -indication, route) can have this exact pattern; never assume a title string -alone is a unique key without checking for a disambiguating qualifier line. - -### 19. Table column headers can be bold + all-caps + short — identical shape to a real title -**What it looks like:** confirmed real example, found via a whole-book -`assemble()` run raising a duplicate-drug_id error: "HSV" and "CMV" each -appear twice as bold, all-caps, short (3-char) spans on physical page 698 — -not drug names at all, but **column headers in a dosing-by-renal-function -table** inside the "Foscarnet natri" monograph ("Liều đối với HSV / HSV / -CMV / CMV"). Bold+all-caps+short is exactly the monograph-title signal -(item 10/12d), so this is a genuine detector ambiguity, not a coding bug. -**Why it matters:** unlike item 12d's part-divider titles (a small, -enumerable, fixed set of known strings), a table's column headers are -unbounded and content-dependent (any future table could use "HSV", "CMV", -or something else entirely as a header) — an exclusion list approach -doesn't generalize here the way it did for part-dividers. -**Handling:** require a **structural anchor** rather than a text exclusion -list: a real monograph title is always followed shortly by at least one -recognized section heading from the vocabulary (in practice, always "Tên -chung quốc tế" first) before the next title-shaped candidate. A -table-header false positive is not — the table's own cells are numbers/ -plain text, matching no vocabulary entry. Implemented as -`assembler._filter_false_positive_titles` (lookahead of 6 events, checked -against the same coalesced event stream already built for assembly — no -separate detection pass, no duplicated logic). -**Generalizes:** yes, more broadly than item 12d — any document where -section/entity boundaries are marked by a *shape* (bold+caps+short) that a -table, list, or figure caption could coincidentally also match should -verify a **structural follow-on anchor**, not just a shape match or a -denylist of known bad strings, since the space of possible false-shaped -content (table headers, figure labels, pull-quotes) is unbounded while the -space of "what a real boundary is followed by" is small and known. - -### 20. Section headings are not consistently bold across monographs — some combine label+value in one plain span -**What it looks like:** confirmed real example, found by investigating why -a whole-book `assemble()` run showed 48 monographs with zero ATC codes and -not stated-absent (far more than the ~13-14 the original spot-check -extrapolated). AMITRIPTYLIN's real "Mã ATC:" field is a **single, plain -(non-bold)** span containing the label AND value together: `"Mã ATC: -N06AA09."` — unlike Abacavir's equivalent, which is a bold `"Mã ATC: "` -label span followed by a separate plain `"J05AF06."` value span. Both -render visually similar but have completely different span/style -structure. Given the book's own foreword states it was "biên soạn bởi -nhiều tác giả" (written by many authors), this kind of per-author styling -inconsistency across ~700 individually-authored monographs is plausible -and, once checked, confirmed real — not a one-off. -**Why it matters:** a detector that requires `span.bold` to recognize a -section heading (reasonable-looking given every *title* is confirmed bold) -silently drops entire sections for a meaningful fraction of the corpus — -this directly caused undercounted ATC codes (and, structurally, would -equally affect any other section) for monographs using this looser style. -**Handling:** match section headings by **vocabulary text**, not by -boldness — the same "don't gate on a styling attribute, only content is -reliable" lesson as item 10 (font size), now applied to boldness. Also -handle the "label + value combined in one span" shape explicitly (a prefix -match: does the span start with a known label followed by ":", with the -remainder treated as the section's inline value) rather than assuming -label and value are always separate spans. -**Generalizes:** yes — any print-authored reference work assembled from -many individual authors/editors over a long production process should -expect inconsistent low-level styling of nominally-identical structural -elements; verify a structural signal (styling) against the *content* it's -supposed to correlate with, across a large real sample, before trusting it -as a universal discriminator — the same methodology lesson as item 10, -found again independently here. - -### 21. "All-caps" is not 100% reliable either — and a class-level monograph's own internal sub-headings can masquerade as new monographs -**What it looks like:** two distinct confirmed real findings from the same -investigation: -1. The class-level monograph "CÁC CHẤT ỨC CHẾ HMG-CoA REDUCTASE" embeds the - mixed-case abbreviation "CoA" (Coenzyme A) inside an otherwise all-caps - title. A strict `text.isupper()` check requires *zero* lowercase - letters, so this single embedded abbreviation caused the entire - monograph to be silently dropped from the corpus — found only by - directly checking whether this specific, previously-known (outlier item - 12a) class-level monograph was present in a real whole-book `assemble()` - run, and discovering it was not. -2. Within that same class-level monograph, individual statin names - ("SIMVASTATIN", "LOVASTATIN", "PRAVASTATIN", "FLUVASTATIN") appear as - their own bold+all-caps+short sub-headings, each introducing its own - "Liều lượng và cách dùng" sub-section — shape-identical to a real - monograph title, and (after fix 1 above made the loosened "any known - section" anchor check pass) briefly became a second false-positive - category alongside item 19's table headers, since these sub-headings - *are* followed by a recognized section, just never by "Tên chung quốc - tế" specifically (that section belongs only to the parent). -**Why it matters:** together these show that neither "all-caps" nor "loosen -the anchor to any section" is safe in isolation — the fix for one false -positive (item 19, HSV/CMV) reopened a different one (SIMVASTATIN) until -the anchor check was tightened back to the *specific* section the book's -own template guarantees is always first for a genuine top-level monograph. -**Handling:** `detector._is_mostly_upper` uses a **lowercase-letter ratio** -(≤10%), not an absolute count — an earlier absolute-count version (≤2 -lowercase letters) let a real regression through: "Mã ATC:" has only 1 -lowercase letter (a normal Vietnamese diacritic, 'ã') but that's 20% of its -5 letters, correctly rejected by the ratio while HMG-CoA's 1/27 ≈ 3.7% -correctly passes. `assembler._has_anchor_ahead` -requires specifically the "ten_chung_quoc_te" section key, not just any -recognized section, since that is the one invariant the book's documented -template actually guarantees is unique to real top-level monographs. -**Generalizes:** yes — (1) don't assume a styling/casing convention holds -with zero exceptions across an entire corpus, even one confirmed exception -matters at whole-corpus scale; (2) when a document has nested substructure -that mimics top-level structure (a class monograph containing per-item -sub-entries), the anchor used to confirm a real boundary must be the most -*specific* invariant available, not just "some known follow-on content" — -a looser check that fixes one false positive can silently reopen another. - -### 22. Running-header boilerplate was never actually stripped, despite item 13's warning — measured whole-corpus at 98.4% of monographs affected -**What it looks like:** the running header at the top of every physical page -("DTQGVN 2" + printed page number + the current monograph's name, e.g. -physical page 1008's "DTQGVN 2" / "1009" / "Morphin sulfat", tagged -`column="full_width"` by `extract/spans.py`) matches no section heading and -isn't a real all-caps title, so it fell through every classification branch -in `assembler._classify` into plain body text — splicing itself into the -*middle* of whatever section is open when a physical page turns. Real -example, MORPHIN SULFAT's `liều lượng và cách dùng`: `"...Nếu\nDTQGVN 2\n -1009\nMorphin sulfat\nuống viên thuốc..."` — the header text lands inside a -real dosing sentence. -**Why it matters:** item 13 (above) already *warned* "strip the fixed -boilerplate before parsing content" back when the extraction layer was -first built, but that step was never actually implemented in `assembler.py` -— the warning existed in the catalog without a corresponding code path or -test enforcing it, and nothing caught the gap until a whole-corpus -measurement was actually run. Measured: **1,374 of 11,409 sections (12.0%) -contained a literal "DTQGVN" string mid-text; 671 of 682 monographs (98.4%) -had at least one affected section** — this is not a rare edge case, it's -the default outcome for any section whose text happens to cross a physical -page boundary (i.e. most sections longer than about half a page). Left -unfixed, boilerplate gets baked into chunks and embeddings and can surface -mid-sentence in a citation shown to a doctor/pharmacist. -**Handling:** `assembler._is_page_boilerplate` drops any span with -`column == "full_width"` and `y0 < HEADER_BAND_Y` (the same header-band -threshold `page_map.py` already uses to read the folio) before it reaches -any other classification branch. Whole-corpus re-measurement after the fix: -0 of 11,409 sections contain "DTQGVN". Regression test uses the exact real -MORPHIN SULFAT span shape. -**Generalizes:** a documented risk in this catalog is not the same as a -verified-fixed risk — "we know this could happen" needs a whole-corpus -measurement (not just a warning paragraph) before it can be crossed off, -and ideally a regression test that would fail if the fix were ever reverted. - -### 23. PyMuPDF's raw block order doesn't reliably sequence left-column-before-right-column — confirmed wrong on 12 of 1398 pages -**What it looks like:** `extract/spans.py` originally trusted PyMuPDF's own -block iteration order to already emit left-column content before -right-column content, validated only against one example page during ADR -0003. On physical page 1100 (the OXYBUTYNIN/OXYMETAZOLIN monograph -boundary) and 11 other pages, PyMuPDF's raw block order emits the *right* -column first. Since `assembler.assemble` appends section content to -whichever monograph is currently open, this silently attributed -OXYMETAZOLIN's right-column sections (Chống chỉ định, Thận trọng, Thời kỳ -mang thai, Thời kỳ cho con bú, ADR, Hướng dẫn xử trí ADR, Liều lượng và -cách dùng) to the still-open OXYBUTYNIN monograph — overwriting -OXYBUTYNIN's real sections and leaving OXYMETAZOLIN missing all 7. -**Why it matters:** medically relevant — wrong contraindication/ADR content -silently attached to the wrong drug. Found via a whole-document -(1668-page) character-similarity diff against an independent parser -(`opendataloader-pdf`), not from a sample; confirmed by rendering the page -to an image and reading it directly, then confirmed again in the actual -`assemble()` output. -**Handling:** `extract.spans._sort_blocks_reading_order` explicitly sorts -each page's blocks by (full_width header band first, then left column, -then right column) and then by y-position, instead of trusting raw PyMuPDF -order. Whole-range (99-1496) re-scan after the fix: 0 pages with the -reversed-order signature (was 12). Directly verified OXYBUTYNIN's and -OXYMETAZOLIN's `assemble()`-produced sections are now distinct and -drug-appropriate. -**Generalizes:** don't trust an upstream library's element ordering just -because it happened to be correct on the one page checked during initial -validation — for a whole-corpus pipeline, explicitly sort by the actual -signal you care about (here: visual column position) rather than an -implicit "the library probably does this right" assumption. - -### 24. Some text exists only as vector outlines — no text extractor can read it, and single dropped glyphs corrupt otherwise-clean sentences -**What it looks like:** physical page 714 prints 17 full lines of ordinary -GATIFLOXACIN prose that `page.get_text()` does not return, `page.search_for()` -cannot find, and neither `pdfplumber` nor `opendataloader-pdf` returns either. -`page.get_drawings()` shows why: each line is a filled path of 1,126-1,831 -items, shaped exactly like one line of type and filled with the body-text -colour. The same defect occurs at glyph granularity (39-45 path items), and -that form is far more dangerous — a single Vietnamese diacritic character -drops out of a line that otherwise extracts perfectly: `Độ ổn định` extracts -as `Độ n định`, `≥ 1 tuổi` as `≥ 1 tu i`, `tại chỗ` as `tại ch `. The result -reads as ordinary text, so no structural check, no count and no cross-tool -comparison notices it. -**Why it matters:** this is silent loss of clinical prose in a drug -formulary, and it is invisible to every check that asks a text layer a -question. It survived a whole-document span-coverage ledger reporting -`unassigned = 0`, because the spans that existed were all routed correctly — -the missing content was never a span at all. -**Check:** render the page, white out every extracted span's bbox, and look -at the ink that survives (`ingestion/validation/residual_ink.py`, ~0.06 -s/page). Confirm with `page.get_drawings()`: a filled path with ≥30 items -whose box is 3-20pt tall is type, not decoration (real decoration on this -book carries 1-2 items). -**Handling:** `ingestion/extract/outlined_text.py` detects the runs; -recovery cannot be automatic because the paths carry no character codes, so -each run was rendered and transcribed by reading it, into -`ingestion/data/verified/outlined_text_transcriptions.json` with page, bbox, -and the extracted line it belongs to. Whole-document scope: **51 runs on 5 -pages** (714 ×31, 736 ×16, 1373, 1444, 1445 ×2), 1,116 characters. -**Generalizes:** yes — any PDF produced by a layout tool that converts -selected text to outlines (common when a font cannot be embedded) has this. -Never treat "the text layer returned something for this page" as evidence -the page was fully extracted; compare against the rendered pixels. - -### 25. A fraction can be printed with no fraction bar at all, so no geometric detector can find it -**What it looks like:** ADENOSIN (physical page 147) prints its infusion-rate -formula as three plain lines — `Tốc độ truyền dịch (ml/phút) = 0,140 -(mg/kg/phút) × trọng lượng cơ thể (kg)` / `Nồng độ adenosin (3 mg/ml).` — -with **no rule drawn between numerator and denominator**, confirmed by -rendering the region and reading it. Extracted linearly it reads as a -multiplication chain, i.e. the division silently disappears. -**Why it matters:** it defeats the detector that catches every other 2D -formula in this book. The fraction-bar signal (item 8, and -`residual_ink.py`'s `fraction_bar_candidate`) finds ink; there is no ink to -find here. It was caught only because a prose-leak gate matched its text. -**Check:** there is no cheap automatic check. Treat any line ending in a -unit-bearing quantity immediately followed by a line that is itself a -unit-bearing quantity as a division candidate for human review. -**Handling:** quarantined via the verified region list with -`source_prints_no_bar: true`. The count of bar-less formulas in this book is -**unmeasured** — recorded as `recall_limit` in -`ingestion/data/verified/formula_regions_2d.json` so the bar scan is never -mistaken for complete formula coverage. -**Generalizes:** yes — measured precision of the fraction-bar rule on this -book is **16/23 = 69.6%**, and its recall is unknown. A geometric heuristic -finds candidates; it never proves absence. - -### 26. Exact section vocabulary can occur as wrapped prose or inside tables; context must precede label matching -**What it looks like:** several unrelated defects shared one cause. A wrapped -body sentence can put `chống chỉ định.` alone on the next visual line -(NADROPARIN, physical page 1016); a dosing-table cell can literally be named -`Chỉ định` (WARFARIN p1485 and IOBITRIDOL p826); and a verified fraction band -widened to capture its numerator can geometrically overlap prose in the other -column (NETILMICIN p1042). Exact vocabulary matching alone classified these as -structure or quarantined content. - -**Why it matters:** the output remains grammatical while moving or deleting a -clinically decisive phrase, assigning a dosing table to indications, or hiding -a cross-reference. Aggregate “all spans assigned” and section-level provenance -gates all passed before these defects were found. - -**Handling:** classify out-of-scope spans and known table regions before title/ -section matching; treat a non-bold exact label as prose when it is the adjacent -line of an unterminated span in the same PDF block; require a formula region's -column to agree with the source span's column; and validate source-span IDs on -every individual part. Confirmed aliases (`Tên chung quốc tế và mã ATC`, `Dạng -bào chế và hàm lượng`, and the tetanus-toxoid dosing heading) are recorded in -the open vocabulary. - -**Whole-corpus result:** 684 monographs (was 683), maximum monograph range 7 -pages (was the false 164-page ZOLPIDEM range), 11,974 sections, 151 quarantined -blocks, 15,066 chunks, 0 unassigned spans, and every readiness gate passing. - -**Generalizes:** vocabulary is evidence, not sufficient context. Apply known -geometric scope (page, table, column, visual-line continuity) before interpreting -a label-shaped string as document structure. - -### 27. One physical table can be non-contiguous in PDF block order -**What it looks like:** a table is contiguous on the rendered page, but the PDF -content stream interleaves a visually later section heading between its cells. -This split CAPECITABIN p308 and IMATINIB p795 into multiple blocks with the same -region ID and conflicting section owners. CAPECITABIN p309 adds a second case: -two explicitly captioned dose-adjustment tables are printed after the ordinary -`Tên thương mại` field without repeating the dosage heading. - -**Why it matters:** sorting or classifying one extracted span at a time makes a -single physical object acquire several meanings. The flattened text remains -plausible, so ordinary text and coverage gates do not expose the defect. - -**Handling:** collect all spans belonging to a verified region before semantic -classification and emit the region atomically at its first occurrence. A narrow -caption rule maps only `Bảng N. Điều chỉnh liều ...` appendices to -`lieu_luong_va_cach_dung`; generic occurrences of the word “liều” are not used. -A readiness gate now requires unique physical-region IDs. - -**Verification:** all **151/151 unique regions** were rendered and read against -the PDF. The regenerated corpus has 151 blocks, 151 unique IDs, and zero -duplicate-ID gate failures; CAPECITABIN p309 tables are both owned by dosage. - -**Generalizes:** physical-region identity must outrank text-stream adjacency for -tables, formulas, figures, and other layout objects. - -### 28. A bar-less formula needs an asymmetric band, but geometry cannot prove its operator -**What it looks like:** ADENOSIN p147 prints a wrapped numerator followed by -`Nồng độ adenosin (3 mg/ml).` with no horizontal fraction rule. The generic -symmetric formula band captured the numerator only, making a plausible but -incomplete source crop. - -**Why it matters:** the missing denominator changes the calculation. Visual -review of all reconstructed sandbox crops found the defect even though ordinary -readiness and block-count gates passed. - -**Handling:** verified bar-less regions use a 31pt lower margin from the -synthetic anchor. On this page the denominator ends about 29pt below the anchor; -the following `Ví dụ:` begins immediately after the new boundary. A regression -requires the denominator boundary and excludes that prose. The reconstructed -record still sets `requires_human_operator_confirmation`: layout supplies no -bar from which multiplication versus division can be proven. - -**Generalizes:** expand a verified crop to preserve all visible operands, but -never invent a mathematical operator that the source geometry does not encode. - -## Not yet investigated (flagged for future work, not silently ignored) - -- **Footnote-style superscript reference markers** (seen as `a, b, c, d` in - one table) — not yet checked for whether the footnote text stays - correctly associated with its marker/row during extraction. -- **How many bar-less formulas exist** (item 25) — one confirmed, total - unmeasured; no geometric signal can bound it. -- **Production 2D grid reconstruction** (item 7) — the 100-page sandbox now - reconstructs grids and logical cross-page tables, but merged-cell semantics - and whole-book recall are not yet production gates. -- **Exact shortest monograph name+page** — a quick unmerged crude scan (no - multi-line title merge) gave a different longest-monograph ranking than - the already-documented authoritative one (item 12e: "AMOXICILIN VÀ KALI - CLAVULANAT", 45,623 chars), meaning the crude scan's numbers are not - reliable enough to name an exact shortest monograph — deferred to the real - Phase 1.2 detector (with proper multi-line merge and back-index-validated - boundaries), which will produce a trustworthy number as a side effect of - its own validation run, rather than trusting today's quick, differently- - scoped script.