Files
duocthu/docs/progress-log.md
T

14 KiB

Progress Log

Chronological record of work done on this project, newest entry on top. The goal is continuity across sessions: if a work session ends unexpectedly (context/token limit, interruption), whoever picks this up next — human or Claude — should be able to read the latest entry and know exactly what's done and what's next, without having to reconstruct it from git history.

Convention: add a new entry at the top before ending a session whenever meaningful progress was made, and proactively the moment it looks like the session might run out of context/tokens mid-task — don't wait until the very end if that risk is showing.


2026-07-30 — Whole-corpus structural survey (not just anecdotes)

Done (direct pushback: "I feel like you're minimizing how complex this PDF really is — go find another 10-30 outliers, not just Vitamin D"):

  • Built a real per-monograph structural survey across all 680 detected monographs (not 2 anecdotes) — computed ATC-code count, known-section count, and character length for every one.
  • Multi-ATC monographs are NOT rare: 173/680 (25.4%) have more than one ATC code — INSULIN has 20, BETAMETHASON and DEXAMETHASON 11 each, PREDNISOLON 10, HYDROCORTISON 9. The earlier "found 2 examples" framing badly understated this. Even 25.4% is a floor (see next point).
  • Investigated the 22 apparent "zero ATC" monographs (spot-checked 14): found two distinct real causes of false negatives — stray internal whitespace splitting an ATC code ("J04A C01" instead of "J04AC01") and digit/letter confusion ("NO3AX12" instead of "N03AX12") — 9 of 14 resolved as real ATC codes hidden by extraction noise (one of them, TRIAMCINOLON, turned out to have 5 ATC codes, meaning the true multi-ATC percentage is higher than 25.4%). The remaining ~5 genuinely say "Mã ATC: Chưa có." (not yet assigned) — a valid data state, not an error.
  • Found and confirmed a false-positive monograph boundary: the part-divider title "CÁC CHUYÊN LUẬN THUỐC" (Part 2's own section title, not a drug) was detected as if it were a monograph.
  • Measured real structural variance: monograph length ranges 2,331-45,623 characters (~20x spread), detected section count ranges 8-20.
  • All findings added to docs/pdf-parsing-outlier-catalog.md (items 12a revised with real numbers, 12c, 12d, 12e — new).
  • Verified one of my own debugging steps was itself wrong (read raw page text from the top instead of the correctly-bounded monograph segment, which briefly looked like a segmentation bug before being traced back to a debugging mistake, not a real defect) — corrected before reporting.

Not done yet / next up:

  • Full-corpus re-count with the relaxed ATC regex (whitespace-tolerant, O/0-aware) not yet run — only 14/22 zero-ATC cases spot-checked, and the 173/680 multi-ATC count still uses the strict (undercounting) regex.
  • Phase 1 real implementation still pending overall (see earlier entries).

2026-07-30 — Confirmed class-level monographs and a real source typo

Done (direct follow-up: "have you checked drug-class entries like Vitamin D, or actual spelling/font-size errors?"):

  • Found and confirmed a second real example of a class-level monograph covering multiple ATC codes/substances: "VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ" (7 ATC codes, one per specific vitamin D analogue) — same pattern as the earlier GONADOTROPIN finding, confirming this is recurring, not a one-off.
  • Found and confirmed a genuine spelling/capitalization typo in the source PDF itself: the running header on this monograph's continuation pages reads "Vitamin d..." (lowercase d) vs the correct ALL-CAPS heading "VITAMIN D...". Verified via font/bbox inspection that this is a real source-text inconsistency, not an extraction artifact. The detection heuristic still worked correctly here (the typo'd header isn't all-caps so it's correctly rejected), but this was incidental, not a designed defense against typos.
  • Added both findings to docs/pdf-parsing-outlier-catalog.md (items 12a, 12b), with the general lesson: rely on multiple independent structural signals, not any single text match, since real source typos do occur.
  • Added CLAUDE.md with a standing rule: never fabricate or bluff a claim (number, test result, capability estimate) — verify before stating, explicitly flag estimates as estimates. Grounded in concrete incidents from this investigation (the size-threshold bug, the scope-gap bug).

Not done yet / next up:

  • No systematic scan yet for other class-level (multi-ATC) monographs beyond the two found incidentally — Phase 1's data model should assume ATC code is a list per monograph regardless, rather than trying to enumerate every class-level entry in advance.
  • Phase 1 real implementation still pending overall (see earlier entries).

2026-07-30 — Comprehensive PDF outlier catalog (tables, formulas, columns)

Done (in response to direct follow-up questions about table/formula handling and full-book coverage):

  • Found and confirmed a table split across a page break loses its header on the continuation page — real example: "Bảng 4" (ARV rash management table) ends with an orphaned, header-less data row on the next page when extracted with pdfplumber.
  • Found the same header-loss risk also happens across a column boundary within a single page (no page break needed) — real example: "Bảng 6".
  • Found and confirmed 2D grid/nomogram tables are not linearly recoverable — the body-surface-area lookup table (appendix) extracts as scrambled bare numbers with no row/column association.
  • Found two different formula-rendering outcomes: a simple inline- exponent formula (Du Bois BSA) extracts cleanly as text; a stacked- fraction formula (Cockcroft-Gault) extracts as disordered fragments — confirmed the determining factor is 1D vs 2D visual layout, not "formulas are always broken."
  • Found and confirmed a full-width table that breaks out of the normal two-column page grid (bbox spans nearly the full page width).
  • Checked whether front-matter "committee list" pages are genuinely multi-column (the user suspected 3 columns) — confirmed via bbox inspection they are not true structural columns, just single wide text blocks with internal whitespace padding between names.
  • Consolidated all outlier findings from this investigation (this entry and the previous one) into a single, reusable, generalized reference: docs/pdf-parsing-outlier-catalog.md — written so it can guide parsing of other similarly-structured PDFs, not just this book.

Not done yet / next up:

  • No automatic detector exists yet for (a) 2D-formula regions, or (b) 2D grid-table reconstruction — both flagged as open items in the catalog, not silently skipped.
  • Table-continuation re-attachment (page-break and column-break cases) has no implementation yet — needed before Phase 1 can trust any multi-row table content.
  • Phase 1 real implementation still pending overall (see previous entry).

2026-07-30 — PDF parsing strategy validated empirically (pre-Phase-1)

Done:

  • Investigated the real PDF structure before writing any ingestion code (previous scaffold's assumptions about doc.get_toc() turned out wrong).
  • Confirmed: 1668 pages, no bookmark/outline (0 TOC entries), tagged-PDF structure tree exists but is too shallow to use (~29 elements only).
  • Cross-tested 3 extraction tools on real sample pages: PyMuPDF (correct reading order — kept as primary), pdfplumber (scrambled reading order on this layout — demoted to table-extraction-only use), opendataloader-pdf (correct reading order, useful independent font-metadata cross-check, but inconsistent heading classification — not trusted as sole signal). Docling install hit a numpy/pyarrow ABI conflict in the global Python env; tested in an isolated .venv_docling_test/ (gitignored) instead of risking the global environment — see whether that resolved before relying on it.
  • Found the real structural ground truth: every section/monograph heading is a bold font span in the PDF (confirmed at the PyMuPDF span level AND independently by opendataloader's own font metadata — two tools agreeing). Font size is not reliable (10.0pt and 9.5pt both occur for genuine monograph titles) — an early size-based threshold silently dropped ~15% of real monographs; caught and fixed via whole-document validation, not spot-checking.
  • Found the real ground truth for validation: the back-of-book "Mục lục tra cứu" (page ~1528 onward) has exact page numbers per drug — much stronger than the front-matter drug list (which has no page numbers). Also found the book's own contents page states individual monographs run printed pages 99-1496 exactly.
  • Ran automated whole-document (1668-page, ~20-50s per run) validation against that page-verified ground truth: 91.7% recall (665/725), with the remaining gap traced to one concrete, fixable cause (multi-line wrapped ALL-CAPS titles not yet merged across lines) rather than a flaw in the bold-span signal itself.
  • Documented the full methodology and results in docs/adr/0003-pdf-parsing-strategy.md and updated the ingestion section of docs/architecture.md to match reality (removed the incorrect TOC-preference assumption).

Also validated (in response to direct user questions about correctness):

  • No real duplicate drug monographs found across the full 1405-page monograph range. The one apparent collision ("GONADOTROPIN" at 2 pages) is a detector artifact from the known multi-line-title bug (a different monograph's wrapped title fragment collided with it), not real content duplication.
  • Confirmed the PDF is genuinely two-column (bounding-box verified: left column x≈44-299, right column x≈308-562). PyMuPDF's reading order across columns is correct (already implied by earlier validation).
  • Found and precisely characterized one real data-corruption defect: a single text run on physical page 1373 has reversed (right-to-left) glyph order, producing scrambled text — confirmed by reversing the string, which recovers the correct Vietnamese sentence. A full scan of all 1405 monograph pages (grouping fragments into visual rows, checking for descending x-order) found this exact 1 occurrence and no others — rare, isolated, but real, and now has a cheap (~16s) automated detector.
  • Full details, methodology, and exact numbers added to docs/adr/0003-pdf-parsing-strategy.md under "Follow-up validation."
  • Caught a real scope gap: the glyph-reversal scan above was initially run on the monograph range only (1405 of 1668 pages), leaving ~260 pages (front matter, appendices, back index) unchecked. Re-ran across the full 1668 pages: still exactly 1 defect (same page, 1373) — confirmed isolated, not hiding elsewhere. Also found 6 near-empty pages (3, 37, 99, 1495, 1497, 1666), all of which land exactly on major section-transition boundaries — intentional print blank pages, not lost content.

Not done yet / next up:

  • Resolve/confirm docling status in the isolated venv (numpy/pyarrow conflict was fixed by using a separate venv; install completed — actual parsing comparison against the sample pages still pending).
  • Phase 1 real implementation: build ingestion/ for real using the validated bold-span detector (not the exploratory scratch scripts) as one continuous cross-page stream (not per-page silos), fix the multi-line heading-merge gap, add the glyph-order sanity check as a mandatory pre-ingestion pass, re-run the validation script to confirm improved recall, then proceed to chunking + embedding + Qdrant upsert.
  • Decide and implement chunking strategy for the non-monograph parts of the book (general chapters pages 37-98, appendices 1497-1528) — needed so the full book (page 0 to last) ends up captured in the RAG corpus in some appropriate form, per the user's explicit requirement that no content be silently dropped.
  • Clean up exploratory scratch_* files from the repo root as they accumulate during investigation (routinely deleted after findings are persisted to docs — not left in git history).

2026-07-30 — Initial monorepo scaffold

Done:

  • Designed the microservices architecture (see docs/architecture.md): Python/FastAPI ai-service for RAG, NestJS for api-gateway/auth-service/ user-service/chat-service, Next.js web, Qdrant for vectors, Postgres for relational data, Redis reserved for caching/queues.
  • Scaffolded the full monorepo directory tree (apps/, packages/, ingestion/, infra/, docs/) with baseline config (package.json/ pyproject.toml stubs, pnpm workspace, docker-compose topology stub).
  • Moved duoc-thu-quoc-gia-viet-nam-2018.pdf into ingestion/data/raw/.
  • Decided vector DB: Qdrant over pgvector (docs/adr/0001-vector-db-qdrant.md).
  • Decided deployment: GitOps via the team's existing ArgoCD instance, not a custom push-based CD pipeline (docs/adr/0002-argocd-gitops.md, infra/argocd/). CI's job is build/test/push image + bump the Helm values image tag; ArgoCD does the actual sync.
  • git init + initial commit (this scaffold).
  • Created a private GitHub repo (BaoVu2k4/vsf-duocthu, default branch master) and pushed the initial commit; fixed targetRevision in the ArgoCD Application manifests to master to match.

Not done yet / next up (Phase 1 of the build roadmap in docs/architecture.md):

  • No business logic exists yet anywhere — this was scaffold only.
  • Phase 1: build the ingestion/ pipeline for real (PDF extraction via PyMuPDF, monograph/section segmentation, section-aware chunking, OpenAI embeddings, Qdrant upsert) and validate retrieval quality via the ingestion/notebooks/ QA step.
  • Still pending/TBD: which cloud provider (AWS/GCP/Azure) for Terraform (infra/terraform/README.md), and the team's ArgoCD instance's actual cluster/server + project details (infra/argocd/README.md TODOs).