7.5 KiB
ADR 0004: Chunking strategy for drug monographs — validated against real per-section measurements
Status
Accepted for the monograph range (printed pp. 99-1496) only. General chapters (pp. 37-98) and appendices (pp. 1497-1528) are explicitly out of scope — see Consequences.
Context
docs/architecture.md's original "Chunking" paragraph specified (drug, section) as the chunk unit, a ~500-800 token budget, and a 400-token/
50-overlap sliding window for oversized sections. Those numbers were written
before segmentation existed — a plausible guess, never checked against real
per-section text length.
Phase 1 (extract → segment → validate) is now real, tested code producing
682 real monographs from the full 1668-page source PDF. This session ran
python -m ingestion.cli run for real and measured actual per-section
length across the whole corpus with a temporary investigation script
(ingestion/scratch/chunking_stats_survey.py, deleted after this ADR
captured its findings, per this project's investigation-script rule) —
something that had never been measured before this ADR.
What was actually measured (whole corpus, 682 monographs)
- Sections per monograph: min 11, median 17, max 19 (of ~18-19 known
section keys in
segment/vocab.py's open taxonomy). - Whole-monograph length: median 11,480 chars, p90 19,068 chars, max 38,786 chars.
- Per-section length, converted to a chars/4 token estimate — an
estimate, not a real tokenizer count:
- Most of the ~18 section types sit comfortably under 800 estimated
tokens even at their p90 (e.g.
chi_dinhp90≈268 tok,dang_thuoc_va_ ham_luongp90≈115 tok,tac_dung_khong_mong_muonp90≈481 tok). - Two sections routinely exceed 800 tokens:
duoc_ly_va_co_che_tac_dung(242 of 678 monographs that have this section, 35.7%, max ≈3542 tok) andlieu_luong_va_cach_dung(200 of 675, 29.6%, max ≈3631 tok). - A smaller tail also exceeds it:
than_trong(25/680, 3.7%),tuong_tac_thuoc(22/642, 3.4%). - This means: the original 800-token ceiling is directionally correct (it clears ~16 of 18 section types at their p90 with room to spare), but "sub-chunk in that case" is not a rare hedge as originally implied — it is the routine path for roughly a third of all monographs, on two specific, named, high-clinical-importance sections (mechanism of action and dosing).
- Most of the ~18 section types sit comfortably under 800 estimated
tokens even at their p90 (e.g.
A separate, blocking bug was found while gathering this data, not fixed
by this ADR (out of scope — belongs to extract/segment, owned by a
parallel session at the time of writing): running header/footer
boilerplate ("DTQGVN 2" + page number + repeated drug name, tagged
column="full_width" in extract/spans.py) is never filtered out of
section body text before it reaches SectionSpan.text. Measured:
1,374 of 11,409 sections (12.0%) contain a literal "DTQGVN" string
mid-text; 671 of 682 monographs (98.4%) have at least one affected section
(e.g. MORPHIN SULFAT's lieu_luong_va_cach_dung: "...Nếu\nDTQGVN 2\n1009\n Morphin sulfat\nuống viên thuốc..."). This is docs/pdf-parsing-outlier- catalog.md item 13's known risk, measured whole-corpus for the first time
here. Chunking must not run against real data until this is fixed —
otherwise boilerplate is baked into embeddings and can surface mid-sentence
in a chunk shown to a doctor or pharmacist.
Decision
- Chunk unit stays
(drug_id, section_key)— matchessegment/models.py's existingMonograph.sections: Dict[str, SectionSpan], matches how a doctor/pharmacist would query ("what does it say about liều dùng"), and lets a citation point at one clinical section rather than a whole 2,000-19,000-char monograph. - Token budget: keep the 800-token ceiling (chars/4 estimate) as the split trigger. Below it, a section is one chunk, verbatim. This is now a validated choice, not a guess.
- Sub-chunking only applies to the long-tail sections above (~30-36%
of monographs for the two named sections, a few percent for the rest).
Method: sentence-boundary-aware sliding window, replacing the
originally-guessed fixed-character window. Target ~600-700 tokens per
sub-chunk (headroom under the 800 ceiling), ~1 sentence / 50-80 token
overlap between adjacent sub-chunks. Split only at a sentence boundary
(
./;/:followed by whitespace + capital letter), explicitly not treating a Vietnamese decimal comma (e.g. "0,425") as a boundary. - Why sentence-aware, not line- or character-based:
assembler.pyjoinsbody_linesone line per PyMuPDF span, i.e. one PDF visual line-wrap point — not a semantic paragraph or sentence boundary. A blind character/line window can split a sentence mid-way. This is a real, measured risk here, not theoretical: outlier-catalog item 17 found adult/child dosing splits ("Người lớn"/"Trẻ em") appear on 1,121 of ~1,400 monograph-range pages — a chunk boundary landing inside one of those sentences would be a patient-safety-relevant defect, not a cosmetic one. - Chunk metadata / provenance (extends the existing
drug_name, section_type, source_page_range, chunk_idlist indocs/architecture.md— per CLAUDE.md's provenance rule):chunk_id({drug_id}__{section_key}__{part_index}),drug_id,drug_name,section_key,section_display_name,atc_codes(inherited from the monograph — enables ATC-class-filtered retrieval),source_page_range(monograph-level, see Consequences),part_index/part_count(0/1for un-split sections, keeps the schema uniform across all chunks).
Consequences
- Scope: this decision covers the monograph range only. General
chapters and appendices contain real tables and 2D stacked-fraction
formulas (
docs/document-profile.md, investigation in progress as of this ADR) that need their own structural survey before any chunking rule can be designed for them — do not extend this ADR's rules to those ranges without a fresh investigation. - Hard prerequisite: the boilerplate-leakage bug described above must
be fixed in
extract/segmentbefore this chunking design is run against real data for ingestion. This ADR does not fix it. - Known gap — sub-compound tagging inside class-level monographs: 25.5%
of the corpus has more than one ATC code per monograph (outlier item
12a), e.g. "VITAMIN D VÀ CÁC THUỐC TƯƠNG TỰ" documents dosing for 7
different analogues inside one
lieu_luong_va_cach_dungsection. No reliable structural signal was found in sampled text to split a section by sub-compound — a chunk from this section is tagged with the class name only, not the specific analogue a query might target. Deferred to golden-dataset-driven eval rather than guessed at now. - Known gap — sub-chunk page precision:
source_page_rangeis monograph-level, not sub-chunk-exact. A sub-chunk from late in a multi-page section inherits the whole monograph's page range rather than its own precise page, because per-line page tracking doesn't currently exist inSectionSpan/Heading. The monograph + section-heading page is still real, checkable provenance, but this is a known precision gap, not full sub-chunk traceability. Flagged as a future improvement. - Not yet built: the Vietnamese sentence-boundary splitter itself
(abbreviation handling, decimal-comma handling, ATC-code-period handling)
is specified here as a rule, not implemented or unit-tested. Building and
testing it is a separate, later task (
ingestion/ingestion/chunk/, which does not exist yet).