Files
duocthu/docs/06-document-model-and-chunking.md
T

7.8 KiB

06 — Document model and chunking

Entity model

erDiagram
    MONOGRAPH ||--o{ SECTIONSPAN : sections
    MONOGRAPH ||--o{ TABLEBLOCK : tables
    MONOGRAPH ||--o{ SECTIONPART : preamble
    SECTIONSPAN ||--|| HEADING : heading
    SECTIONSPAN ||--o{ SECTIONPART : parts
    SECTIONSPAN ||--o{ CHUNK : "prose chunks"
    TABLEBLOCK ||--|| CHUNK : "1 block_descriptor chunk"
    CHUNK ||--o{ CHUNKATTACHMENT : attachments
    CHUNK ||--|| VECTORPOINT : "uuid5(chunk_id)"

    MONOGRAPH {
        string drug_id PK
        string drug_name
        int_list source_page_range
        string_list atc_codes
        bool atc_stated_absent
    }
    SECTIONSPAN {
        string key
        string display_name
        string text
    }
    SECTIONPART {
        string kind "prose|table"
        string text
        int physical_page
        float_list bbox
        string_list source_span_ids
        bool quarantined
    }
    TABLEBLOCK {
        string table_id PK
        string shape
        int physical_page
        float_list bbox
        string section_key
        bool quarantined
    }
    CHUNK {
        string chunk_id PK
        string drug_id FK
        string section_key
        string text
        string source_text
        int_list source_page_range
        int_list printed_page_range
        int part_index
        int part_count
        string chunk_kind
        bool has_quarantined_content
        int schema_version
    }
    CHUNKATTACHMENT {
        string block_id FK
        string kind "table|formula"
        int physical_page
        float_list bbox
        int printed_page
        bool quarantined
    }

Source: ingestion/segment/models.py, ingestion/chunk/models.py, ingestion/load/models.py.

The 19 section keys

Book order, as defined in apps/ai-service/rag/sections.py::SECTION_ORDER (18 entries — ten_thuong_mai exists in the vocabulary but not in the ordering tuple) and rag/understanding.py::SECTION_KEYS (all 19):

ten_chung_quoc_te, ten_thuong_mai, ma_atc, loai_thuoc, dang_thuoc_va_ham_luong, duoc_ly_va_co_che_tac_dung, chi_dinh, chong_chi_dinh, than_trong, thoi_ky_mang_thai, thoi_ky_cho_con_bu, tac_dung_khong_mong_muon, huong_dan_xu_tri_adr, lieu_luong_va_cach_dung, tuong_tac_thuoc, qua_lieu_va_xu_tri, do_on_dinh_va_bao_quan, tuong_ky, thong_tin_quy_che.

All 19 appear in the loaded corpus. Chunk counts per section (counted this session over chunks.jsonl):

Section Chunks
duoc_ly_va_co_che_tac_dung 1,896
lieu_luong_va_cach_dung 1,873
than_trong 927
tac_dung_khong_mong_muon 857
tuong_tac_thuoc 810
chi_dinh 710
dang_thuoc_va_ham_luong 691
ten_chung_quoc_te 684

The two largest sections being pharmacology and dosage is exactly why rag/sections.py exists — see 09-retrieval-pipeline.md.

Chunking strategy (ADR 0004)

Unit: (drug_id, section_key). A section under the token ceiling becomes one chunk, verbatim. Only the long tail is sub-chunked.

CEILING_TOKENS = 800    # above this, sub-chunk
TARGET_TOKENS  = 650    # packing target
OVERLAP_TOKENS = 65     # sliding-window overlap

Token counting uses tiktoken cl100k_base when available, and an estimate otherwise — cli chunk prints which one it used.

Sub-chunking

  1. Atomise (_atoms): split into sentences (chunk/sentences.py, which treats : as a boundary). A "sentence" longer than TARGET_TOKENS that contains commas is split on commas — needed because a drug-interaction list is one grammatical sentence hundreds of names long: VORICONAZOL's tương tác thuốc produced 981- and 888-token parts, and a truncated interaction list reads as "this drug is not listed", a false negative in the dangerous direction.
  2. Pack (_pack_parts): greedily fill to TARGET_TOKENS, then overlap the tail by up to OVERLAP_TOKENS.

The clinical-context rules inside the packer

These are the non-obvious part, and each exists for a measured defect:

  • Never end a part on a label. "Người lớn: 500 mg mỗi 8 giờ." splits after the colon; flushing there would leave a chunk ending "Người lớn:" with the dose in the next one. Measured before the rule: 38 such chunks. A dose separated from the population it applies to is a patient-safety defect.
  • Carry the governing label forward. contexts / scope_contexts / context_chain() track the active label and its parent scope per atom, so a population label that fell out of both the 650-token buffer and the 65-token overlap several parts ago is repeated at the seam.
  • Split a trailing label off compound atoms. _split_trailing_label handles "7,5 mg … .\nBước 5:" so the dose at the atom's start does not lose Bước 4.
  • Repeated labels are marked as context, not source. Chunk.text may contain a prepended label; Chunk.source_text is the exact contiguous source material. Provenance and reassembly use source_text; the gate chunk_source_text_not_unique enforces that it maps uniquely back to its section.

oversized

A single pathological atom (a label glued to a very long sentence) can exceed the ceiling. The chunker sets oversized=True and flags it rather than cutting mid-dose. cli chunk prints the count; gate chunk_over_token_ceiling targets zero.

Chunk record (schema v4)

Field Type Notes
chunk_id str {drug_id}__{section_key}__{part_index} or {drug_id}__{section_key}__block__{table_id}
drug_id, drug_name str
section_key, section_display_name str
text str What is embedded. May carry repeated context labels.
source_text str Exact contiguous source material
context_labels str[] Labels repeated into text for retrieval only
heading_physical_page int
source_page_range [int,int] Physical (0-indexed PyMuPDF)
printed_page_range [int,int] The folio a clinician reads
atc_codes str[]
part_index, part_count int Position within the section
est_tokens, oversized int, bool
chunk_kind prose | block_descriptor
attachments ChunkAttachment[] Lifted tables/formulas
has_quarantined_content bool Derivable from attachments; stored anyway
schema_version int Must be exactly 4 at load time

The loader's REQUIRED_CHUNK_FIELDS check rejects a record missing any of chunk_id, drug_id, drug_name, section_key, text, source_text, heading_physical_page, source_page_range, printed_page_range, chunk_kind. _is_missing treats 0 and False as present and only None or an empty collection as absent — physical page 0 and has_quarantined_content=False are both legitimate.

Two-page addressing

Every citation carries both:

  • printed page — the folio printed in the book, what a clinician cites;
  • physical page — PyMuPDF's 0-indexed page in the PDF file, for the viewer (#page= fragments need +1).

packages/shared-types/src/dto/chat.ts documents this distinction on the Citation interface, and apps/web/app/api/chat/route.ts keeps a quarantined block's own physical page separate (quarantinePhysicalPage) because a table often sits on the page after the paragraph that mentions it — verified on real data, per the code comment.

Parent/child hydration

RetrievalDocument.parent_id and ParentDocument exist in the retrieval domain, and RetrievalService._hydrate will fetch a parent and use its text when a matched child names one. No chunk in the current corpus sets parent_idingestion/chunk/models.py has no such field, so the payload never carries it. The parent path is therefore currently inert for the loaded corpus; it is exercised only by tests and by the in-memory eval store.