7.8 KiB
06 — Document model and chunking
Entity model
erDiagram
MONOGRAPH ||--o{ SECTIONSPAN : sections
MONOGRAPH ||--o{ TABLEBLOCK : tables
MONOGRAPH ||--o{ SECTIONPART : preamble
SECTIONSPAN ||--|| HEADING : heading
SECTIONSPAN ||--o{ SECTIONPART : parts
SECTIONSPAN ||--o{ CHUNK : "prose chunks"
TABLEBLOCK ||--|| CHUNK : "1 block_descriptor chunk"
CHUNK ||--o{ CHUNKATTACHMENT : attachments
CHUNK ||--|| VECTORPOINT : "uuid5(chunk_id)"
MONOGRAPH {
string drug_id PK
string drug_name
int_list source_page_range
string_list atc_codes
bool atc_stated_absent
}
SECTIONSPAN {
string key
string display_name
string text
}
SECTIONPART {
string kind "prose|table"
string text
int physical_page
float_list bbox
string_list source_span_ids
bool quarantined
}
TABLEBLOCK {
string table_id PK
string shape
int physical_page
float_list bbox
string section_key
bool quarantined
}
CHUNK {
string chunk_id PK
string drug_id FK
string section_key
string text
string source_text
int_list source_page_range
int_list printed_page_range
int part_index
int part_count
string chunk_kind
bool has_quarantined_content
int schema_version
}
CHUNKATTACHMENT {
string block_id FK
string kind "table|formula"
int physical_page
float_list bbox
int printed_page
bool quarantined
}
Source: ingestion/segment/models.py, ingestion/chunk/models.py,
ingestion/load/models.py.
The 19 section keys
Book order, as defined in apps/ai-service/rag/sections.py::SECTION_ORDER
(18 entries — ten_thuong_mai exists in the vocabulary but not in the ordering
tuple) and rag/understanding.py::SECTION_KEYS (all 19):
ten_chung_quoc_te, ten_thuong_mai, ma_atc, loai_thuoc,
dang_thuoc_va_ham_luong, duoc_ly_va_co_che_tac_dung, chi_dinh,
chong_chi_dinh, than_trong, thoi_ky_mang_thai, thoi_ky_cho_con_bu,
tac_dung_khong_mong_muon, huong_dan_xu_tri_adr, lieu_luong_va_cach_dung,
tuong_tac_thuoc, qua_lieu_va_xu_tri, do_on_dinh_va_bao_quan, tuong_ky,
thong_tin_quy_che.
All 19 appear in the loaded corpus. Chunk counts per section (counted this
session over chunks.jsonl):
| Section | Chunks |
|---|---|
duoc_ly_va_co_che_tac_dung |
1,896 |
lieu_luong_va_cach_dung |
1,873 |
than_trong |
927 |
tac_dung_khong_mong_muon |
857 |
tuong_tac_thuoc |
810 |
chi_dinh |
710 |
dang_thuoc_va_ham_luong |
691 |
ten_chung_quoc_te |
684 |
The two largest sections being pharmacology and dosage is exactly why
rag/sections.py exists — see 09-retrieval-pipeline.md.
Chunking strategy (ADR 0004)
Unit: (drug_id, section_key). A section under the token ceiling becomes
one chunk, verbatim. Only the long tail is sub-chunked.
CEILING_TOKENS = 800 # above this, sub-chunk
TARGET_TOKENS = 650 # packing target
OVERLAP_TOKENS = 65 # sliding-window overlap
Token counting uses tiktoken cl100k_base when available, and an estimate
otherwise — cli chunk prints which one it used.
Sub-chunking
- Atomise (
_atoms): split into sentences (chunk/sentences.py, which treats:as a boundary). A "sentence" longer thanTARGET_TOKENSthat contains commas is split on commas — needed because a drug-interaction list is one grammatical sentence hundreds of names long:VORICONAZOL'stương tác thuốcproduced 981- and 888-token parts, and a truncated interaction list reads as "this drug is not listed", a false negative in the dangerous direction. - Pack (
_pack_parts): greedily fill toTARGET_TOKENS, then overlap the tail by up toOVERLAP_TOKENS.
The clinical-context rules inside the packer
These are the non-obvious part, and each exists for a measured defect:
- Never end a part on a label.
"Người lớn: 500 mg mỗi 8 giờ."splits after the colon; flushing there would leave a chunk ending"Người lớn:"with the dose in the next one. Measured before the rule: 38 such chunks. A dose separated from the population it applies to is a patient-safety defect. - Carry the governing label forward.
contexts/scope_contexts/context_chain()track the active label and its parent scope per atom, so a population label that fell out of both the 650-token buffer and the 65-token overlap several parts ago is repeated at the seam. - Split a trailing label off compound atoms.
_split_trailing_labelhandles"7,5 mg … .\nBước 5:"so the dose at the atom's start does not loseBước 4. - Repeated labels are marked as context, not source.
Chunk.textmay contain a prepended label;Chunk.source_textis the exact contiguous source material. Provenance and reassembly usesource_text; the gatechunk_source_text_not_uniqueenforces that it maps uniquely back to its section.
oversized
A single pathological atom (a label glued to a very long sentence) can exceed
the ceiling. The chunker sets oversized=True and flags it rather than cutting
mid-dose. cli chunk prints the count; gate chunk_over_token_ceiling targets
zero.
Chunk record (schema v4)
| Field | Type | Notes |
|---|---|---|
chunk_id |
str | {drug_id}__{section_key}__{part_index} or {drug_id}__{section_key}__block__{table_id} |
drug_id, drug_name |
str | |
section_key, section_display_name |
str | |
text |
str | What is embedded. May carry repeated context labels. |
source_text |
str | Exact contiguous source material |
context_labels |
str[] | Labels repeated into text for retrieval only |
heading_physical_page |
int | |
source_page_range |
[int,int] | Physical (0-indexed PyMuPDF) |
printed_page_range |
[int,int] | The folio a clinician reads |
atc_codes |
str[] | |
part_index, part_count |
int | Position within the section |
est_tokens, oversized |
int, bool | |
chunk_kind |
prose | block_descriptor |
|
attachments |
ChunkAttachment[] | Lifted tables/formulas |
has_quarantined_content |
bool | Derivable from attachments; stored anyway |
schema_version |
int | Must be exactly 4 at load time |
The loader's REQUIRED_CHUNK_FIELDS check rejects a record missing any of
chunk_id, drug_id, drug_name, section_key, text, source_text,
heading_physical_page, source_page_range, printed_page_range,
chunk_kind. _is_missing treats 0 and False as present and only None or
an empty collection as absent — physical page 0 and
has_quarantined_content=False are both legitimate.
Two-page addressing
Every citation carries both:
- printed page — the folio printed in the book, what a clinician cites;
- physical page — PyMuPDF's 0-indexed page in the PDF file, for the viewer
(
#page=fragments need+1).
packages/shared-types/src/dto/chat.ts documents this distinction on the
Citation interface, and apps/web/app/api/chat/route.ts keeps a quarantined
block's own physical page separate (quarantinePhysicalPage) because a table
often sits on the page after the paragraph that mentions it — verified on real
data, per the code comment.
Parent/child hydration
RetrievalDocument.parent_id and ParentDocument exist in the retrieval
domain, and RetrievalService._hydrate will fetch a parent and use its text
when a matched child names one. No chunk in the current corpus sets
parent_id — ingestion/chunk/models.py has no such field, so the payload
never carries it. The parent path is therefore currently inert for the loaded
corpus; it is exercised only by tests and by the in-memory eval store.