Files
duocthu/docs-legacy/06-document-model-and-chunking.md
T

206 lines
7.8 KiB
Markdown

# 06 — Document model and chunking
## Entity model
```mermaid
erDiagram
MONOGRAPH ||--o{ SECTIONSPAN : sections
MONOGRAPH ||--o{ TABLEBLOCK : tables
MONOGRAPH ||--o{ SECTIONPART : preamble
SECTIONSPAN ||--|| HEADING : heading
SECTIONSPAN ||--o{ SECTIONPART : parts
SECTIONSPAN ||--o{ CHUNK : "prose chunks"
TABLEBLOCK ||--|| CHUNK : "1 block_descriptor chunk"
CHUNK ||--o{ CHUNKATTACHMENT : attachments
CHUNK ||--|| VECTORPOINT : "uuid5(chunk_id)"
MONOGRAPH {
string drug_id PK
string drug_name
int_list source_page_range
string_list atc_codes
bool atc_stated_absent
}
SECTIONSPAN {
string key
string display_name
string text
}
SECTIONPART {
string kind "prose|table"
string text
int physical_page
float_list bbox
string_list source_span_ids
bool quarantined
}
TABLEBLOCK {
string table_id PK
string shape
int physical_page
float_list bbox
string section_key
bool quarantined
}
CHUNK {
string chunk_id PK
string drug_id FK
string section_key
string text
string source_text
int_list source_page_range
int_list printed_page_range
int part_index
int part_count
string chunk_kind
bool has_quarantined_content
int schema_version
}
CHUNKATTACHMENT {
string block_id FK
string kind "table|formula"
int physical_page
float_list bbox
int printed_page
bool quarantined
}
```
Source: `ingestion/segment/models.py`, `ingestion/chunk/models.py`,
`ingestion/load/models.py`.
## The 19 section keys
Book order, as defined in `apps/ai-service/rag/sections.py::SECTION_ORDER`
(18 entries — `ten_thuong_mai` exists in the vocabulary but not in the ordering
tuple) and `rag/understanding.py::SECTION_KEYS` (all 19):
`ten_chung_quoc_te`, `ten_thuong_mai`, `ma_atc`, `loai_thuoc`,
`dang_thuoc_va_ham_luong`, `duoc_ly_va_co_che_tac_dung`, `chi_dinh`,
`chong_chi_dinh`, `than_trong`, `thoi_ky_mang_thai`, `thoi_ky_cho_con_bu`,
`tac_dung_khong_mong_muon`, `huong_dan_xu_tri_adr`, `lieu_luong_va_cach_dung`,
`tuong_tac_thuoc`, `qua_lieu_va_xu_tri`, `do_on_dinh_va_bao_quan`, `tuong_ky`,
`thong_tin_quy_che`.
All 19 appear in the loaded corpus. Chunk counts per section (counted this
session over `chunks.jsonl`):
| Section | Chunks |
|---|---|
| `duoc_ly_va_co_che_tac_dung` | 1,896 |
| `lieu_luong_va_cach_dung` | 1,873 |
| `than_trong` | 927 |
| `tac_dung_khong_mong_muon` | 857 |
| `tuong_tac_thuoc` | 810 |
| `chi_dinh` | 710 |
| `dang_thuoc_va_ham_luong` | 691 |
| `ten_chung_quoc_te` | 684 |
The two largest sections being pharmacology and dosage is exactly why
`rag/sections.py` exists — see [09-retrieval-pipeline.md](09-retrieval-pipeline.md).
## Chunking strategy (ADR 0004)
**Unit: `(drug_id, section_key)`.** A section under the token ceiling becomes
**one chunk, verbatim**. Only the long tail is sub-chunked.
```python
CEILING_TOKENS = 800 # above this, sub-chunk
TARGET_TOKENS = 650 # packing target
OVERLAP_TOKENS = 65 # sliding-window overlap
```
Token counting uses `tiktoken` `cl100k_base` when available, and an estimate
otherwise — `cli chunk` prints which one it used.
### Sub-chunking
1. **Atomise** (`_atoms`): split into sentences (`chunk/sentences.py`, which
treats `:` as a boundary). A "sentence" longer than `TARGET_TOKENS` that
contains commas is split on commas — needed because a drug-interaction list
is one grammatical sentence hundreds of names long: `VORICONAZOL`'s
`tương tác thuốc` produced 981- and 888-token parts, and a truncated
interaction list reads as *"this drug is not listed"*, a false negative in
the dangerous direction.
2. **Pack** (`_pack_parts`): greedily fill to `TARGET_TOKENS`, then overlap the
tail by up to `OVERLAP_TOKENS`.
### The clinical-context rules inside the packer
These are the non-obvious part, and each exists for a measured defect:
- **Never end a part on a label.** `"Người lớn: 500 mg mỗi 8 giờ."` splits after
the colon; flushing there would leave a chunk ending `"Người lớn:"` with the
dose in the next one. Measured before the rule: 38 such chunks. A dose
separated from the population it applies to is a patient-safety defect.
- **Carry the governing label forward.** `contexts` / `scope_contexts` /
`context_chain()` track the active label *and* its parent scope per atom, so a
population label that fell out of both the 650-token buffer and the 65-token
overlap several parts ago is repeated at the seam.
- **Split a trailing label off compound atoms.** `_split_trailing_label` handles
`"7,5 mg … .\nBước 5:"` so the dose at the atom's start does not lose
`Bước 4`.
- **Repeated labels are marked as context, not source.** `Chunk.text` may
contain a prepended label; `Chunk.source_text` is the exact contiguous source
material. Provenance and reassembly use `source_text`; the gate
`chunk_source_text_not_unique` enforces that it maps uniquely back to its
section.
### `oversized`
A single pathological atom (a label glued to a very long sentence) can exceed
the ceiling. The chunker sets `oversized=True` and flags it rather than cutting
mid-dose. `cli chunk` prints the count; gate `chunk_over_token_ceiling` targets
zero.
## Chunk record (schema v4)
| Field | Type | Notes |
|---|---|---|
| `chunk_id` | str | `{drug_id}__{section_key}__{part_index}` or `{drug_id}__{section_key}__block__{table_id}` |
| `drug_id`, `drug_name` | str | |
| `section_key`, `section_display_name` | str | |
| `text` | str | What is embedded. May carry repeated context labels. |
| `source_text` | str | Exact contiguous source material |
| `context_labels` | str[] | Labels repeated into `text` for retrieval only |
| `heading_physical_page` | int | |
| `source_page_range` | [int,int] | Physical (0-indexed PyMuPDF) |
| `printed_page_range` | [int,int] | The folio a clinician reads |
| `atc_codes` | str[] | |
| `part_index`, `part_count` | int | Position within the section |
| `est_tokens`, `oversized` | int, bool | |
| `chunk_kind` | `prose` \| `block_descriptor` | |
| `attachments` | ChunkAttachment[] | Lifted tables/formulas |
| `has_quarantined_content` | bool | Derivable from `attachments`; stored anyway |
| `schema_version` | int | Must be exactly `4` at load time |
The loader's `REQUIRED_CHUNK_FIELDS` check rejects a record missing any of
`chunk_id`, `drug_id`, `drug_name`, `section_key`, `text`, `source_text`,
`heading_physical_page`, `source_page_range`, `printed_page_range`,
`chunk_kind`. `_is_missing` treats `0` and `False` as present and only `None` or
an empty collection as absent — physical page 0 and
`has_quarantined_content=False` are both legitimate.
## Two-page addressing
Every citation carries both:
- **printed page** — the folio printed in the book, what a clinician cites;
- **physical page** — PyMuPDF's 0-indexed page in the PDF file, for the viewer
(`#page=` fragments need `+1`).
`packages/shared-types/src/dto/chat.ts` documents this distinction on the
`Citation` interface, and `apps/web/app/api/chat/route.ts` keeps a quarantined
block's *own* physical page separate (`quarantinePhysicalPage`) because a table
often sits on the page after the paragraph that mentions it — verified on real
data, per the code comment.
## Parent/child hydration
`RetrievalDocument.parent_id` and `ParentDocument` exist in the retrieval
domain, and `RetrievalService._hydrate` will fetch a parent and use its text
when a matched child names one. **No chunk in the current corpus sets
`parent_id`** — `ingestion/chunk/models.py` has no such field, so the payload
never carries it. The parent path is therefore currently inert for the loaded
corpus; it is exercised only by tests and by the in-memory eval store.