# 04 — Ingestion pipeline Offline batch. **Never** part of the live request path (`ingestion/README.md`, and no import of `ingestion` exists anywhere in `apps/`). The pipeline has already been run. The artifacts below exist on disk and the corpus is loaded into Qdrant. ## Entrypoints | Command | Module | What it does | |---|---|---| | `python -m ingestion.cli run --pdf ` | `cli.py::_cmd_run` | extract → segment → `monographs.jsonl` | | `python -m ingestion.cli detect-tables --pdf ` | `_cmd_detect_tables` | locate + classify table regions → `table_regions.json` (slow, cached) | | `python -m ingestion.cli chunk --monographs … --pdf …` | `_cmd_chunk` | monographs → `chunks.jsonl` | | `python -m ingestion.cli chunk-ready --monographs … --chunks …` | `_cmd_chunk_ready` | run every acceptance gate; exit 1 on any failure | | `python -m ingestion.cli validate --pdf ` | `_cmd_validate` | recall/precision vs. the back-of-book index | | `python -m ingestion.cli coverage --pdf ` | `_cmd_coverage` | span-level ledger: where every span ended up | | `python -m ingestion.cli residual-ink --pdf ` | `_cmd_residual_ink` | ink on the page no extracted span accounts for | | `python -m ingestion.load.run --provider cohere-v4 --collection duocthu_v1` | `load/run.py::main` | embed (cached) + upsert + manifest | | `visual-diff`, `scaffold-golden` | `_cmd_not_implemented` | **`NotImplementedError`** — declared, never built | Note the split: `cli.py` stops at chunking. Embedding and loading live in a separate entrypoint precisely because that step spends money. ## Pipeline ```mermaid flowchart TD A[/"data/raw/*.pdf — 1,668 pages"/] B["extract_spans(doc)
extract/spans.py"] B2["load_transcribed_runs + merge_outlined_runs
extract/outlined_text.py, repair.py"] C["scan_glyph_order / scan_reading_order
extract/glyph_order.py — reports, does not correct"] D["_region_index()
table_regions.json + verified/formula_regions_2d.json"] E["segment.assemble(spans, table_index)
segment/assembler.py"] F[/"monographs.jsonl — 684"/] G["build_page_map(doc)
physical → printed folio"] H["chunk_all(monographs, header_rows, printed_page_map)
chunk/chunker.py"] I[/"chunks.jsonl — 15,100, schema v4"/] J["validation.evaluate + evaluate_chunks
named gates"] K["CachingEmbeddingProvider(BedrockCohere)
embed/cache.py, embed/bedrock_cohere.py"] L[/"embeddings cache — sha256-keyed"/] M["ChunkLoader.load()
load/upsert.py"] N[("duocthu_v1")] O[("duocthu_v1__manifest")] A --> B --> B2 --> E A --> C D --> E E --> F --> H --> I A --> G --> H F --> J I --> J I --> K --> L --> M --> N M --> O ``` ## Stage detail ### 1. Span extraction — `extract/spans.py` PyMuPDF (`fitz`) yields text spans in reading order with font flags, bbox, physical page and the printed folio resolved by `extract/page_map.py`. `extract/page_map.py` maps physical → printed folio by reading the isolated numeric token in each page's top 60pt header band. It does **not** hard-code the empirically constant `+1` offset, and it refuses to guess when two same-size candidates conflict (returns `None`). It prefers the largest-font candidate, because a real confirmed case — physical page 1243, `RIBOFLAVIN (Vitamin B2)` — had the title's subscript "2" fall into the header band next to the real folio, which previously dropped the entire monograph. ### 2. Vector-outlined text repair — `extract/outlined_text.py`, `repair.py` 51 runs of text in this PDF exist **only as vector paths**, so no extractor returns them: `"Độ ổn định"` came out as `"Độ n định"`. Human-transcribed runs in `data/verified/outlined_text_transcriptions.json` are merged back into the span stream by `_extracted_and_repaired_spans()`. Every command that builds monographs calls that same helper — the CLI comment says why: otherwise the coverage ledger would describe a different pipeline than the one producing the output. ### 3. Region index — tables and formulas `_region_index()` merges `data/processed/table_regions.json` (from `detect-tables`) with `data/verified/formula_regions_2d.json`, keyed by physical page. Spans falling inside a region are lifted out of prose. ### 4. Segmentation — `segment/assembler.py` (654 lines) See [05-document-parsing.md](05-document-parsing.md) for boundary detection. `assemble()` walks the classified event stream and emits `Monograph` objects with `sections`, `tables`, `preamble` and `atc_codes`. It raises `DuplicateDrugIdError` rather than silently merging two drugs with the same slug. `assemble()` optionally fills a `ledger` list — one row per span with a state (`prose`, `table`, `quarantined`, `boilerplate`, `unassigned`, …). That ledger is what `coverage` reports on. ### 5. Chunking — `chunk/chunker.py` See [06-document-model-and-chunking.md](06-document-model-and-chunking.md). `chunk_all()` **raises** if `printed_page_map` is `None`: > refusing to emit an embedding corpus without printed-page provenance ### 6. Gates — `validation/readiness.py` `chunk-ready` prints every gate with its count and target and exits non-zero if any fails. Gates on monographs: `outlined_run_not_merged`, `known_corruption_string`, `formula_fragment_in_prose`, `pua_char`, `replacement_char_ufffd`, `empty_section`, `section_without_provenance`, `part_without_source_span_ids`, `unflagged_quarantine_block`, `duplicate_table_id`, `duplicate_drug_id`, `monograph_without_page_range`. Gates on chunks (ADR 0006): `chunk_over_token_ceiling`, `chunk_without_printed_page_range`, `chunk_schema_version_not_supported`, `prose_without_source_text`, `chunk_source_text_not_unique`, `chunk_physical_range_not_exact`, `descriptor_range_not_attachment_page`, `attachment_without_printed_page`, `context_label_missing_from_text`, `section_not_reassemblable_from_chunks`, `section_block_without_chunk_reference`, `attachment_block_id_unknown`, `attachment_without_page_or_bbox`, `block_text_leaked_into_chunk_text`, `attachment_header_row_present`, `descriptor_with_unverified_columns`, `descriptor_chunk_without_attachment`, `descriptor_count_vs_block_count`. The command's own closing text names what the gates do **not** prove: > Not proven by these gates: content accuracy against the source (no > whole-document human-reviewed ground truth exists), table row/column > reconstruction, and recall for borderless tables and bar-less formulas. **Status: the gate values were not re-run in this documentation pass.** The gates exist and are tested (`ingestion/tests/test_validation_readiness.py`); the last recorded run is in `docs/progress-log.md`. ### 7. Embed + load — `load/run.py` ``` python -m ingestion.load.run \ --chunks data/processed/chunks.jsonl \ --provider cohere-v4 \ --collection duocthu_v1 \ --qdrant-url http://localhost:6333 \ [--embed-only] ``` - Texts are embedded in slices of 960 with 3 attempts and exponential backoff. - `CachingEmbeddingProvider` keys vectors by `(model_id, input_kind, sha256(text))`, so an interrupted run resumes and an unrelated chunk edit re-embeds only what changed. - `--embed-only` stops before the vector store. - The loader computes `corpus_sha256` over the whole `chunks.jsonl` and refuses to write into a collection built from a different corpus, model, dimension count or input kind (`load/manifest.py::assert_compatible`). - Exit code is `0` only if `collection_count == points_upserted`. ## Artifacts on disk | File | Size | Content | |---|---|---| | `data/raw/duoc-thu-quoc-gia-viet-nam-2018.pdf` | 37 MB | Source, committed | | `data/processed/monographs.jsonl` | 31 MB | 684 monographs | | `data/processed/chunks.jsonl` | 30 MB | 15,100 chunks, all `schema_version=4` | | `data/processed/coverage_ledger.json` | 52 MB | Per-span state ledger | | `data/processed/table_regions.json` | 136 KB | Classified table regions | | `data/processed/residual_ink.json` | 558 KB | Unaccounted-for ink regions | | `data/processed/glyph_extraction_ratio.json` | 40 KB | Per-page glyph accounting | | `data/processed/embeddings/` | — | Embedding cache | | `data/verified/drug_entities.json` | — | 684 entities, 10,164 aliases | | `data/verified/formula_regions_2d.json` | — | Human-verified 2-D formula regions | | `data/verified/outlined_text_transcriptions.json` | — | 51 transcribed vector-path runs | | `data/reconstruction/crops/*.png` | — | Crops of quarantined blocks | Verified this session by counting the files directly: ``` chunks: 15100 kinds: {'prose': 14949, 'block_descriptor': 151} schema_version: {4: 15100} distinct drug_id: 684 distinct section_key: 19 monographs: 684 drug entities: 684 / aliases: 10164 ``` ## Invariants the implementation actually enforces Each of these is a code path or a gate, not an aspiration: | Invariant | Enforced by | |---|---| | A chunk cannot be emitted without a printed-page range | `chunker.py::_page_ranges` raises; `load/models.py::_validate_page_range` raises | | Quarantined block text never appears in a prose chunk's `text` | `assembler.py` lifts region spans out; gate `block_text_leaked_into_chunk_text` | | A block descriptor's text is built from metadata only, never cell values | `chunker.py::describe_block`; `_attachment()` forces `header_row=[]` | | Every section must be reassemblable from its chunks | gate `section_not_reassemblable_from_chunks` | | A chunk's `source_text` must occur exactly once in its section | `_supporting_pages` raises otherwise; gate `chunk_source_text_not_unique` | | Two drugs cannot share a `drug_id` | `DuplicateDrugIdError`; gate `duplicate_drug_id` | | The same chunk always lands on the same Qdrant point | `point_id_for = uuid5(POINT_NAMESPACE, chunk_id)` | | A collection cannot mix two corpora or two models | `load/manifest.py::assert_compatible` → `CorpusMismatch` | | A collection with points but no manifest is refused | same function | ## Incremental processing Only the embedding step is incremental (content-hash cache). `run`, `chunk`, `detect-tables`, `coverage` and `residual-ink` are full-document passes with no caching between them beyond the JSON artifacts they write.