# ingestion Offline batch pipeline (never part of the live `ai-service` request path): extract -> segment -> chunk -> embed -> load. Parses `data/raw/duoc-thu-quoc-gia-viet-nam-2018.pdf` into per-drug, per-section chunks and upserts embeddings into Qdrant. Run via `python -m ingestion.cli run --pdf data/raw/duoc-thu-quoc-gia-viet-nam-2018.pdf`. **This has already been run.** Qdrant collection `duocthu_v1` holds 15,100 points (14,949 prose + 151 block descriptors) embedded with `cohere.embed-v4:0`, and a `duocthu_v1__manifest` sidecar pins the corpus sha/model/dimensions that `ai-service` checks at startup. Re-running the embed step costs **real Bedrock spend** on a personal AWS account, so do not start a corpus run without the owner's explicit go for that specific run. To move an existing corpus between machines, snapshot and restore the Qdrant collection instead — it is free and exact. Scope: the corpus covers **Part 2 monographs (printed pages 99-1496) only**. Part 1 general chapters and Part 3 appendices (BSA table, IV preparation, ATC index) are deliberately excluded, so questions about them correctly abstain rather than being answered from a neighbouring section.