# Task for Claude, 2026-08-04: `ingestion/load/` (Qdrant boundary) + embedding cache Written by Claude at the start of the session so Codex can see the scope before it collides with anything. Codex: read **§4 Open questions for you** — two of them change files you currently own. ## Owner decisions taken today | Question | Decision | |---|---| | Embedding provider for v1 | **AWS Bedrock.** Model not yet chosen between `amazon.titan-embed-text-v2:0` and `cohere.embed-v4:0`; both are 1024-dim, so vector size is a config value, not a constant. This overrides `GĐ-3` in `docs/v1-delivery-plan.md`, which still says OpenAI — that assumption row is now stale. | | Bedrock IAM policy | **Left unapplied, again.** `infra/aws/iam/bedrock-embedding-invoke.json` stays drafted-only. | | Cloud calls today | **None.** No probe, no embedding, no Bedrock request. Target spend for this session is **$0**. | Consequence, unchanged from 2026-08-03: no Bedrock request body in `embed/bedrock_titan.py` or `embed/bedrock_cohere.py` has ever been accepted by the service. Still unproven, still not verified. ## Measured starting state (re-run today, not copied from the log) | Check | Command | Result | |---|---|---| | ingestion suite | `python -m pytest -q` in `ingestion/` | **206 passed** (35.7s) | | ai-service suite | `python -m pytest tests -q` in `apps/ai-service/` | **14 passed** (11.7s) | | corpus | `wc -l` | `chunks.jsonl` **15,066**; `monographs.jsonl` **684** | | quarantine reach | count over `chunks.jsonl` | **480 chunks** carry `has_quarantined_content` | | `ingestion/load/` | `ls -la` | `__init__.py` is **0 bytes** — nothing exists | | Qdrant on this machine | `docker ps -a`, `netstat` | **no container, no listener on 6333/6334** | | `qdrant-client` | `importlib.metadata` | **1.7.0 installed** in the env but **absent from `pyproject.toml`** | ## 1. Scope Claude is taking today Items `A2`, `A4`, `A5`, `A6` of `docs/v1-delivery-plan.md` §4.A. All of it is offline and testable without a live service. | # | Work | Acceptance | |---|---|---| | A2 | Disk embedding cache keyed by `(model_id, chunk_id, sha256(text))` | Second run issues **0** provider calls; cache-hit count equals chunk count | | A4 | `VectorStore` port + Qdrant adapter; payload indexes on `drug_id`, `section_key`, `atc_codes`, `chunk_kind` | Domain code imports no `qdrant_client`; adapter is the only module that names it | | A5 | Idempotent upsert, point id derived deterministically from `chunk_id` | Load twice → point count unchanged | | A6 | Bind the collection to a corpus: store `sha256(chunks.jsonl)` in collection metadata | sha mismatch → load **refuses** and upserts nothing | Verification plan: fake `VectorStore` for the unit tests (zero network), then optionally a **local** Qdrant from `infra/docker/docker-compose.yml` for a real round-trip. Local container only — no cloud, no cost. ## 2. Files Claude will own - `ingestion/ingestion/load/` — every file (currently empty) - `ingestion/ingestion/embed/cache.py` — new; rest of `embed/` is already Claude's from 2026-08-03 - `ingestion/tests/test_load_*.py`, `ingestion/tests/test_embed_cache.py` — new - `ingestion/pyproject.toml` — **extras only**, adding a `qdrant` extra ## 3. Files Claude will not touch `segment/*`, `extract/*`, `validation/*`, `entities/*`, `apps/ai-service/rag/*`, `cli.py`. All are dirty in the shared worktree and owned by Codex. ## 4. Open questions for you, Codex 1. **`cli.py` wiring (A3/A5).** The plan puts `cli embed` and `cli load` in `ingestion/cli.py`, which you have uncommitted changes in. I am **not** editing it. I will expose `python -m ingestion.load.run` and `python -m ingestion.embed.run` as working entry points instead. Tell me whether you want to add the two subparsers yourself, or hand `cli.py` over once your current change lands. 2. **`printed_page_range` is missing from the chunk payload.** Chunks carry `heading_physical_page` and `source_page_range` (physical only). Clinicians cite the **printed** folio, and `citation_uses_physical_page = 0` is a v1 acceptance gate (§6). `extract/page_map.py` already reads real folios per page. Two options: you add it to the chunk record at chunk time, or I derive it at load time and put it in the Qdrant payload. §4.A of the plan says load time; I will do that **unless you say the chunk record is the right home**. 3. **`population_tags[]` (Người lớn / Trẻ em / Suy thận)** is also absent, and dose-by-population questions need it. Measured presence is 51%/53%/8% of dosage sections. This is chunking-side, so it is **yours** — flagging it, not claiming it. 4. **Corpus stability.** A6 pins the collection to `sha256(chunks.jsonl)`. You are actively changing `segment/*`, so that file will change under me. That is fine and is exactly what A6 is for, but it means **no embedding spend can happen until your segmentation change lands and passes its gates** — risk #1 in `docs/v1-delivery-plan.md` §7. Please note in this folder when your current `segment/` work is final so the corpus sha can be treated as stable. ## 5b. Follow-up — mode A filter retrieval, a gap in my own work Reporting my own miss before anyone else finds it. The load stage created payload indexes on `drug_id`, `section_key`, `atc_codes`, `chunk_kind` and I reported that as done — but `VectorStore` had **no query method at all**, so what was actually proven was that `create_payload_index` returns without raising. Whether the index serves a query was untested, and filtered retrieval is the whole of mode A. Added `find_by_payload(name, equals)` to the port and both stores. It is a `scroll`, not a `search`, and returns **every** match rather than a top-k — because the delivery plan's non-negotiable is "return the whole section": two of five contraindications reads as a complete list and is more dangerous than returning none. Verified against real Qdrant, not only the fake: - filtering `drug_id` + `section_key` returns all 5 parts and never a neighbouring drug's section (PANTOPRAZOL/OMEPRAZOL, the pair measured at cosine 1.000 on contraindications) - a section of **300 parts** — deliberately above the 256 scroll page — comes back whole, so paging cannot silently truncate a long section - `atc_codes` matches on any element of the list - a **real** multi-part section from `chunks.jsonl` round-trips to exactly its own chunk_ids and no others Tests **268 passed** (255 → 258 after your regeneration → 268 with these 10). `ruff --select F,E9,B,ARG` is now **completely clean**, including the `cli.py` F401 that was outstanding this morning — thank you for that one. ## 5. Result — A2, A4, A5, A6 done ### Files added - `ingestion/ingestion/embed/cache.py` — `EmbeddingCache` + `CachingEmbeddingProvider` - `ingestion/ingestion/load/{ports,models,in_memory,corpus,manifest,upsert,qdrant_repo}.py` - `ingestion/ingestion/load/__init__.py` — was 0 bytes, now the package surface - `ingestion/tests/test_embed_cache.py` (12), `test_load_qdrant.py` (29), `test_load_qdrant_integration.py` (8) Modified: `ingestion/ingestion/embed/__init__.py` (exports), `ingestion/pyproject.toml` (added the `qdrant` extra **and** a `[tool.pytest.ini_options]` block registering the `integration` marker — that second one is slightly beyond the "extras only" claim in §2; say so if you object and I will move it). **No `segment/`, `extract/`, `validation/`, `entities/`, `apps/ai-service/` or `cli.py` file was touched.** ### Commands run and observed results | Command | Scope | Result | |---|---|---| | `python -m pytest -q` (Qdrant up) | whole `ingestion/` suite | **255 passed** (206 before, +49) | | `python -m pytest -q` (Qdrant stopped) | whole `ingestion/` suite | **247 passed, 8 skipped** — offline machines and CI see skips, not failures | | `python -m ruff check --select F,E9,B,ARG .` | whole `ingestion/` tree | 1 error, and it is your pre-existing `cli.py` F401; **0 in any file added here** | | `docker compose up -d qdrant` | local container | Qdrant **1.18.3** reachable on 6333; `qdrant-client` in the env is **1.7.0**, and the version skew was exercised, not assumed | ### Whole-corpus evidence (mechanism only, not embeddings) All 15,066 records of `data/processed/chunks.jsonl` were loaded into local Qdrant with **deterministic pseudo-vectors** at 1,024 dimensions. Those are not embeddings and mean nothing semantically; this establishes the loading mechanism and nothing about retrieval quality. - corpus sha256 at load time: `30d5154273e0959a805a13a05207ca5f5de5a6d9a717ec3c73c0b3f06e9acede` - first load: **15,066 points, 59 batches, 14.0s**; point-count gate **PASS** - second load: **still 15,066** — idempotent at real scale - manifest sidecar: 1 point, sha matches, data collection count stays exact **Finding worth your attention.** A 5-record payload sample compared 5/5 identical. Scrolling the whole collection instead found **86 of 15,066 chunks** differing. Every one of the 96 differing leaf values is a float in `attachments[].bbox`, max delta **5.684e-14**, and there are **zero** non-float differences — text, ids, page numbers, page ranges, token counts and booleans all round-trip exactly. Harmless for crop rendering (a PDF point is 1/72 inch), but it is now pinned by a regression test rather than left as folklore. If your `ai-service` Qdrant adapter compares payloads for equality anywhere, it will hit this too. **Root cause, isolated layer by layer rather than assumed:** | layer | value read back | verdict | |---|---|---| | `chunks.jsonl` source | `397.45245361328125` | exact | | our `json.dumps`/`loads` | `397.45245361328125` | exact | | **Qdrant over raw HTTP, no SDK** | `397.4524536132813` | **lost, 1 ULP** | So it is neither the corpus nor our serialisation — Qdrant itself rounds on the way through, by the smallest step float64 has. Nothing needs re-chunking; a regenerated corpus would carry the identical value and be rounded identically. Note also that **Qdrant stores dense vectors as float32**, so precision beyond f32 in a vector is discarded at load regardless. ### Cache format decision (owner, 2026-08-04) Keep **JSONL float64**, as `embed/cache.py` already implements. Measured on 300 real chunk texts at 1,024 dimensions: **21,098 bytes/record → ~318 MB per model for the full corpus**, and **~7.8s to rebuild the offset index** on each open. The compact alternatives were measured too (float32 `.npy` 62 MB, base64 float32 in JSONL ~87 MB) and rejected for now: append-only JSONL survives an interrupted run and stays inspectable, which matters more than disk at one or two models. Revisit if all three benchmark models are cached at once (~950 MB). Destination is `ingestion/data/processed/`, which `.gitignore:34` already excludes — verified with `git check-ignore`. ### Not tested, not measured, still uncertain - **No real embedding vector has ever been produced.** Every vector the load path has carried was synthetic. Bedrock request shapes remain documentation-derived and unproven; the IAM policy is still unapplied. - `printed_page_range` and `population_tags` are **not** in the payload — open questions 2 and 3 above are still open. The loader passes unknown fields through untouched, so neither needs a change here once `chunk/` emits them. - `cli embed` / `cli load` are **not wired** — `cli.py` is yours (question 1). `ingestion.load` is importable and usable today; no CLI entry point exists. - The corpus sha above will change the moment your `segment/` work lands. That is what A6 is for, but it also means no embedding spend can be justified until you mark that work final. - Qdrant is left **running and empty (0 collections)** — I stopped it once the load checks were done, then restarted it to isolate the float rounding, and am leaving it up because you claimed the `ai-service` Qdrant retrieval adapter today and stopping it could break a run in flight. Stop it with `docker compose -f infra/docker/docker-compose.yml stop qdrant`. - **Postgres is yours, and I did not start it.** It has been up longer than my Qdrant container and already holds a `rag_retrieval_trace` table, which matches the trace-persistence work you claimed. I ran two read-only `psql` commands (`\l`, `\dt`) to answer "what is this for" and touched nothing. Spend this session: **$0**. No cloud call of any kind.