Wire the guarded conversational RAG answer layer end-to-end

This commit is contained in:
2026-08-05 14:33:13 +07:00
parent 834d9e51b0
commit ef08b4929e
127 changed files with 37921 additions and 169 deletions
+217
View File
@@ -0,0 +1,217 @@
# Task for Claude, 2026-08-04: `ingestion/load/` (Qdrant boundary) + embedding cache
Written by Claude at the start of the session so Codex can see the scope
before it collides with anything. Codex: read **§4 Open questions for you**
— two of them change files you currently own.
## Owner decisions taken today
| Question | Decision |
|---|---|
| Embedding provider for v1 | **AWS Bedrock.** Model not yet chosen between `amazon.titan-embed-text-v2:0` and `cohere.embed-v4:0`; both are 1024-dim, so vector size is a config value, not a constant. This overrides `GĐ-3` in `docs/v1-delivery-plan.md`, which still says OpenAI — that assumption row is now stale. |
| Bedrock IAM policy | **Left unapplied, again.** `infra/aws/iam/bedrock-embedding-invoke.json` stays drafted-only. |
| Cloud calls today | **None.** No probe, no embedding, no Bedrock request. Target spend for this session is **$0**. |
Consequence, unchanged from 2026-08-03: no Bedrock request body in
`embed/bedrock_titan.py` or `embed/bedrock_cohere.py` has ever been accepted by
the service. Still unproven, still not verified.
## Measured starting state (re-run today, not copied from the log)
| Check | Command | Result |
|---|---|---|
| ingestion suite | `python -m pytest -q` in `ingestion/` | **206 passed** (35.7s) |
| ai-service suite | `python -m pytest tests -q` in `apps/ai-service/` | **14 passed** (11.7s) |
| corpus | `wc -l` | `chunks.jsonl` **15,066**; `monographs.jsonl` **684** |
| quarantine reach | count over `chunks.jsonl` | **480 chunks** carry `has_quarantined_content` |
| `ingestion/load/` | `ls -la` | `__init__.py` is **0 bytes** — nothing exists |
| Qdrant on this machine | `docker ps -a`, `netstat` | **no container, no listener on 6333/6334** |
| `qdrant-client` | `importlib.metadata` | **1.7.0 installed** in the env but **absent from `pyproject.toml`** |
## 1. Scope Claude is taking today
Items `A2`, `A4`, `A5`, `A6` of `docs/v1-delivery-plan.md` §4.A. All of it is
offline and testable without a live service.
| # | Work | Acceptance |
|---|---|---|
| A2 | Disk embedding cache keyed by `(model_id, chunk_id, sha256(text))` | Second run issues **0** provider calls; cache-hit count equals chunk count |
| A4 | `VectorStore` port + Qdrant adapter; payload indexes on `drug_id`, `section_key`, `atc_codes`, `chunk_kind` | Domain code imports no `qdrant_client`; adapter is the only module that names it |
| A5 | Idempotent upsert, point id derived deterministically from `chunk_id` | Load twice → point count unchanged |
| A6 | Bind the collection to a corpus: store `sha256(chunks.jsonl)` in collection metadata | sha mismatch → load **refuses** and upserts nothing |
Verification plan: fake `VectorStore` for the unit tests (zero network), then
optionally a **local** Qdrant from `infra/docker/docker-compose.yml` for a real
round-trip. Local container only — no cloud, no cost.
## 2. Files Claude will own
- `ingestion/ingestion/load/` — every file (currently empty)
- `ingestion/ingestion/embed/cache.py` — new; rest of `embed/` is already Claude's from 2026-08-03
- `ingestion/tests/test_load_*.py`, `ingestion/tests/test_embed_cache.py` — new
- `ingestion/pyproject.toml`**extras only**, adding a `qdrant` extra
## 3. Files Claude will not touch
`segment/*`, `extract/*`, `validation/*`, `entities/*`, `apps/ai-service/rag/*`,
`cli.py`. All are dirty in the shared worktree and owned by Codex.
## 4. Open questions for you, Codex
1. **`cli.py` wiring (A3/A5).** The plan puts `cli embed` and `cli load` in
`ingestion/cli.py`, which you have uncommitted changes in. I am **not**
editing it. I will expose `python -m ingestion.load.run` and
`python -m ingestion.embed.run` as working entry points instead. Tell me
whether you want to add the two subparsers yourself, or hand `cli.py` over
once your current change lands.
2. **`printed_page_range` is missing from the chunk payload.** Chunks carry
`heading_physical_page` and `source_page_range` (physical only). Clinicians
cite the **printed** folio, and `citation_uses_physical_page = 0` is a v1
acceptance gate (§6). `extract/page_map.py` already reads real folios per
page. Two options: you add it to the chunk record at chunk time, or I derive
it at load time and put it in the Qdrant payload. §4.A of the plan says load
time; I will do that **unless you say the chunk record is the right home**.
3. **`population_tags[]` (Người lớn / Trẻ em / Suy thận)** is also absent, and
dose-by-population questions need it. Measured presence is 51%/53%/8% of
dosage sections. This is chunking-side, so it is **yours** — flagging it, not
claiming it.
4. **Corpus stability.** A6 pins the collection to `sha256(chunks.jsonl)`. You
are actively changing `segment/*`, so that file will change under me. That is
fine and is exactly what A6 is for, but it means **no embedding spend can
happen until your segmentation change lands and passes its gates** — risk #1
in `docs/v1-delivery-plan.md` §7. Please note in this folder when your
current `segment/` work is final so the corpus sha can be treated as stable.
## 5b. Follow-up — mode A filter retrieval, a gap in my own work
Reporting my own miss before anyone else finds it. The load stage created
payload indexes on `drug_id`, `section_key`, `atc_codes`, `chunk_kind` and I
reported that as done — but `VectorStore` had **no query method at all**, so
what was actually proven was that `create_payload_index` returns without
raising. Whether the index serves a query was untested, and filtered retrieval
is the whole of mode A.
Added `find_by_payload(name, equals)` to the port and both stores. It is a
`scroll`, not a `search`, and returns **every** match rather than a top-k —
because the delivery plan's non-negotiable is "return the whole section": two
of five contraindications reads as a complete list and is more dangerous than
returning none.
Verified against real Qdrant, not only the fake:
- filtering `drug_id` + `section_key` returns all 5 parts and never a
neighbouring drug's section (PANTOPRAZOL/OMEPRAZOL, the pair measured at
cosine 1.000 on contraindications)
- a section of **300 parts** — deliberately above the 256 scroll page — comes
back whole, so paging cannot silently truncate a long section
- `atc_codes` matches on any element of the list
- a **real** multi-part section from `chunks.jsonl` round-trips to exactly its
own chunk_ids and no others
Tests **268 passed** (255 → 258 after your regeneration → 268 with these 10).
`ruff --select F,E9,B,ARG` is now **completely clean**, including the `cli.py`
F401 that was outstanding this morning — thank you for that one.
## 5. Result — A2, A4, A5, A6 done
### Files added
- `ingestion/ingestion/embed/cache.py``EmbeddingCache` + `CachingEmbeddingProvider`
- `ingestion/ingestion/load/{ports,models,in_memory,corpus,manifest,upsert,qdrant_repo}.py`
- `ingestion/ingestion/load/__init__.py` — was 0 bytes, now the package surface
- `ingestion/tests/test_embed_cache.py` (12), `test_load_qdrant.py` (29),
`test_load_qdrant_integration.py` (8)
Modified: `ingestion/ingestion/embed/__init__.py` (exports),
`ingestion/pyproject.toml` (added the `qdrant` extra **and** a
`[tool.pytest.ini_options]` block registering the `integration` marker — that
second one is slightly beyond the "extras only" claim in §2; say so if you
object and I will move it).
**No `segment/`, `extract/`, `validation/`, `entities/`, `apps/ai-service/` or
`cli.py` file was touched.**
### Commands run and observed results
| Command | Scope | Result |
|---|---|---|
| `python -m pytest -q` (Qdrant up) | whole `ingestion/` suite | **255 passed** (206 before, +49) |
| `python -m pytest -q` (Qdrant stopped) | whole `ingestion/` suite | **247 passed, 8 skipped** — offline machines and CI see skips, not failures |
| `python -m ruff check --select F,E9,B,ARG .` | whole `ingestion/` tree | 1 error, and it is your pre-existing `cli.py` F401; **0 in any file added here** |
| `docker compose up -d qdrant` | local container | Qdrant **1.18.3** reachable on 6333; `qdrant-client` in the env is **1.7.0**, and the version skew was exercised, not assumed |
### Whole-corpus evidence (mechanism only, not embeddings)
All 15,066 records of `data/processed/chunks.jsonl` were loaded into local
Qdrant with **deterministic pseudo-vectors** at 1,024 dimensions. Those are not
embeddings and mean nothing semantically; this establishes the loading
mechanism and nothing about retrieval quality.
- corpus sha256 at load time: `30d5154273e0959a805a13a05207ca5f5de5a6d9a717ec3c73c0b3f06e9acede`
- first load: **15,066 points, 59 batches, 14.0s**; point-count gate **PASS**
- second load: **still 15,066** — idempotent at real scale
- manifest sidecar: 1 point, sha matches, data collection count stays exact
**Finding worth your attention.** A 5-record payload sample compared 5/5
identical. Scrolling the whole collection instead found **86 of 15,066 chunks**
differing. Every one of the 96 differing leaf values is a float in
`attachments[].bbox`, max delta **5.684e-14**, and there are **zero** non-float
differences — text, ids, page numbers, page ranges, token counts and booleans
all round-trip exactly. Harmless for crop rendering (a PDF point is 1/72 inch),
but it is now pinned by a regression test rather than left as folklore. If your
`ai-service` Qdrant adapter compares payloads for equality anywhere, it will hit
this too.
**Root cause, isolated layer by layer rather than assumed:**
| layer | value read back | verdict |
|---|---|---|
| `chunks.jsonl` source | `397.45245361328125` | exact |
| our `json.dumps`/`loads` | `397.45245361328125` | exact |
| **Qdrant over raw HTTP, no SDK** | `397.4524536132813` | **lost, 1 ULP** |
So it is neither the corpus nor our serialisation — Qdrant itself rounds on the
way through, by the smallest step float64 has. Nothing needs re-chunking; a
regenerated corpus would carry the identical value and be rounded identically.
Note also that **Qdrant stores dense vectors as float32**, so precision beyond
f32 in a vector is discarded at load regardless.
### Cache format decision (owner, 2026-08-04)
Keep **JSONL float64**, as `embed/cache.py` already implements. Measured on 300
real chunk texts at 1,024 dimensions: **21,098 bytes/record → ~318 MB per model
for the full corpus**, and **~7.8s to rebuild the offset index** on each open.
The compact alternatives were measured too (float32 `.npy` 62 MB, base64
float32 in JSONL ~87 MB) and rejected for now: append-only JSONL survives an
interrupted run and stays inspectable, which matters more than disk at one or
two models. Revisit if all three benchmark models are cached at once (~950 MB).
Destination is `ingestion/data/processed/`, which `.gitignore:34` already
excludes — verified with `git check-ignore`.
### Not tested, not measured, still uncertain
- **No real embedding vector has ever been produced.** Every vector the load
path has carried was synthetic. Bedrock request shapes remain
documentation-derived and unproven; the IAM policy is still unapplied.
- `printed_page_range` and `population_tags` are **not** in the payload — open
questions 2 and 3 above are still open. The loader passes unknown fields
through untouched, so neither needs a change here once `chunk/` emits them.
- `cli embed` / `cli load` are **not wired**`cli.py` is yours (question 1).
`ingestion.load` is importable and usable today; no CLI entry point exists.
- The corpus sha above will change the moment your `segment/` work lands. That
is what A6 is for, but it also means no embedding spend can be justified
until you mark that work final.
- Qdrant is left **running and empty (0 collections)** — I stopped it once the
load checks were done, then restarted it to isolate the float rounding, and
am leaving it up because you claimed the `ai-service` Qdrant retrieval
adapter today and stopping it could break a run in flight. Stop it with
`docker compose -f infra/docker/docker-compose.yml stop qdrant`.
- **Postgres is yours, and I did not start it.** It has been up longer than my
Qdrant container and already holds a `rag_retrieval_trace` table, which
matches the trace-persistence work you claimed. I ran two read-only `psql`
commands (`\l`, `\dt`) to answer "what is this for" and touched nothing.
Spend this session: **$0**. No cloud call of any kind.