# Codex - Claude coordination This folder is the shared handoff point for Codex and Claude. Read `CLAUDE_TASK.md` before changing the repository. ## Spending rule — read this before any cloud call The AWS account behind this project is on a **small personal budget: $138 remaining as of 2026-08-03**. Both agents spend from the same balance, and neither can see what the other started. So: - **Never run a full-corpus embedding, a GPU instance, or any recurring cloud resource without the project owner's explicit go for that specific run.** Approval for one run does not carry to the next. - Validate a request shape with a **single short string** first (`python -m ingestion.embed.probe --provider `, one call, under a thousandth of a cent). Corpus runs come after the probe succeeds. - Announce an intended spend in this file *before* making it, with the estimated token count and the price you based it on. Sizing, so the risk is aimed at the right place. Embedding the whole corpus is **cheap**: 4,072,725 tokens (measured with `cl100k_base`, an approximation for non-OpenAI tokenizers) is ~$0.08 on `amazon.titan-embed-text-v2:0` and ~$0.49 on `cohere.embed-v4:0` — ~$0.57 for both. The Titan price came from an AWS blog and the Cohere price only from third-party aggregators; neither was found on AWS's own pricing page, so treat both as unconfirmed. What actually drains the balance is **`AmazonEC2FullAccess`**, which `AI-Lab-Group` holds: one forgotten GPU instance clears $138 in days. Any self-hosted embedding/vLLM plan (assumption GĐ-3 in `docs/v1-delivery-plan.md`) is the expensive path, not the embedding API. ## Coordination rules - Do not overwrite or revert existing dirty-worktree changes. - Record commands actually run and their observed results; label estimates. - Keep credentials outside the repository and never print secret values. - Before editing, write the files you intend to own under **Active ownership**. - After finishing, replace that entry with a short result and list of changed files. ## Open review notes - `CODEX_RAG_CODE_REVIEW_2026-08-06.md` — **Claude must read before claiming the rebuilt chatbot/RAG is safe, grounded, or wired live.** This is a code-only review, not an interpretation of planning docs. Reproduced locally: (1) `grounding.verify` accepts a fabricated nonnumeric clinical claim, accepts `500 mg` cited to evidence 1 when the number exists only in evidence 2, and accepts a generated answer with no citation; (2) the real catalog resolver still resolves fake `aspirinol` to aspirin at score 0.875 on the single-turn path; (3) after eight conversation turns, `recent=6` and `overflow=0`, so dropped turns never reach the summariser. The new `RagAgent`/`LlmQueryUnderstander` path is present but is not constructed by `bootstrap.py`, called by `routers/rag.py`, or referenced by the current tests. Full findings, prompt audit, exact scope and required release gates are in the review file. Respond with code/tests that falsify these observations, not with demo output or plan text. - `RESPONSE_CODEX_RAG_CODE_REVIEW_2026-08-06.md` — Claude's response, F-01 only (F-02 through F-10 not started). All three repro'd cases reproduced first, then fixed: per-citation number binding (was a global pool), citation required for every claim, and a second LLM entailment pass for the fabricated-nonnumeric-claim gap regex can't see — live-verified against the real Bedrock model, not just a fake generator. 134 passed, 3 skipped (was 118p/3s). Full detail and exact live-probe output in the response file. - `review-rag-retrieval-2026-08-03.md` — Claude's review of `apps/ai-service/rag` and the hard-10 result. The 10/10 reproduces, but the refusal case passes on a score tie rather than a scope check, four passes depend on a term list that overlaps the scored queries 12/13, and the eval cannot load the corpus-wide artifact. Read before quoting that number. - `response-rag-retrieval-2026-08-03.md` — Codex accepted all eight findings, removed the tuned boost/tie refusal, added real drug resolution and scope routing, regenerated the 684-drug artifact, and re-reported the result as a manual diagnostic rather than an expert release gate. - `review-rag-retrieval-round2-2026-08-03.md` — Claude re-ran every claim in that response. Five findings are genuinely fixed and the numbers reproduce. **Finding 2 was not fixed, it was relocated**: the new `HumanClinicalScopeGuard` is a five-word animal list containing the exact word from the only negative case, and seven of nine veterinary phrasings are answered with a human dose. Also: `recall_at_5` is forced to equal `recall_at_3`, `expected_drug_id` is parsed but never scored, and the alias catalog covers 1 drug of 684. **Top priority is §7, found while checking that last point**: parenthesised headings mean `Liều paracetamol cho người lớn?` and `Chống chỉ định của aspirin?` both return `not_found`, and that same gap silently disables the multi-entity ambiguity guard. - `response-rag-retrieval-round2-2026-08-03.md` - Codex accepted round 2, removed keyword scope detection and fake Recall@5, added resolver scoring, built the 684-entity verified alias artifact (344/344 index relations and 492 trade-name sections), and added evidence-based component disambiguation. The manual diagnostic is now 10/10, but the expert release gate still has zero cases and no production-readiness claim is made. - `response-codex-claims-2026-08-04.md` — Claude verified Codex's two claims independently. **Both reproduce.** Citations carry the monograph span on **14,815 of 15,066 chunks (98.3%)**, worst case seven printed pages for a one-line field. The two ARSENIC TRIOXYD descriptors do carry data-row cells, found by an independent detector rather than by looking where pointed. A **third** case is added: `foscarnet_natri` p698_t0 is a multi-level header labelled `SHAPE_SIMPLE`, in a renal-**dosing** section — harmless this time, but the shape classifier was wrong. Claude agrees with the descriptor embargo and would widen it to all 151 descriptors, since the detector has blind spots and only a visual check of the 71 `Cột:` descriptors would settle it. ## Active ownership - Claude: **2026-08-11** — see `CLAUDE_CLAIM_2026-08-11.md` for the full claim and reasoning. Five production bugs found by driving `https://realvuxbaro.me` (not by reading docs), fixed, deployed and re-verified over 37 live cases: the 25s client abort that was discarding correct grounded answers, availability failures mislabelled as `unsupported_claim`/`incomplete_answer`, the pediatric clarify question re-asking for fields the user had just given, a no-op entailment retry loop, and identical-looking citation chips. Touched `rag/answer.py`, `rag/agent.py`, their tests, `ChatPanel.tsx`, `ChatBubble.tsx`. **Deliberately NOT changed**: the pediatric gate still requires both age and weight, chips are not collapsed, the completeness judge was not relaxed. 230 passed (was 219). Commits `93aa322`, `4e78363`. - **Note on reading status text here, 2026-08-11**: ownership entries in this file and in `CLAUDE_HANDOFF_2026-08-10.md` are written at a point in time and can fall behind — five commits landed on 2026-08-10 between 17:09 and 17:27 after the entries below were written. `git log` is the reliable source for current state; these entries are useful for intent and reasoning. - Codex parallel session: **STOPPED, 2026-08-10** — owner ended the session. (Commits `9c3acd0` … `4438c5f` landed after this line was written.) Left `rag/expansion.py` and `rag/context.py` finished and tested but not wired into any live retrieval path; `rag/fusion.py`/`tests/test_fusion.py` (a third, separate ChatGPT session's work, per Codex's own note above) likewise finished-but-unwired. Claude took over this scope at the owner's explicit direction same session — see `project_production_deployment_live` memory and `docs/progress-log.md` for what got wired in and why. - Claude: **PAUSED END OF SESSION, 2026-08-06** — worked the full correction order from `CODEX_RAG_CODE_REVIEW_2026-08-06.md`. **F-01 through F-07 and F-09 done; F-08 and F-10 done for their core finding, with a named remainder** (see `docs/progress-log.md` top entry, "Status at end of today's session", for the exact scope line per item). Every completed item live-verified against the real running server (not only unit tests) — F-02 scoped down by explicit owner correction (`intent`/`QueryIntent.RECOMMENDATION` is deliberately NOT gated, this product is for doctors/pharmacists, not lay users). **Remaining, next session**: F-08's full request-deadline object (Postgres half already fixed), F-10's comprehensive adversarial battery (one solid end-to-end case now exists and passes, per `RUN_INTEGRATION=1`), a real mg/kg dosing calculator, and `symptom_to_drug` reverse lookup. `apps/ai-service`: **184 passed, 4 skipped** (was 118p/3s at the start of today). **F-03**: `rag/agent.py`'s `RagAgent` (built last session, never constructed/called by anything live — Codex's exact finding) is now built by `bootstrap.py` and called by `routers/rag.py` for both single- and multi-turn requests. The old resolver/routing/conversational stack is NOT deleted yet (still used for autocomplete + the no-generator-configured fallback, still unit-tested) — full removal is gated on F-10's parity suite per Codex's own ask. Drove the real running server (not just unit tests with fakes) and found + fixed two live bugs: `retrieve_framed` had no bare-name/overview case and was sending entire ~29-section monographs as evidence; the new entailment check (F-01) is noisier than one call suggests and needed a same-claim retry. Full detail, including a residual known limitation left deliberately unresolved (owner capped further retry/token spend on one narrow interaction-evidence edge case), in `docs/progress-log.md`. 162 passed, 3 skipped. Claiming: `apps/ai-service/rag/{grounding,answer,prompt,agent,service, policy}.py`, `apps/ai-service/bootstrap.py`, `apps/ai-service/routers/ rag.py`, `apps/ai-service/adapters/prometheus.py`, `apps/ai-service/tests/*` (RAG-answer/generation/retrieval/agent/api tests), `apps/web/app/api/chat/route.ts`. Not touching `ingestion/`, `cli.py`, or anything Codex is mid-investigation on (the monograph-count entry just added to `docs/progress-log.md` — read-only, not editing). - Codex: **done, 2026-08-06** — code-only review of the current RAG runtime and prompt path. Added `coordination/CODEX_RAG_CODE_REVIEW_2026-08-06.md`; no application, ingestion, prompt, test, or infrastructure file was changed. - Claude: **IN PROGRESS, 2026-08-05** — making the live demo path survive a reviewer typing into the UI. Both Codex entries below read *done, 2026-08-04*, so nothing was taken out from under anyone. Claiming: `apps/ai-service/rag/{ports,service}.py`, `apps/ai-service/adapters/embedding.py`, `apps/ai-service/config.py`, `apps/ai-service/tests/*` (additions), `apps/web/**`, `packages/api-client/src/*`. **Not touching** `ingestion/`, `cli.py`, `segment/`, `extract/`, or the two untracked files `ingestion/ingestion/embed/benchmark_local.py` and `ingestion/tests/test_embed_benchmark_local.py`, which are Codex's and still uncommitted. Measured today before editing: ingestion **296 passed**; ai-service **37 passed, 3 skipped**; `duocthu_v1` holds **15,100 points** at 1024-dim Cosine. Bedrock is still closed — verified live today, because the API returned `AccessDeniedException` on `InvokeModel` from a real request. **No cloud call, no spend.** - Codex: **done, 2026-08-04** — `apps/ai-service/` API RAG, Qdrant retrieval adapter, PostgreSQL trace persistence, guardrails and printed-page citations. Claiming `apps/ai-service/{main.py,config.py,adapters/,routers/}`, additions under `apps/ai-service/rag/`, its tests/migrations and dependency declarations. Codex will not edit Claude's `ingestion/load/*`, `ingestion/embed/cache.py`, load/cache tests, or `pyproject.toml` extras. Added verified printed folios to chunk schema v3 and regenerated 15,066 chunks; corpus SHA is `e474c83790b450d3262f532e81abf6526a485e3a98e376413247da23f4619c38`. `chunk_without_printed_page_range = 0`; population tags and `cli embed/load` remain pending. No Bedrock calls, corpus embedding, IAM changes, commit, or push. - Claude: **done, 2026-08-04** — `ingestion/load/` (Qdrant boundary) and `embed/cache.py`, items A2/A4/A5/A6 of `docs/v1-delivery-plan.md` §4.A. Full scope, owner decisions and **four open questions addressed to Codex** are in `CLAUDE_TASK_2026-08-04.md` — read that before touching `cli.py`, the chunk payload, or `segment/`. Claiming: `ingestion/ingestion/load/*` (empty today), `ingestion/ingestion/embed/cache.py`, `ingestion/tests/test_load_*.py`, `ingestion/tests/test_embed_cache.py`, and `ingestion/pyproject.toml` extras only. **Not touching** `segment/`, `extract/`, `validation/`, `entities/`, `apps/ai-service/rag/`, or `cli.py` — all dirty and owned by Codex. The later project-owner instruction keeps runtime provider-agnostic and limits Bedrock to research/benchmarking. The Bedrock IAM policy stays **unapplied**; **no cloud call today**, measured spend **$0**. - Claude: **done, 2026-08-03** — AWS Bedrock embedding setup, items 1-4 of `CLAUDE_TASK.md`. Item 5 (live calls) is blocked on an IAM policy that was drafted but deliberately not applied. Full handoff at the end of `CLAUDE_TASK.md`. Owned and changed: `ingestion/ingestion/embed/*` (all files), `ingestion/tests/test_embed_providers.py`, `infra/aws/iam/*`, `ingestion/pyproject.toml` (extras only). No parser, segmentation, table, formula, chunking or `cli.py` file was touched.