Files
duocthu/coordination/CLAUDE_HANDOFF_2026-08-10.md
T

129 lines
6.9 KiB
Markdown

# Claude handoff — 2026-08-10, in case of context/token cutoff
Read this before touching `apps/ai-service/rag/answer.py`, `rag/prompt.py`,
`adapters/bedrock_claude.py`, or any test file under `apps/ai-service/tests/`
that references the answer-generation schema. A structured-claims refactor
is **IN PROGRESS AND NOT YET FULLY GREEN**.
## What's done and committed (pushed, deployed, live-verified)
- Production live at `https://realvuxbaro.me` (EC2 + Docker Compose + Caddy
SSL + GitHub Actions CI/CD). See `project_production_deployment_live`
memory (Claude's own memory dir, not readable by Codex — this file is the
Codex-readable version of the relevant parts).
- Section-neighbour lexical pooling (`rag/service.py::_pooled_neighbour_hits`,
`adapters/qdrant.py::search_lexical`) — fixes 2 of 3 persistent audit
abstains (Aspirin+loét dạ dày, Vancomycin rapid-infusion). Committed,
deployed, live-verified.
- Token-budget packing wired into the overview/rerank fallback
(`rag/context.py::pack_evidence`, was dead code, now used in
`rag/service.py`). Committed, deployed.
- Entailment verification changed from "accept on any single True out of 3"
to MAJORITY VOTE (2-of-3) — `rag/answer.py::_verify_entailment`. Committed,
deployed, live-verified no regression. Full reasoning (measured math,
adversarial spot-check results) is in the commit message and the
function's own docstring — read that before changing it again.
- `Composer.tsx` autocomplete: fixed matching the whole sentence instead of
the last word being typed, added a race guard on the debounced fetch.
Committed, deployed.
## What's IN PROGRESS, NOT committed, NOT deployed (as of this handoff)
**Structured-claims output** (`rag/prompt.py` ANSWER_SCHEMA changed from
`{answer: string, evidence_sufficient, clarifying_question}` to
`{claims: [{text, citations}], evidence_sufficient, clarifying_question}`;
`rag/answer.py` parses `claims` and deterministically assembles the display
string via `_assemble_answer` — same `text [n]` format the frontend already
renders, so `grounding.verify` and the frontend need NO changes).
**Modified, uncommitted**: `apps/ai-service/adapters/bedrock_claude.py`,
`apps/ai-service/rag/answer.py`, `apps/ai-service/rag/prompt.py`,
`apps/ai-service/tests/test_grounded_generation.py` (this one IS finished —
25/25 pass).
**Still broken as of this handoff** (`python -m pytest -q` in
`apps/ai-service`, 4 failures, 208 passed):
- `tests/test_agent.py::test_a_generous_budget_does_not_change_normal_behaviour`
— 1 fake generator payload at ~line 493 still uses the old
`{"answer": "...", ...}` shape, needs converting to
`{"claims": [{"text": "...", "citations": [...]}], ...}` (see
`test_grounded_generation.py`'s already-converted tests for the pattern).
- `tests/test_citation_and_intro.py` — 3 failures, same root cause (old-shape
fake payloads not yet converted): `test_only_cited_sources_are_returned`,
`test_sufficiency_check_outage_fails_open_to_generation_not_abstain`,
`test_list_mode_skips_the_sufficiency_clarify`.
- Have NOT yet checked `tests/test_live_datastores.py` or
`tests/test_bedrock_converse.py` for old-shape payloads — grep for
`"answer":` across `apps/ai-service` to find any remaining.
**Conversion pattern** (mechanical, already applied ~15 times in
`test_grounded_generation.py`):
```python
# OLD:
{"answer": "Người lớn uống 500 mg [1].", "evidence_sufficient": True}
# NEW:
{"claims": [{"text": "Người lớn uống 500 mg", "citations": [1]}], "evidence_sufficient": True}
```
For `evidence_sufficient: False` payloads, old `{"answer": "...", ...}`
becomes `{"claims": [], "evidence_sufficient": False}``_attempt_generation`
now requires `claims` to be a present list even when insufficient, or it's
misclassified as `malformed_output` instead of `evidence_insufficient`.
**After all tests are green**: run full local live-verify (restart
ai-service, hit `/v1/rag/query` for a few real drugs, confirm answers still
read normally and citations still work) before committing. Then commit,
push, let CI/CD deploy, live-verify on `https://realvuxbaro.me` too.
## Task #4 — real BM25 via Qdrant native sparse vectors (NOT STARTED)
Owner gave a detailed 12-point spec, paraphrased:
1. Keep deterministic section routing as the fast path, unchanged.
2. Only for free-form / no-section-matched / low-confidence queries: run
dense + Qdrant native sparse (BM25) in parallel, both filtered to the
resolved `drug_id`, fuse via RRF, feed the existing reranker, then
token-budget pack.
3. Never run hybrid for a query the deterministic route already answered.
4/5. Dense failure falls back to sparse-only; sparse failure falls back to
dense-only.
6. No hardcoding to specific drugs/sections/questions.
7. Proper Vietnamese tokenization — do not blindly reuse English
stemming/stopword defaults.
8. Trace must record dense hits, sparse hits, RRF score, reranker score,
route taken, and per-step latency.
9. Run an ablation on the golden set: dense-only / sparse-only /
dense+sparse RRF / dense+sparse+reranker.
10. Report Recall@K, MRR/nDCG, citation correctness, latency per variant.
11. **Verify Qdrant server AND client library versions support native sparse
vectors/Query API BEFORE designing anything further.**
12. Never touch/lose the existing `duocthu_v1` collection — new
collection/version with a rollback path. Do not deploy before testing
and reporting results.
**Version check already done** (2026-08-10): Qdrant SERVER is 1.18.3 (full
native sparse-vector + Query API/RRF support). Installed `qdrant-client`
PYTHON package is 1.7.0 (has basic `SparseVector` model but NOT the newer
Query API/`FusionQuery` — that needs a client upgrade to roughly 1.10+).
`pyproject.toml`'s `qdrant-client>=1.7,<2` already permits upgrading within
range, no constraint change needed. **Nothing sparse-related has been built
yet** — no sparse index, no corpus indexing, no real sparse query has run.
Do not report "BM25 exists" until all of that is actually done and verified
— explicit owner instruction, PostgreSQL ts_rank/tsvector does NOT count as
BM25 (different formula, no term-frequency saturation / doc-length norm).
## Hard constraint, repeat for emphasis
**No commit message, code comment, memory file, or project doc may
reference the competitor pipeline material the owner showed via
screenshots earlier this session, or say anything is "based on"/"dựa
theo" it.** Justify every design choice from this codebase's own live
findings or public, generically-cited RAG research only. Already checked
clean through commit `df55af4`; keep checking every future commit before
pushing.
## Coordination note
Codex's session was explicitly stopped by the owner this same day; Claude
took over `rag/**` scope at the owner's direction (see
`coordination/README.md`'s "Active ownership" section, already updated).
If Codex resumes, read this file and `coordination/README.md` first.