Files
duocthu/coordination/CLAUDE_HANDOFF_2026-08-10.md
T

7.8 KiB

Claude handoff — 2026-08-10, in case of context/token cutoff

Update 2026-08-11 — two items below have moved on since this was written. Checked against the code and against live production.

  1. The structured-claims refactor described below as in progress shipped the same day (dfdbf52, then 9c3acd0); the conversion is complete and the suite is at 230 passing.
  2. Entailment majority vote (2-of-3) is no longer in the code. df55af4 introduced it and 9c3acd0 replaced it with a single pass (_ENTAILMENT_MAX_ATTEMPTS = 1); _verify_entailment's docstring gives the reasoning — repeating an identical temperature-0 prompt is a correlated retry rather than an independent vote.

Task #4 (real BM25 via Qdrant native sparse vectors) further down is still accurate and still not started. For current state see coordination/CLAUDE_CLAIM_2026-08-11.md and the 2026-08-11 entry in docs/progress-log.md.

Read this before touching apps/ai-service/rag/answer.py, rag/prompt.py, adapters/bedrock_claude.py, or any test file under apps/ai-service/tests/ that references the answer-generation schema. A structured-claims refactor is IN PROGRESS AND NOT YET FULLY GREEN.

What's done and committed (pushed, deployed, live-verified)

  • Production live at https://realvuxbaro.me (EC2 + Docker Compose + Caddy SSL + GitHub Actions CI/CD). See project_production_deployment_live memory (Claude's own memory dir, not readable by Codex — this file is the Codex-readable version of the relevant parts).
  • Section-neighbour lexical pooling (rag/service.py::_pooled_neighbour_hits, adapters/qdrant.py::search_lexical) — fixes 2 of 3 persistent audit abstains (Aspirin+loét dạ dày, Vancomycin rapid-infusion). Committed, deployed, live-verified.
  • Token-budget packing wired into the overview/rerank fallback (rag/context.py::pack_evidence, was dead code, now used in rag/service.py). Committed, deployed.
  • Entailment verification changed from "accept on any single True out of 3" to MAJORITY VOTE (2-of-3) — rag/answer.py::_verify_entailment. Committed, deployed, live-verified no regression. Full reasoning (measured math, adversarial spot-check results) is in the commit message and the function's own docstring — read that before changing it again.
  • Composer.tsx autocomplete: fixed matching the whole sentence instead of the last word being typed, added a race guard on the debounced fetch. Committed, deployed.

What's IN PROGRESS, NOT committed, NOT deployed (as of this handoff)

Structured-claims output (rag/prompt.py ANSWER_SCHEMA changed from {answer: string, evidence_sufficient, clarifying_question} to {claims: [{text, citations}], evidence_sufficient, clarifying_question}; rag/answer.py parses claims and deterministically assembles the display string via _assemble_answer — same text [n] format the frontend already renders, so grounding.verify and the frontend need NO changes).

Modified, uncommitted: apps/ai-service/adapters/bedrock_claude.py, apps/ai-service/rag/answer.py, apps/ai-service/rag/prompt.py, apps/ai-service/tests/test_grounded_generation.py (this one IS finished — 25/25 pass).

Still broken as of this handoff (python -m pytest -q in apps/ai-service, 4 failures, 208 passed):

  • tests/test_agent.py::test_a_generous_budget_does_not_change_normal_behaviour — 1 fake generator payload at ~line 493 still uses the old {"answer": "...", ...} shape, needs converting to {"claims": [{"text": "...", "citations": [...]}], ...} (see test_grounded_generation.py's already-converted tests for the pattern).
  • tests/test_citation_and_intro.py — 3 failures, same root cause (old-shape fake payloads not yet converted): test_only_cited_sources_are_returned, test_sufficiency_check_outage_fails_open_to_generation_not_abstain, test_list_mode_skips_the_sufficiency_clarify.
  • Have NOT yet checked tests/test_live_datastores.py or tests/test_bedrock_converse.py for old-shape payloads — grep for "answer": across apps/ai-service to find any remaining.

Conversion pattern (mechanical, already applied ~15 times in test_grounded_generation.py):

# OLD:
{"answer": "Người lớn uống 500 mg [1].", "evidence_sufficient": True}
# NEW:
{"claims": [{"text": "Người lớn uống 500 mg", "citations": [1]}], "evidence_sufficient": True}

For evidence_sufficient: False payloads, old {"answer": "...", ...} becomes {"claims": [], "evidence_sufficient": False}_attempt_generation now requires claims to be a present list even when insufficient, or it's misclassified as malformed_output instead of evidence_insufficient.

After all tests are green: run full local live-verify (restart ai-service, hit /v1/rag/query for a few real drugs, confirm answers still read normally and citations still work) before committing. Then commit, push, let CI/CD deploy, live-verify on https://realvuxbaro.me too.

Task #4 — real BM25 via Qdrant native sparse vectors (NOT STARTED)

Owner gave a detailed 12-point spec, paraphrased:

  1. Keep deterministic section routing as the fast path, unchanged.
  2. Only for free-form / no-section-matched / low-confidence queries: run dense + Qdrant native sparse (BM25) in parallel, both filtered to the resolved drug_id, fuse via RRF, feed the existing reranker, then token-budget pack.
  3. Never run hybrid for a query the deterministic route already answered. 4/5. Dense failure falls back to sparse-only; sparse failure falls back to dense-only.
  4. No hardcoding to specific drugs/sections/questions.
  5. Proper Vietnamese tokenization — do not blindly reuse English stemming/stopword defaults.
  6. Trace must record dense hits, sparse hits, RRF score, reranker score, route taken, and per-step latency.
  7. Run an ablation on the golden set: dense-only / sparse-only / dense+sparse RRF / dense+sparse+reranker.
  8. Report Recall@K, MRR/nDCG, citation correctness, latency per variant.
  9. Verify Qdrant server AND client library versions support native sparse vectors/Query API BEFORE designing anything further.
  10. Never touch/lose the existing duocthu_v1 collection — new collection/version with a rollback path. Do not deploy before testing and reporting results.

Version check already done (2026-08-10): Qdrant SERVER is 1.18.3 (full native sparse-vector + Query API/RRF support). Installed qdrant-client PYTHON package is 1.7.0 (has basic SparseVector model but NOT the newer Query API/FusionQuery — that needs a client upgrade to roughly 1.10+). pyproject.toml's qdrant-client>=1.7,<2 already permits upgrading within range, no constraint change needed. Nothing sparse-related has been built yet — no sparse index, no corpus indexing, no real sparse query has run. Do not report "BM25 exists" until all of that is actually done and verified — explicit owner instruction, PostgreSQL ts_rank/tsvector does NOT count as BM25 (different formula, no term-frequency saturation / doc-length norm).

Hard constraint, repeat for emphasis

No commit message, code comment, memory file, or project doc may reference the competitor pipeline material the owner showed via screenshots earlier this session, or say anything is "based on"/"dựa theo" it. Justify every design choice from this codebase's own live findings or public, generically-cited RAG research only. Already checked clean through commit df55af4; keep checking every future commit before pushing.

Coordination note

Codex's session was explicitly stopped by the owner this same day; Claude took over rag/** scope at the owner's direction (see coordination/README.md's "Active ownership" section, already updated). If Codex resumes, read this file and coordination/README.md first.