129 lines
6.9 KiB
Markdown
129 lines
6.9 KiB
Markdown
# Claude handoff — 2026-08-10, in case of context/token cutoff
|
|
|
|
Read this before touching `apps/ai-service/rag/answer.py`, `rag/prompt.py`,
|
|
`adapters/bedrock_claude.py`, or any test file under `apps/ai-service/tests/`
|
|
that references the answer-generation schema. A structured-claims refactor
|
|
is **IN PROGRESS AND NOT YET FULLY GREEN**.
|
|
|
|
## What's done and committed (pushed, deployed, live-verified)
|
|
|
|
- Production live at `https://realvuxbaro.me` (EC2 + Docker Compose + Caddy
|
|
SSL + GitHub Actions CI/CD). See `project_production_deployment_live`
|
|
memory (Claude's own memory dir, not readable by Codex — this file is the
|
|
Codex-readable version of the relevant parts).
|
|
- Section-neighbour lexical pooling (`rag/service.py::_pooled_neighbour_hits`,
|
|
`adapters/qdrant.py::search_lexical`) — fixes 2 of 3 persistent audit
|
|
abstains (Aspirin+loét dạ dày, Vancomycin rapid-infusion). Committed,
|
|
deployed, live-verified.
|
|
- Token-budget packing wired into the overview/rerank fallback
|
|
(`rag/context.py::pack_evidence`, was dead code, now used in
|
|
`rag/service.py`). Committed, deployed.
|
|
- Entailment verification changed from "accept on any single True out of 3"
|
|
to MAJORITY VOTE (2-of-3) — `rag/answer.py::_verify_entailment`. Committed,
|
|
deployed, live-verified no regression. Full reasoning (measured math,
|
|
adversarial spot-check results) is in the commit message and the
|
|
function's own docstring — read that before changing it again.
|
|
- `Composer.tsx` autocomplete: fixed matching the whole sentence instead of
|
|
the last word being typed, added a race guard on the debounced fetch.
|
|
Committed, deployed.
|
|
|
|
## What's IN PROGRESS, NOT committed, NOT deployed (as of this handoff)
|
|
|
|
**Structured-claims output** (`rag/prompt.py` ANSWER_SCHEMA changed from
|
|
`{answer: string, evidence_sufficient, clarifying_question}` to
|
|
`{claims: [{text, citations}], evidence_sufficient, clarifying_question}`;
|
|
`rag/answer.py` parses `claims` and deterministically assembles the display
|
|
string via `_assemble_answer` — same `text [n]` format the frontend already
|
|
renders, so `grounding.verify` and the frontend need NO changes).
|
|
|
|
**Modified, uncommitted**: `apps/ai-service/adapters/bedrock_claude.py`,
|
|
`apps/ai-service/rag/answer.py`, `apps/ai-service/rag/prompt.py`,
|
|
`apps/ai-service/tests/test_grounded_generation.py` (this one IS finished —
|
|
25/25 pass).
|
|
|
|
**Still broken as of this handoff** (`python -m pytest -q` in
|
|
`apps/ai-service`, 4 failures, 208 passed):
|
|
- `tests/test_agent.py::test_a_generous_budget_does_not_change_normal_behaviour`
|
|
— 1 fake generator payload at ~line 493 still uses the old
|
|
`{"answer": "...", ...}` shape, needs converting to
|
|
`{"claims": [{"text": "...", "citations": [...]}], ...}` (see
|
|
`test_grounded_generation.py`'s already-converted tests for the pattern).
|
|
- `tests/test_citation_and_intro.py` — 3 failures, same root cause (old-shape
|
|
fake payloads not yet converted): `test_only_cited_sources_are_returned`,
|
|
`test_sufficiency_check_outage_fails_open_to_generation_not_abstain`,
|
|
`test_list_mode_skips_the_sufficiency_clarify`.
|
|
- Have NOT yet checked `tests/test_live_datastores.py` or
|
|
`tests/test_bedrock_converse.py` for old-shape payloads — grep for
|
|
`"answer":` across `apps/ai-service` to find any remaining.
|
|
|
|
**Conversion pattern** (mechanical, already applied ~15 times in
|
|
`test_grounded_generation.py`):
|
|
```python
|
|
# OLD:
|
|
{"answer": "Người lớn uống 500 mg [1].", "evidence_sufficient": True}
|
|
# NEW:
|
|
{"claims": [{"text": "Người lớn uống 500 mg", "citations": [1]}], "evidence_sufficient": True}
|
|
```
|
|
For `evidence_sufficient: False` payloads, old `{"answer": "...", ...}`
|
|
becomes `{"claims": [], "evidence_sufficient": False}` — `_attempt_generation`
|
|
now requires `claims` to be a present list even when insufficient, or it's
|
|
misclassified as `malformed_output` instead of `evidence_insufficient`.
|
|
|
|
**After all tests are green**: run full local live-verify (restart
|
|
ai-service, hit `/v1/rag/query` for a few real drugs, confirm answers still
|
|
read normally and citations still work) before committing. Then commit,
|
|
push, let CI/CD deploy, live-verify on `https://realvuxbaro.me` too.
|
|
|
|
## Task #4 — real BM25 via Qdrant native sparse vectors (NOT STARTED)
|
|
|
|
Owner gave a detailed 12-point spec, paraphrased:
|
|
1. Keep deterministic section routing as the fast path, unchanged.
|
|
2. Only for free-form / no-section-matched / low-confidence queries: run
|
|
dense + Qdrant native sparse (BM25) in parallel, both filtered to the
|
|
resolved `drug_id`, fuse via RRF, feed the existing reranker, then
|
|
token-budget pack.
|
|
3. Never run hybrid for a query the deterministic route already answered.
|
|
4/5. Dense failure falls back to sparse-only; sparse failure falls back to
|
|
dense-only.
|
|
6. No hardcoding to specific drugs/sections/questions.
|
|
7. Proper Vietnamese tokenization — do not blindly reuse English
|
|
stemming/stopword defaults.
|
|
8. Trace must record dense hits, sparse hits, RRF score, reranker score,
|
|
route taken, and per-step latency.
|
|
9. Run an ablation on the golden set: dense-only / sparse-only /
|
|
dense+sparse RRF / dense+sparse+reranker.
|
|
10. Report Recall@K, MRR/nDCG, citation correctness, latency per variant.
|
|
11. **Verify Qdrant server AND client library versions support native sparse
|
|
vectors/Query API BEFORE designing anything further.**
|
|
12. Never touch/lose the existing `duocthu_v1` collection — new
|
|
collection/version with a rollback path. Do not deploy before testing
|
|
and reporting results.
|
|
|
|
**Version check already done** (2026-08-10): Qdrant SERVER is 1.18.3 (full
|
|
native sparse-vector + Query API/RRF support). Installed `qdrant-client`
|
|
PYTHON package is 1.7.0 (has basic `SparseVector` model but NOT the newer
|
|
Query API/`FusionQuery` — that needs a client upgrade to roughly 1.10+).
|
|
`pyproject.toml`'s `qdrant-client>=1.7,<2` already permits upgrading within
|
|
range, no constraint change needed. **Nothing sparse-related has been built
|
|
yet** — no sparse index, no corpus indexing, no real sparse query has run.
|
|
Do not report "BM25 exists" until all of that is actually done and verified
|
|
— explicit owner instruction, PostgreSQL ts_rank/tsvector does NOT count as
|
|
BM25 (different formula, no term-frequency saturation / doc-length norm).
|
|
|
|
## Hard constraint, repeat for emphasis
|
|
|
|
**No commit message, code comment, memory file, or project doc may
|
|
reference the competitor pipeline material the owner showed via
|
|
screenshots earlier this session, or say anything is "based on"/"dựa
|
|
theo" it.** Justify every design choice from this codebase's own live
|
|
findings or public, generically-cited RAG research only. Already checked
|
|
clean through commit `df55af4`; keep checking every future commit before
|
|
pushing.
|
|
|
|
## Coordination note
|
|
|
|
Codex's session was explicitly stopped by the owner this same day; Claude
|
|
took over `rag/**` scope at the owner's direction (see
|
|
`coordination/README.md`'s "Active ownership" section, already updated).
|
|
If Codex resumes, read this file and `coordination/README.md` first.
|