# Disease / Condition → Medication: audit and minimal design Date: 2026-08-11 Scope: the current worktree and the local `duocthu_v1` Qdrant collection. Code and runtime observations take precedence over older ADR/progress-log statements. ## Phase 1 — Current production capability ### Verified request path ```text apps/web/app/api/chat/route.ts -> POST /v1/rag/query -> routers/rag.py::query_rag -> rag/agent.py::RagAgent.handle -> rag/understanding.py::LlmQueryUnderstander.understand -> rag/agent.py::RagAgent._route -> rag/service.py::RetrievalService -> adapters/qdrant.py::QdrantRetriever -> rag/answer.py::GroundedAnswerService.answer_from_result -> rag/grounding.py::verify + LLM entailment check -> citations built from retrieved metadata ``` The web BFF currently always sends `subject_scope="human"` and `intent="fact_lookup"`. The live `RagAgent` deliberately does not trust/use that client intent for routing; its `QueryFrame.turn_type` drives the dispatch. The legacy `QueryRoutingService` remains the no-agent/retrieval-only fallback. ### Query understanding and routing `rag/understanding.py` already has one LLM-driven structured pass. Its current closed turn taxonomy is: ```text drug_attribute, drug_overview, interaction, symptom_to_drug, dosing_calc, smalltalk, out_of_scope ``` It validates drug ids against a deterministic catalog-bounded candidate set and extracts section, population, weight, age, indication, route, clarification state, and a standalone-query rewrite. It does not yet represent a normalized disease, the requested disease↔drug relation, comorbidities, allergies, current medicines, renal/hepatic state, pregnancy/breastfeeding, labs, or a clinical case boundary. `rag/agent.py::RagAgent._route` already dispatches `symptom_to_drug` to `_symptom_to_drug`. However, the repository's most recent measured local run recorded 5/5 disease/symptom queries being misrouted into clarification and never reaching the reverse lookup. This is a query-understanding/routing failure, not an absence of indication data. ### Retrieval actually present - Exact metadata retrieval: `find_by_section(drug_id, section_key)` scrolls the complete section, sorted by `part_index`. - Drug overview: `find_by_drug` scrolls prose for one monograph. - Dense: `search` within one drug and `search_indication` across `section_key=chi_dinh`. - Lexical: `search_lexical` is normalized token-overlap over Qdrant's text index; it is BM25-style but not a true sparse-vector/BM25 ranking. - Hybrid: `rag/fusion.py::reciprocal_rank_fusion` exists and is unit-tested, but it is not wired into the live retrieval service. - Reranker: Cohere rerank is configurable and used for overview/similarity. It is off by default and the current indication route does not call it. - Reverse indication: `find_by_indication` performs contiguous normalized phrase matching over prose `chi_dinh` chunks; `search_indication` is the dense fallback. Both exclude contraindication, ADR, precaution, and interaction sections. The current reverse lookup deduplicates to one hit per drug inside the adapter, but stops at the first matching chunk and caps before any entity-level reranking. Qdrant scroll order is not a clinical ranking, so the current top-N is arbitrary among exact matches. It also retains only one evidence chunk rather than an explicit drug-level aggregate. ### Qdrant and chunk schema The local runtime collection was queried directly during this audit: - collection `duocthu_v1`: green, 15,100 points, cosine vectors, 1,024 dimensions; - payload indexes: `chunk_id`, `drug_id`, `section_key`, `atc_codes`, `chunk_kind`, `has_quarantined_content`, plus multilingual `text` index; - a live `chi_dinh` point contains `drug_id`, `drug_name`, `section_key`, `section_display_name`, `text`, `source_text`, physical and printed page ranges, part index/count, ATC code, attachments, and quarantine flag. `ingestion/ingestion/chunk/models.py::Chunk` and `ingestion/ingestion/load/models.py` confirm those fields. `parent_id` and explicit `source_refs` are supported by the AI-service retrieval model/adapter, but the current ingestion `Chunk` contract does not emit `parent_id`; the live sample also has no parent id. Parent hydration is therefore reusable compatibility machinery, not an active parent-child hierarchy in the current v4 corpus. Provenance is at chunk page-range/attachment-region precision; there is no character-offset span. ### Grounding, claims, citations, and generation - `rag/prompt.py::ANSWER_SCHEMA` requires structured claims with citation indices. - `rag/grounding.py::verify` rejects invalid citations, uncited claims, and numbers absent from the specifically cited evidence. - `GroundedAnswerService` additionally runs an LLM entailment/completeness check. - API citations are built from retrieved `SourceRef`, never model-authored prose. - `list_mode` exists for reverse indication and asks generation to enumerate the retrieved drugs without calling any one first-line/preferred. There is no deterministic candidate-set field on generated claims today. A model that names an extra drug should be rejected by semantic entailment, but there is no direct `generated_drugs - retrieved_drugs` set check. Citation responses expose chunk id and page data; drug/section labels are currently reconstructed in the web BFF by splitting `chunk_id`, rather than carried explicitly as provenance. ### Conversation state Raw conversation lines are persisted by `adapters/postgres.py::PostgresConversationStore`. `RagAgent._last_frame` is the only normalized state and is in-process only. `_merge_with_prior_frame` only has a code-level merge backstop for an open clarification. Ordinary multi-turn patient facts are otherwise re-derived by the LLM from raw history and can be lost; no explicit new-patient/case boundary exists. ### Tests and evaluations Baseline command run before feature changes: ```text python -m pytest -q --ignore=tests/test_api.py --ignore=tests/test_live_datastores.py 243 passed in 2.08s ``` Existing tests cover the primitive reverse indication route, its section filter, one-hit-per-drug behavior, no-result abstention, list-mode prompting, grounding, and raw history isolation. They do not cover disease normalization/ambiguity, relation confusion, patient context, second-stage safety retrieval, candidate assessment, or unsupported-drug rate. `rag/evaluation.py`/`run_eval.py` are drug-first and do not calculate the requested condition-to-drug metrics. ## Phase 2 — Gap analysis | Requirement | State | Existing implementation | Minimal proposed change | |---|---|---|---| | Disease intent | Partial | `symptom_to_drug` frame and agent branch | Rename/accept `condition_to_drug`; retain old value as compatibility alias; add requested relation | | Drug→condition / dose / contraindication / interaction distinction | Partial | turn type + `attribute` | Add explicit `drug_to_condition` and relation-safe reverse categories without replacing section taxonomy | | Condition extraction/normalization | Missing | free-text `indication` only | Add `ConditionQuery`; deterministic conservative alias normalization plus LLM structured output; preserve original | | Ambiguity | Partial | generic `needs_clarify` | Add condition ambiguity fields and deterministic guard for known broad category-only queries | | Indication-only reverse retrieval | Exists | both indication methods filter `chi_dinh` | Keep filter; add ranked candidate pool and aggregate at drug level | | Drug-level aggregation/rerank | Partial | one first hit per drug | Aggregate all candidate hits by `drug_id`, then rerank/cap entities, never count chunks as votes | | Dense/sparse/hybrid | Partial | dense + lexical; RRF not live | Reuse exact lexical-first and dense fallback initially; keep fusion seam, avoid unmeasured full-stack rewrite | | Patient context | Missing | population/age/weight only | Add structured `PatientContext`, optional and field-preserving | | Comorbidities/current medicines/allergy | Missing | direct interaction supports 2 named drugs | Make first-class context and trigger targeted second-stage retrieval | | Renal/hepatic/pregnancy/age | Partial | sections exist for drug-centric queries | Select only relevant safety sections for top candidates, using lexical seed then whole-section hydration | | Candidate assessment | Missing | raw evidence pool only | Add evidence-only `MedicationCandidateAssessment` grouped by drug and safety facet/status | | Candidate-set hallucination guard | Partial | grounding + entailment | Require candidate `drug_id` on list-mode claims and validate it/cited evidence deterministically | | Provenance | Partial | pages + chunk id | Carry drug name, section key/title, and corpus source explicitly through Evidence/Citation/API | | Structured conversation state | Partial | raw Postgres history + in-memory last frame | Merge `PatientContext` only on explicit case continuation; reset on new case/topic; keep raw history fallback | | Guideline distinction | Missing in prompt | corpus is Part 2 monographs only | Add prompt contract: indication evidence is not first-line/preferred/treatment-of-choice evidence | | Metrics/eval | Missing for this feature | generic recall/resolution eval | Add deterministic feature eval cases/metrics including unsupported-drug rate and section/relation correctness | ## Phase 3 — Minimal architecture ```text LlmQueryUnderstander -> QueryFrame(condition + relation + optional PatientContext + case action) -> deterministic condition normalization / ambiguity backstop -> RagAgent condition_to_drug route -> RetrievalService.retrieve_by_indication -> chi_dinh lexical candidates (dense only as fallback) -> group by drug_id -> entity-level rerank/cap -> indication evidence -> if patient context exists: top candidates × context -> lexical selection among relevant safety facets -> hydrate only selected whole sections -> MedicationCandidateAssessment per drug -> candidate-set validator -> existing structured generation + grounding + entailment -> explicit drug/section/page/source citations ``` ### Files to modify - `apps/ai-service/rag/understanding.py`: frame/schema/prompt/parser and bounded conversation-state merge. - `apps/ai-service/rag/agent.py`: relation-safe dispatch, ambiguity response, optional patient-specific second stage, candidate assessments. - `apps/ai-service/rag/service.py`: drug-level indication aggregation/rerank and targeted patient-safety retrieval. - `apps/ai-service/adapters/qdrant.py`: return a wider, scored indication candidate pool without first-match/scroll-order ranking. - `apps/ai-service/rag/models.py`, `rag/answer.py`, `rag/prompt.py`: evidence provenance and deterministic candidate-set claim validation. - `apps/ai-service/routers/rag.py`, `apps/web/app/api/chat/route.ts`, shared types: expose explicit provenance without parsing chunk ids. - instrumentation and tests/evals for new routes and metrics. ### File to create - `apps/ai-service/rag/clinical.py`: small domain-only schemas and conservative condition/context normalization. It contains no disease→drug knowledge. - focused tests/eval fixture for condition-to-drug and patient safety. ### Explicit non-goals No ingestion rewrite, knowledge graph, internet access, guideline subsystem, autonomous diagnosis, agent loop, new service, or hard-coded disease→drug map. The Part 2 monograph corpus can prove an indication and drug-specific safety text; it cannot by itself prove first-line/preferred regimens.