Add read-only production runtime audit

This commit is contained in:
2026-08-17 11:17:40 +07:00
parent 057d4ed9dc
commit a1de4715a4
106 changed files with 6869 additions and 1782 deletions
+216
View File
@@ -0,0 +1,216 @@
# 09 — Retrieval pipeline
Implementation: `apps/ai-service/rag/service.py` (`RetrievalService`, 741 lines),
`apps/ai-service/adapters/qdrant.py` (501 lines), `rag/sections.py`,
`rag/context.py`.
Tests: `tests/test_retrieval_service.py`, `tests/test_section_routing.py`,
`tests/test_qdrant_adapter.py`, `tests/test_rerank_overview.py`,
`tests/test_section_order.py`.
## What retrieval is here
**Similarity is the fallback, not the default.** Measured 2026-08-04: letting
vector similarity choose the section gives hit@1 **0.544** overall and **0.05**
on `chong_chi_dinh`, because `duoc_ly_va_co_che_tac_dung` is the largest section
and sits close to almost any question about the drug. When the question names
the section it wants, a payload filter answers it exactly.
That single measurement is the reason the architecture looks the way it does.
## Retrieval routes
```mermaid
flowchart TD
IN["retrieve_framed(drug_id, section_key, query, is_overview)"]
S{section_key given?}
SEC["find_by_section(drug_id, section_key)<br/>Qdrant scroll, payload filter, NO vector<br/>score = 1.0 by construction"]
POOL["_pooled_neighbour_hits<br/>only when section == than_trong"]
OV["find_by_drug(drug_id)<br/>every prose section, book order"]
ISOV{is_overview?}
INTRO["keep INTRO_SECTIONS only:<br/>ten_chung_quoc_te, loai_thuoc,<br/>chi_dinh, duoc_ly_va_co_che_tac_dung"]
RR["_rerank(query, hits)<br/>Cohere rerank-v3.5, top_k=6, fail-open"]
PACK["pack_evidence(max_tokens=6000)"]
DEC["_decide(evidence)"]
IN --> S
S -->|yes| SEC --> POOL --> DEC
S -->|no| OV
OV -->|None| AB["ABSTAIN insufficient_retrieval_score"]
OV --> ISOV
ISOV -->|yes| INTRO --> DEC
ISOV -->|no| RR --> PACK --> DEC
```
### 1. Section route (the primary path)
`find_by_section` is a **`scroll`, not a `search`** — it must not be a top-k.
Paging continues until the offset is exhausted, because Qdrant's default page is
256 and a long section silently truncated would read as a complete answer.
Results are re-sorted by `part_index` (see
[07-indexing-and-storage.md](07-indexing-and-storage.md) for why). No evidence
limit is applied — the whole section is the answer, and a truncated list of
contraindications reads as a complete one.
Score is `1.0` because the match is exact by construction. It is **not** a
similarity and must not be compared to one.
### 2. Bounded cross-section pooling
`_pooled_neighbour_hits` uses `search_lexical` to find another section of the
*same drug* whose text matches the query strongly. It exists for one measured
case: a `thận trọng` question about a specific condition (loét dạ dày) whose
real answer was filed only under `chống chỉ định`.
It is deliberately narrow:
```python
_LEXICAL_POOL_ENABLED_SECTIONS = {"than_trong"} # only this route
_LEXICAL_POOL_EXCLUDED_SECTIONS = {"duoc_ly_va_co_che_tac_dung"} # the known attractor
_LEXICAL_POOL_MIN_SCORE = 5.0
_MAX_LEXICAL_POOLED_SECTIONS = 2
```
The excluded section is excluded outright rather than by score margin: on the
exact query that motivated the mechanism, the true positive scored 7 matched
terms and that attractor scored 6 — too close for a threshold to separate.
Applying pooling to every section leaked a lexically-overlapping interaction
section into a dosage answer, so it stayed opt-in.
### 3. Drug overview (a bare drug name)
`find_by_drug` scrolls every **prose** chunk of the drug (block descriptors stay
out of a text answer), orders by `SECTION_ORDER` then `part_index`, and prefixes
each section's first chunk with `【display name】`.
For `turn_type == "drug_overview"` only the four `INTRO_SECTIONS` are kept.
Without that split a bare drug name sent the entire ~29-section monograph as
evidence for every generation call — wrong retrieval, and an answer long enough
to intermittently fail generation outright.
### 4. Free-form question about a resolved drug
The full monograph is reranked to `rerank_top_k=6`, then packed to a **token**
budget rather than a flat count:
```python
max_context_tokens = 6000 # pack_evidence, rag/context.py
```
`pack_evidence` packs whole blocks in retrieval order and never truncates
clinical text; anything that does not fit is recorded in
`omitted_evidence_ids`. The cap is applied even when rerank is disabled or
fails open — an ordering aid must not also remove the size bound.
### 5. Reverse lookup: condition/indication → drugs
`retrieve_by_indication` is a two-stage lookup, keyword first:
1. **`find_by_indication`** — scroll every `chi_dinh` prose chunk and require the
normalized indication to appear as a **contiguous, word-boundary-anchored
phrase**. Not a substring (false positives after diacritic stripping), and
explicitly not a token-subset match: a nonsense phrase built from common
filler words previously false-positived against real `chi_dinh` text and
reached generation before being caught.
Score rewards an early, concise mention:
`1 + 1/(1+position) + 1/(1 + words/40)`.
2. **`search_indication`** — dense fallback, tried only when the keyword pass
finds nothing, filtered to `section_key=chi_dinh` and `chunk_kind=prose`.
**This is the only place in the live path where dense vector search is
actually used** (ADR 0008). A weak top score (`< evidence_minimum_score`)
discards the hits, because dense search always returns its nearest
neighbours — a made-up phrase still got 8 unrelated "matches" live.
The adapter returns a ranked **chunk** pool; `_rank_indication_drugs` groups by
`drug_id`, takes the **max** score per drug (never a sum or count, so a drug with
more chunks does not win), optionally reranks the groups, and the service caps
at 8 drugs × 2 evidence chunks.
### 6. Patient-specific safety evidence (stage 2)
`assess_patient_candidates` / `retrieve_patient_drug_context` never create
candidates. For each already-indicated drug they run separate, relation-specific
lexical searches:
| Facet | Query source | Sections searched |
|---|---|---|
| interaction | `patient.interaction_query()` | `tuong_tac_thuoc` |
| warnings | `patient.warning_query()` | `chong_chi_dinh`, `than_trong` (requires a clinical-anchor match) |
| dosage context | `patient.dosage_context_query()` | `lieu_luong_va_cach_dung` (requires a clinical-anchor match) |
| pregnancy / breastfeeding | direct section route | `thoi_ky_mang_thai`, `thoi_ky_cho_con_bu` |
Keeping the queries separate is the point: a current medicine may select an
interaction chunk only when *that medicine* matches inside the interaction
section — CKD or age terms from another facet cannot make an unrelated
interaction look supported. `_patient_context_matches` requires a real clinical
anchor rather than overlap on generic words like `chức năng`.
Absence of a hit is recorded as `CandidateStatus.INSUFFICIENT_EVIDENCE` — never
as "safe".
## The evidence decision — `_decide`
```python
if not evidence: ABSTAIN "parent_hydration_failed"
if any(not item.source_refs for item in evidence): ABSTAIN "missing_provenance"
if any(item.requires_visual_check ...): VERIFY_PDF "visual_verification_required"
else: ANSWERABLE "grounded_evidence_available"
```
`decide()` is exposed publicly so a caller assembling its own pool across several
retrieve calls — `RagAgent._interaction` — gets the same quarantine and
provenance policy. Bypassing it is precisely how the interaction path once
silently dropped a quarantined drug's evidence instead of surfacing `VERIFY_PDF`.
`requires_visual_check` is read from the payload as
`requires_visual_check OR has_quarantined_content`.
## Policy constants — `EvidencePolicy`
| Setting | Default | Applies to |
|---|---|---|
| `minimum_score` | 0.12 (`EVIDENCE_MINIMUM_SCORE`) | dense routes only |
| `candidate_limit` | 5 | `retrieve()`'s dense search |
| `evidence_limit` | 3 | `_hydrate` default; **not** used by the section route |
| `rerank_top_k` | 6 | overview/free-form rerank |
| `max_context_tokens` | 6000 | overview/free-form packing |
| `indication_candidate_limit` | 8 | drugs shown for a reverse lookup |
| `indication_retrieval_limit` | 40 | chunk pool before grouping |
| `indication_evidence_per_drug` | 2 | |
| `patient_candidate_limit` | 2 | stage-2 safety |
| `safety_hits_per_section` | 1 | |
| `safety_sections_per_candidate` | 4 | |
## What this pipeline is *not*
Stated plainly because the terms get reused loosely:
- **Not BM25.** `search_lexical` scores a hit as *the count of distinct matched
query tokens* — no term frequency, no IDF, no length normalisation. The
docstring calls it "a transparent stand-in for a real BM25 score".
- **Not hybrid search.** `rag/fusion.py` implements reciprocal-rank fusion and is
tested, but **no runtime code calls it**. Dense and lexical results are never
fused.
- **No multi-query / query expansion.** `rag/expansion.py` (sibling expansion)
exists and is tested but has **no runtime caller**. No rewritten-query
retrieval exists anywhere.
- **No parent-child hydration in practice.** The code path exists
(`_hydrate``ParentStore.get`) but no chunk in the loaded corpus carries a
`parent_id`.
- **No filters on `atc_codes`.** The field is indexed and stored; nothing
queries it.
## Section keyword resolver — `rag/sections.py`
Used by the legacy `retrieve()` path (no generator configured). Two rules make
it safe:
- **Longest phrase wins.** All phrases across all sections are sorted by length,
so `chống chỉ định` is tested before `chỉ định` — they differ by one prefix
word and mean opposite things. The same rule keeps `quá liều` from being read
as `liều`.
- **No match is not a guess.** An unrecognised question returns `None` and the
caller falls back to similarity. This layer never picks a section it is unsure
of.
Adding a phrasing means adding an entry to `SECTION_PHRASES`, never editing the
matching code.