Files
duocthu/docs/09-retrieval-pipeline.md
T

217 lines
9.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 09 — Retrieval pipeline
Implementation: `apps/ai-service/rag/service.py` (`RetrievalService`, 741 lines),
`apps/ai-service/adapters/qdrant.py` (501 lines), `rag/sections.py`,
`rag/context.py`.
Tests: `tests/test_retrieval_service.py`, `tests/test_section_routing.py`,
`tests/test_qdrant_adapter.py`, `tests/test_rerank_overview.py`,
`tests/test_section_order.py`.
## What retrieval is here
**Similarity is the fallback, not the default.** Measured 2026-08-04: letting
vector similarity choose the section gives hit@1 **0.544** overall and **0.05**
on `chong_chi_dinh`, because `duoc_ly_va_co_che_tac_dung` is the largest section
and sits close to almost any question about the drug. When the question names
the section it wants, a payload filter answers it exactly.
That single measurement is the reason the architecture looks the way it does.
## Retrieval routes
```mermaid
flowchart TD
IN["retrieve_framed(drug_id, section_key, query, is_overview)"]
S{section_key given?}
SEC["find_by_section(drug_id, section_key)<br/>Qdrant scroll, payload filter, NO vector<br/>score = 1.0 by construction"]
POOL["_pooled_neighbour_hits<br/>only when section == than_trong"]
OV["find_by_drug(drug_id)<br/>every prose section, book order"]
ISOV{is_overview?}
INTRO["keep INTRO_SECTIONS only:<br/>ten_chung_quoc_te, loai_thuoc,<br/>chi_dinh, duoc_ly_va_co_che_tac_dung"]
RR["_rerank(query, hits)<br/>Cohere rerank-v3.5, top_k=6, fail-open"]
PACK["pack_evidence(max_tokens=6000)"]
DEC["_decide(evidence)"]
IN --> S
S -->|yes| SEC --> POOL --> DEC
S -->|no| OV
OV -->|None| AB["ABSTAIN insufficient_retrieval_score"]
OV --> ISOV
ISOV -->|yes| INTRO --> DEC
ISOV -->|no| RR --> PACK --> DEC
```
### 1. Section route (the primary path)
`find_by_section` is a **`scroll`, not a `search`** — it must not be a top-k.
Paging continues until the offset is exhausted, because Qdrant's default page is
256 and a long section silently truncated would read as a complete answer.
Results are re-sorted by `part_index` (see
[07-indexing-and-storage.md](07-indexing-and-storage.md) for why). No evidence
limit is applied — the whole section is the answer, and a truncated list of
contraindications reads as a complete one.
Score is `1.0` because the match is exact by construction. It is **not** a
similarity and must not be compared to one.
### 2. Bounded cross-section pooling
`_pooled_neighbour_hits` uses `search_lexical` to find another section of the
*same drug* whose text matches the query strongly. It exists for one measured
case: a `thận trọng` question about a specific condition (loét dạ dày) whose
real answer was filed only under `chống chỉ định`.
It is deliberately narrow:
```python
_LEXICAL_POOL_ENABLED_SECTIONS = {"than_trong"} # only this route
_LEXICAL_POOL_EXCLUDED_SECTIONS = {"duoc_ly_va_co_che_tac_dung"} # the known attractor
_LEXICAL_POOL_MIN_SCORE = 5.0
_MAX_LEXICAL_POOLED_SECTIONS = 2
```
The excluded section is excluded outright rather than by score margin: on the
exact query that motivated the mechanism, the true positive scored 7 matched
terms and that attractor scored 6 — too close for a threshold to separate.
Applying pooling to every section leaked a lexically-overlapping interaction
section into a dosage answer, so it stayed opt-in.
### 3. Drug overview (a bare drug name)
`find_by_drug` scrolls every **prose** chunk of the drug (block descriptors stay
out of a text answer), orders by `SECTION_ORDER` then `part_index`, and prefixes
each section's first chunk with `【display name】`.
For `turn_type == "drug_overview"` only the four `INTRO_SECTIONS` are kept.
Without that split a bare drug name sent the entire ~29-section monograph as
evidence for every generation call — wrong retrieval, and an answer long enough
to intermittently fail generation outright.
### 4. Free-form question about a resolved drug
The full monograph is reranked to `rerank_top_k=6`, then packed to a **token**
budget rather than a flat count:
```python
max_context_tokens = 6000 # pack_evidence, rag/context.py
```
`pack_evidence` packs whole blocks in retrieval order and never truncates
clinical text; anything that does not fit is recorded in
`omitted_evidence_ids`. The cap is applied even when rerank is disabled or
fails open — an ordering aid must not also remove the size bound.
### 5. Reverse lookup: condition/indication → drugs
`retrieve_by_indication` is a two-stage lookup, keyword first:
1. **`find_by_indication`** — scroll every `chi_dinh` prose chunk and require the
normalized indication to appear as a **contiguous, word-boundary-anchored
phrase**. Not a substring (false positives after diacritic stripping), and
explicitly not a token-subset match: a nonsense phrase built from common
filler words previously false-positived against real `chi_dinh` text and
reached generation before being caught.
Score rewards an early, concise mention:
`1 + 1/(1+position) + 1/(1 + words/40)`.
2. **`search_indication`** — dense fallback, tried only when the keyword pass
finds nothing, filtered to `section_key=chi_dinh` and `chunk_kind=prose`.
**This is the only place in the live path where dense vector search is
actually used** (ADR 0008). A weak top score (`< evidence_minimum_score`)
discards the hits, because dense search always returns its nearest
neighbours — a made-up phrase still got 8 unrelated "matches" live.
The adapter returns a ranked **chunk** pool; `_rank_indication_drugs` groups by
`drug_id`, takes the **max** score per drug (never a sum or count, so a drug with
more chunks does not win), optionally reranks the groups, and the service caps
at 8 drugs × 2 evidence chunks.
### 6. Patient-specific safety evidence (stage 2)
`assess_patient_candidates` / `retrieve_patient_drug_context` never create
candidates. For each already-indicated drug they run separate, relation-specific
lexical searches:
| Facet | Query source | Sections searched |
|---|---|---|
| interaction | `patient.interaction_query()` | `tuong_tac_thuoc` |
| warnings | `patient.warning_query()` | `chong_chi_dinh`, `than_trong` (requires a clinical-anchor match) |
| dosage context | `patient.dosage_context_query()` | `lieu_luong_va_cach_dung` (requires a clinical-anchor match) |
| pregnancy / breastfeeding | direct section route | `thoi_ky_mang_thai`, `thoi_ky_cho_con_bu` |
Keeping the queries separate is the point: a current medicine may select an
interaction chunk only when *that medicine* matches inside the interaction
section — CKD or age terms from another facet cannot make an unrelated
interaction look supported. `_patient_context_matches` requires a real clinical
anchor rather than overlap on generic words like `chức năng`.
Absence of a hit is recorded as `CandidateStatus.INSUFFICIENT_EVIDENCE` — never
as "safe".
## The evidence decision — `_decide`
```python
if not evidence: ABSTAIN "parent_hydration_failed"
if any(not item.source_refs for item in evidence): ABSTAIN "missing_provenance"
if any(item.requires_visual_check ...): VERIFY_PDF "visual_verification_required"
else: ANSWERABLE "grounded_evidence_available"
```
`decide()` is exposed publicly so a caller assembling its own pool across several
retrieve calls — `RagAgent._interaction` — gets the same quarantine and
provenance policy. Bypassing it is precisely how the interaction path once
silently dropped a quarantined drug's evidence instead of surfacing `VERIFY_PDF`.
`requires_visual_check` is read from the payload as
`requires_visual_check OR has_quarantined_content`.
## Policy constants — `EvidencePolicy`
| Setting | Default | Applies to |
|---|---|---|
| `minimum_score` | 0.12 (`EVIDENCE_MINIMUM_SCORE`) | dense routes only |
| `candidate_limit` | 5 | `retrieve()`'s dense search |
| `evidence_limit` | 3 | `_hydrate` default; **not** used by the section route |
| `rerank_top_k` | 6 | overview/free-form rerank |
| `max_context_tokens` | 6000 | overview/free-form packing |
| `indication_candidate_limit` | 8 | drugs shown for a reverse lookup |
| `indication_retrieval_limit` | 40 | chunk pool before grouping |
| `indication_evidence_per_drug` | 2 | |
| `patient_candidate_limit` | 2 | stage-2 safety |
| `safety_hits_per_section` | 1 | |
| `safety_sections_per_candidate` | 4 | |
## What this pipeline is *not*
Stated plainly because the terms get reused loosely:
- **Not BM25.** `search_lexical` scores a hit as *the count of distinct matched
query tokens* — no term frequency, no IDF, no length normalisation. The
docstring calls it "a transparent stand-in for a real BM25 score".
- **Not hybrid search.** `rag/fusion.py` implements reciprocal-rank fusion and is
tested, but **no runtime code calls it**. Dense and lexical results are never
fused.
- **No multi-query / query expansion.** `rag/expansion.py` (sibling expansion)
exists and is tested but has **no runtime caller**. No rewritten-query
retrieval exists anywhere.
- **No parent-child hydration in practice.** The code path exists
(`_hydrate``ParentStore.get`) but no chunk in the loaded corpus carries a
`parent_id`.
- **No filters on `atc_codes`.** The field is indexed and stored; nothing
queries it.
## Section keyword resolver — `rag/sections.py`
Used by the legacy `retrieve()` path (no generator configured). Two rules make
it safe:
- **Longest phrase wins.** All phrases across all sections are sorted by length,
so `chống chỉ định` is tested before `chỉ định` — they differ by one prefix
word and mean opposite things. The same rule keeps `quá liều` from being read
as `liều`.
- **No match is not a guess.** An unrecognised question returns `None` and the
caller falls back to similarity. This layer never picks a section it is unsure
of.
Adding a phrasing means adding an entry to `SECTION_PHRASES`, never editing the
matching code.