Files
duocthu/docs/09-retrieval-pipeline.md
T

9.8 KiB
Raw Blame History

09 — Retrieval pipeline

Implementation: apps/ai-service/rag/service.py (RetrievalService, 741 lines), apps/ai-service/adapters/qdrant.py (501 lines), rag/sections.py, rag/context.py. Tests: tests/test_retrieval_service.py, tests/test_section_routing.py, tests/test_qdrant_adapter.py, tests/test_rerank_overview.py, tests/test_section_order.py.

What retrieval is here

Similarity is the fallback, not the default. Measured 2026-08-04: letting vector similarity choose the section gives hit@1 0.544 overall and 0.05 on chong_chi_dinh, because duoc_ly_va_co_che_tac_dung is the largest section and sits close to almost any question about the drug. When the question names the section it wants, a payload filter answers it exactly.

That single measurement is the reason the architecture looks the way it does.

Retrieval routes

flowchart TD
    IN["retrieve_framed(drug_id, section_key, query, is_overview)"]
    S{section_key given?}
    SEC["find_by_section(drug_id, section_key)<br/>Qdrant scroll, payload filter, NO vector<br/>score = 1.0 by construction"]
    POOL["_pooled_neighbour_hits<br/>only when section == than_trong"]
    OV["find_by_drug(drug_id)<br/>every prose section, book order"]
    ISOV{is_overview?}
    INTRO["keep INTRO_SECTIONS only:<br/>ten_chung_quoc_te, loai_thuoc,<br/>chi_dinh, duoc_ly_va_co_che_tac_dung"]
    RR["_rerank(query, hits)<br/>Cohere rerank-v3.5, top_k=6, fail-open"]
    PACK["pack_evidence(max_tokens=6000)"]
    DEC["_decide(evidence)"]

    IN --> S
    S -->|yes| SEC --> POOL --> DEC
    S -->|no| OV
    OV -->|None| AB["ABSTAIN insufficient_retrieval_score"]
    OV --> ISOV
    ISOV -->|yes| INTRO --> DEC
    ISOV -->|no| RR --> PACK --> DEC

1. Section route (the primary path)

find_by_section is a scroll, not a search — it must not be a top-k. Paging continues until the offset is exhausted, because Qdrant's default page is 256 and a long section silently truncated would read as a complete answer. Results are re-sorted by part_index (see 07-indexing-and-storage.md for why). No evidence limit is applied — the whole section is the answer, and a truncated list of contraindications reads as a complete one.

Score is 1.0 because the match is exact by construction. It is not a similarity and must not be compared to one.

2. Bounded cross-section pooling

_pooled_neighbour_hits uses search_lexical to find another section of the same drug whose text matches the query strongly. It exists for one measured case: a thận trọng question about a specific condition (loét dạ dày) whose real answer was filed only under chống chỉ định.

It is deliberately narrow:

_LEXICAL_POOL_ENABLED_SECTIONS  = {"than_trong"}                     # only this route
_LEXICAL_POOL_EXCLUDED_SECTIONS = {"duoc_ly_va_co_che_tac_dung"}     # the known attractor
_LEXICAL_POOL_MIN_SCORE = 5.0
_MAX_LEXICAL_POOLED_SECTIONS = 2

The excluded section is excluded outright rather than by score margin: on the exact query that motivated the mechanism, the true positive scored 7 matched terms and that attractor scored 6 — too close for a threshold to separate. Applying pooling to every section leaked a lexically-overlapping interaction section into a dosage answer, so it stayed opt-in.

3. Drug overview (a bare drug name)

find_by_drug scrolls every prose chunk of the drug (block descriptors stay out of a text answer), orders by SECTION_ORDER then part_index, and prefixes each section's first chunk with 【display name】.

For turn_type == "drug_overview" only the four INTRO_SECTIONS are kept. Without that split a bare drug name sent the entire ~29-section monograph as evidence for every generation call — wrong retrieval, and an answer long enough to intermittently fail generation outright.

4. Free-form question about a resolved drug

The full monograph is reranked to rerank_top_k=6, then packed to a token budget rather than a flat count:

max_context_tokens = 6000   # pack_evidence, rag/context.py

pack_evidence packs whole blocks in retrieval order and never truncates clinical text; anything that does not fit is recorded in omitted_evidence_ids. The cap is applied even when rerank is disabled or fails open — an ordering aid must not also remove the size bound.

5. Reverse lookup: condition/indication → drugs

retrieve_by_indication is a two-stage lookup, keyword first:

  1. find_by_indication — scroll every chi_dinh prose chunk and require the normalized indication to appear as a contiguous, word-boundary-anchored phrase. Not a substring (false positives after diacritic stripping), and explicitly not a token-subset match: a nonsense phrase built from common filler words previously false-positived against real chi_dinh text and reached generation before being caught. Score rewards an early, concise mention: 1 + 1/(1+position) + 1/(1 + words/40).
  2. search_indication — dense fallback, tried only when the keyword pass finds nothing, filtered to section_key=chi_dinh and chunk_kind=prose. This is the only place in the live path where dense vector search is actually used (ADR 0008). A weak top score (< evidence_minimum_score) discards the hits, because dense search always returns its nearest neighbours — a made-up phrase still got 8 unrelated "matches" live.

The adapter returns a ranked chunk pool; _rank_indication_drugs groups by drug_id, takes the max score per drug (never a sum or count, so a drug with more chunks does not win), optionally reranks the groups, and the service caps at 8 drugs × 2 evidence chunks.

6. Patient-specific safety evidence (stage 2)

assess_patient_candidates / retrieve_patient_drug_context never create candidates. For each already-indicated drug they run separate, relation-specific lexical searches:

Facet Query source Sections searched
interaction patient.interaction_query() tuong_tac_thuoc
warnings patient.warning_query() chong_chi_dinh, than_trong (requires a clinical-anchor match)
dosage context patient.dosage_context_query() lieu_luong_va_cach_dung (requires a clinical-anchor match)
pregnancy / breastfeeding direct section route thoi_ky_mang_thai, thoi_ky_cho_con_bu

Keeping the queries separate is the point: a current medicine may select an interaction chunk only when that medicine matches inside the interaction section — CKD or age terms from another facet cannot make an unrelated interaction look supported. _patient_context_matches requires a real clinical anchor rather than overlap on generic words like chức năng.

Absence of a hit is recorded as CandidateStatus.INSUFFICIENT_EVIDENCE — never as "safe".

The evidence decision — _decide

if not evidence:                                   ABSTAIN  "parent_hydration_failed"
if any(not item.source_refs for item in evidence): ABSTAIN  "missing_provenance"
if any(item.requires_visual_check ...):            VERIFY_PDF "visual_verification_required"
else:                                              ANSWERABLE "grounded_evidence_available"

decide() is exposed publicly so a caller assembling its own pool across several retrieve calls — RagAgent._interaction — gets the same quarantine and provenance policy. Bypassing it is precisely how the interaction path once silently dropped a quarantined drug's evidence instead of surfacing VERIFY_PDF.

requires_visual_check is read from the payload as requires_visual_check OR has_quarantined_content.

Policy constants — EvidencePolicy

Setting Default Applies to
minimum_score 0.12 (EVIDENCE_MINIMUM_SCORE) dense routes only
candidate_limit 5 retrieve()'s dense search
evidence_limit 3 _hydrate default; not used by the section route
rerank_top_k 6 overview/free-form rerank
max_context_tokens 6000 overview/free-form packing
indication_candidate_limit 8 drugs shown for a reverse lookup
indication_retrieval_limit 40 chunk pool before grouping
indication_evidence_per_drug 2
patient_candidate_limit 2 stage-2 safety
safety_hits_per_section 1
safety_sections_per_candidate 4

What this pipeline is not

Stated plainly because the terms get reused loosely:

  • Not BM25. search_lexical scores a hit as the count of distinct matched query tokens — no term frequency, no IDF, no length normalisation. The docstring calls it "a transparent stand-in for a real BM25 score".
  • Not hybrid search. rag/fusion.py implements reciprocal-rank fusion and is tested, but no runtime code calls it. Dense and lexical results are never fused.
  • No multi-query / query expansion. rag/expansion.py (sibling expansion) exists and is tested but has no runtime caller. No rewritten-query retrieval exists anywhere.
  • No parent-child hydration in practice. The code path exists (_hydrateParentStore.get) but no chunk in the loaded corpus carries a parent_id.
  • No filters on atc_codes. The field is indexed and stored; nothing queries it.

Section keyword resolver — rag/sections.py

Used by the legacy retrieve() path (no generator configured). Two rules make it safe:

  • Longest phrase wins. All phrases across all sections are sorted by length, so chống chỉ định is tested before chỉ định — they differ by one prefix word and mean opposite things. The same rule keeps quá liều from being read as liều.
  • No match is not a guess. An unrecognised question returns None and the caller falls back to similarity. This layer never picks a section it is unsure of.

Adding a phrasing means adding an entry to SECTION_PHRASES, never editing the matching code.