9.1 KiB
08 — Query understanding
Implementation: apps/ai-service/rag/understanding.py (1,030 lines),
rag/routing.py::CatalogDrugResolver, rag/clinical.py, rag/policy.py.
Tests: tests/test_understanding.py, tests/test_policy.py,
tests/test_clinical_condition_flow.py.
One LLM call per turn produces a QueryFrame. Nothing here answers a medical
question — the frame is intent only.
Why an LLM replaced the heuristics
The previous front end resolved drugs with difflib.SequenceMatcher and routed
sections with a Vietnamese phrase table. The module docstring lists the measured
failures: aspirinol false-matched to aspirin, the correctly-spelled English
INN amoxicillin tied, and the common word uống was read as a drug.
The safety property: candidates are bounded before the model runs
flowchart LR
T[turn + history lines]
R["CatalogDrugResolver.resolve(line)<br/>exact alias span match"]
S["CatalogDrugResolver.suggest(line, k=5, min_score=0.55)<br/>fuzzy, only when no exact match"]
C["candidate drug_id set"]
P["prompt shows ONLY these drug_ids"]
L[LLM]
V["_resolve_id(): output must be in the shown set<br/>(underscore/space form tolerated)"]
F[QueryFrame.drugs]
U[QueryFrame.unknown_drugs]
T --> R --> C
T --> S --> C
C --> P --> L --> V
V -->|in set| F
V -->|not in set| U
A catalog whitelist alone would not be enough, and the code says why
(finding F-04): validating that an output id is some real drug_id does not
prove it is the one the user's text named — a model could satisfy that whitelist
while mapping an invented name onto any of the other 683 real drugs. Bounding the
candidate set first removes that degree of freedom: amoxicillin → amoxicilin
still works (fuzzy puts it in the set), but aspirinol cannot become aspirin
because nothing about aspirinol fuzzy-matches aspirin.
The same change also bounded token cost — the full 684-drug catalog was previously sent on every turn.
Resolver performance
CatalogDrugResolver.resolve and .suggest are both @lru_cache(maxsize=4096).
The comment records the measurement: over the real ~10,164-alias catalog,
resolve() costs ~0.65–0.7 s and suggest() ~0.94–0.97 s, and
_candidate_ids calls both per history line, every turn. An ordinary
multi-turn conversation was enough to exhaust the request budget before the
first Bedrock call, surfacing as a false "service outage".
Exact matching also enumerates the query's contiguous token spans against an
immutable alias index (_alias_to_drug_ids) instead of compiling ~10k regexes,
making the common path O(q²) in the short query rather than O(catalog).
QueryFrame
| Field | Type | Meaning |
|---|---|---|
turn_type |
one of 10 | The router's primary branch |
drugs |
tuple[str] | Canonical drug_ids, catalog-bounded |
unknown_drugs |
tuple[str] | Named but not in the catalog |
attribute |
section key | None | Validated against SECTION_KEYS |
population |
enum | None | tre_em, nguoi_lon, suy_than, … |
weight_kg |
float | None | Accepted only in (0, 500] |
age_text |
str | None | As stated |
indication |
str | None | |
condition |
ConditionQuery | None |
Normalized condition + subtype + ambiguity |
condition_relation |
indication|adverse_effect|contraindication|unknown |
|
patient_context |
PatientContext |
Comorbidities, allergies, ADRs, current meds, renal, hepatic, pregnancy, labs |
context_action |
none|continue|new |
Case continuity |
route |
enum | None | uong, tiem_tinh_mach, dat_truc_trang, … |
section_overview |
bool | Survey the whole section vs. decide for one patient |
standalone_query |
str | None | Turn rewritten self-contained |
depends_on_previous_turn |
bool | |
needs_clarify, clarify_reason, quick_replies |
Ask-back | |
system_error |
str | None | Set only on a genuine technical failure |
raw |
dict | The model's raw JSON, excluded from equality |
system_error exists because a provider outage and a genuine clarifying
question previously produced the identical downstream
reason="needs_more_info", making a real outage indistinguishable from normal
traffic in the API response and in metrics.
The 10 turn types
drug_attribute, drug_overview, interaction, symptom_to_drug,
condition_to_drug, drug_to_condition, condition_relation, dosing_calc,
smalltalk, out_of_scope.
condition_relation exists specifically so "which drug causes X" and "which
drug is contraindicated in X" are never collapsed into an indication lookup.
Prompt construction
The user message (understand()) is assembled from four blocks:
- Candidate drug list —
drug_id\tnamefor the bounded set, with an explicit note that this is not the whole formulary. - Section keys with glosses —
SECTION_KEY_HINTS. Bare slugs were insufficient: 9/9 live calls for "X cần thận trọng gì?" pickedchong_chi_dinh, answering from the wrong section. Thethan_tronggloss now spells out the distinction in capitals. THÔNG TIN ĐÃ XÁC ĐỊNH TỪ CÁC LƯỢT TRƯỚC— a structured summary of the prior frame (_known_facts_block), so established facts are data rather than something to re-derive from a growing transcript.- History then the current turn.
bootstrap.py::_catalog_names decides which alias to show per drug. It always
shows the drug_id's own name form first: paracetamol has 191 aliases, and the
alphabetically-first three were 0Frezefev, ABAB, Ace kid 80 — none
recognisable — after which the model read an earlier "paracetamol" mention as an
unknown drug and answered "not in the formulary" for a drug that plainly is.
Deterministic post-conditions
The LLM output passes through four narrow, knowledge-free rewrites. Each covers an unambiguous surface form where the model's routing would reverse the requested relation:
| Function | Trigger | Effect |
|---|---|---|
_apply_condition_candidate_cue |
a known condition alias + a candidate cue (dùng thuốc gì, lựa chọn thuốc nào, …) |
force condition_to_drug + indication |
_apply_broad_condition_cue |
a broad disease→drug question with no named drug | force condition_to_drug |
_apply_reverse_relation_cues |
thuốc nào gây …, thuốc nào chống chỉ định … |
force condition_relation + the correct relation |
_apply_named_drug_cues |
an explicitly named drug + có tác dụng gì / có chống chỉ định |
force drug_to_condition / drug_attribute |
None of them contains disease or drug knowledge, and none creates a candidate.
Prior-frame merge
_merge_with_prior_frame is the code-level backstop for the model dropping an
already-known field. It fires only when:
- the turn is continuing a case (
context_action == continueordepends_on_previous_turn), or the prior turn was itself a clarify; and context_action != new; and- this turn's own
drugsagree with the prior frame (empty, or the same).
A turn that resolves a different drug is a genuine topic change and inherits nothing — this is the guard against the reproduced "headache question answered about OMEPRAZOL" bleed.
Validation and fail-closed behaviour
| Failure | Result |
|---|---|
AnswerGenerationUnavailable |
Frame with turn_type="out_of_scope", needs_clarify=True, system_error="understanding_provider_unavailable", logged with the real exception |
| Unparseable JSON | system_error="understanding_malformed_output" |
turn_type not in TURN_TYPES |
falls back to drug_attribute if drugs were resolved, else out_of_scope |
attribute not in SECTION_KEYS |
→ None |
population/route outside the allowed set |
→ None |
weight_kg outside (0, 500] |
→ None |
| A named drug not in the shown candidate set | → unknown_drugs, never a fuzzy substitution |
quick_replies |
max 4 items, max 40 chars each, de-duplicated |
Before F-10 this call site had no error handling at all — a provider outage propagated into an unhandled 500 rather than the graceful abstain every other failure mode gets.
Subject-scope policy — rag/policy.py
Deliberately not an LLM call: this gate runs on every request, so it must be cheap, available during a provider outage, and auditable as a fixed rule.
resolve_subject_scope(query, claimed) takes the more conservative of the
caller's claim and a keyword scan (cho cho, cho meo, thu y, gia suc, …
on diacritic-stripped text). A caller can narrow scope but never widen
it — the shipped web BFF hard-codes subject_scope: "human" on every request
without reading the message, which is exactly the review finding (F-02) this
module answers.
It is a corpus-coverage check, not clinical gatekeeping. The module docstring is
explicit that it must never be extended into restricting what a professional is
allowed to ask; the old QueryIntent.RECOMMENDATION keyword detector was
removed for that reason.
rag/agent.py still keeps its own narrower looks_non_human call as a
deterministic guard before every conversational clarify.