11 KiB
10 — RAG orchestration
Implementation: apps/ai-service/rag/agent.py (RagAgent, 776 lines).
Tests: tests/test_agent.py, tests/test_clinical_condition_flow.py,
tests/test_budget.py.
Decision record: docs/adr/0008-llm-understanding-one-shot-rag.md (supersedes
ADR 0007).
No framework
There is no LangChain, LlamaIndex, Haystack, or agent library anywhere in
the dependency set (apps/ai-service/pyproject.toml and the Dockerfile's
inline pip list both confirm it). Orchestration is a plain Python class with a
hand-written branch table. rag/ imports no SDK at all — the LLM arrives as a
JsonLlm / AnswerGenerator protocol.
The two operating modes
bootstrap.py::build_runtime returns different graphs depending on config:
ANSWER_PROVIDER |
app.state.conversational |
Live path |
|---|---|---|
disabled |
None |
Retrieval-only, single-turn, through GroundedAnswerService.answer() + QueryRoutingService (fuzzy resolver + keyword section router). Evidence is quoted verbatim. |
stub / bedrock-converse / bedrock-claude |
RagAgent |
The full understanding-driven path described below |
With EMBEDDING_PROVIDER=disabled, build_runtime returns (None, None, trace_writer, metrics) and /ready answers 503 only if the embedding provider
was not disabled — so a disabled deployment reports ready while
POST /v1/rag/query returns 503 from the dependency.
RagAgent.handle() — one turn
sequenceDiagram
participant R as routers/rag.py
participant A as RagAgent
participant B as RequestBudget
participant S as PostgresConversationStore
participant U as LlmQueryUnderstander
participant RS as RetrievalService
participant GA as GroundedAnswerService
R->>A: handle(turn, conversation_id)
A->>B: start(40_000 ms, 8 calls)
A->>S: recent(conversation_id, history_turns*2 = 12)
Note over A,S: fail-open — a store outage means this turn has no memory
A->>U: understand(turn, history, budget, prior_frame)
U-->>A: QueryFrame
A->>A: _route(turn, frame, budget)
alt retrieval needed
A->>RS: retrieve_framed / retrieve_by_indication / per-drug interaction
A->>GA: answer_from_result(..., prechecked=True)
end
A->>A: _enforce_clarify_circuit_breaker()
A->>S: append(conversation_id, lines)
A->>A: _last_frame[conversation_id] = frame
A-->>R: AgentReply
The routing table — _route
Order matters; the first match wins.
| # | Condition | Outcome |
|---|---|---|
| 1 | looks_non_human(turn) |
abstain / out_of_scope — a deterministic scope guard before any conversational clarify, so an out-of-scope request never looks recoverable |
| 2 | dosing_calc + drugs + not a section overview, population is None |
clarify / missing_population |
| 3 | same, paediatric and (age_text or weight_kg missing) |
clarify / missing_pediatric_age_or_weight |
| 4 | needs_clarify + reason, and not dosing_calc/condition_to_drug/overview |
clarify / needs_more_info, or abstain / <system_error> if the understanding call itself failed |
| 5 | condition_to_drug / symptom_to_drug, condition ambiguous |
clarify / ambiguous_condition |
| 6 | same, condition_relation != INDICATION |
abstain / unsupported_reverse_relation |
| 7 | same, condition or indication present | _condition_to_drug() |
| 8 | same, neither present | clarify / no_condition or no_indication |
| 9 | condition_relation turn type |
abstain / unsupported_reverse_relation |
| 10 | drug_attribute with drugs but no attribute |
clarify / missing_attribute |
| 11 | smalltalk |
answerable / smalltalk (fixed greeting) |
| 12 | out_of_scope |
abstain / out_of_scope |
| 13 | no drugs, but unknown_drugs |
abstain / drug_not_in_formulary naming them |
| 14 | no drugs at all | clarify / no_drug |
| 15 | interaction with ≥2 drugs |
_interaction() |
| 16 | otherwise | _single_drug() |
Why dosing is a state machine, not a model opinion
The LLM extracts the fields; code decides which are required. Live testing caught the model asking an adult's weight repeatedly after the user had supplied a route, and previously dumping oral + rectal regimens together.
Paediatric turns require both age and weight, because the formulary branches
on both — paracetamol prints an age band (Trẻ em 4-6 tuổi: 240 mg) and a
weight rule (10-50 kg: 15 mg/kg), so answering with only one means picking a
regimen the source does not let you pick.
What changed on 2026-08-11 is the question, not the gate:
_pediatric_clarify_question now asks only for the missing field and echoes back
the known one ("Bé nặng 18 kg, vậy bé bao nhiêu tuổi?"). Reproduced 5/5 before
the fix: "Bé 18 ký …", "Bé nặng 18 kg …" and "Trẻ 5 tuổi …" all received
the same generic sentence.
Route is deliberately not a universal required slot. Retrieval and the answer contract decide from the actual evidence whether omitting it is harmless (one applicable route → answer now) or materially ambiguous (several routes → clarify with model-proposed quick replies). This prevents a chip funnel for a question that was already precise enough.
Clarify circuit breaker
MAX_CONSECUTIVE_CLARIFY = 4
Found live 2026-08-07: the understanding model could re-ask the same clarifying
question forever — reproduced three times independently, one case never
converging after five real answered turns. _merge_with_prior_frame addresses
most of the cause; this is the code-level bound, because nothing otherwise stops
a model that keeps deciding needs_clarify=true. Any non-clarify decision resets
the streak. On trip it returns abstain / clarify_loop_exhausted with an
instruction to restate the whole question or start a new session.
The streak counter is in-process only — see 02-system-architecture.md.
_synthesize_query — the context that reaches generation
GroundedAnswerService.answer_from_result has no conversation history of its
own; the query string it receives is the entire context its generation call
sees. _synthesize_query folds the resolved frame into one self-contained
question:
<turn>. Đối tượng: trẻ em. Tuổi: 5 tuổi. Cân nặng: 18 kg. Đường dùng: uống.
Chỉ định/triệu chứng: …. Bệnh nền: …. Dữ kiện thận: ….
Without it, a reply like "Uống" three turns into a dose conversation would
reach generation as just "Uống" — the two P0s the 2026-08-06 audit named
(population/weight/age/route extracted then discarded downstream) are exactly
this gap. Redundant when the turn is already self-contained; omission is the
failure mode, not repetition.
For a patient-specific candidate list, _patient_generation_query is used
instead. It deliberately withholds the raw patient values from the prompt: those
values have already done their job (selecting safety sections) and are not Dược
thư evidence, so restating them inside a cited claim would be — correctly —
rejected by the numeric grounding guard.
Interaction path
For each named drug, retrieve its tuong_tac_thuoc section; keep parts whose
decision is ANSWERABLE or VERIFY_PDF; then pass the combined pool through
RetrievalService.decide().
Keeping VERIFY_PDF parts is deliberate. Previously only ANSWERABLE parts were
kept, so a quarantined drug's evidence — and the "table exists, verify PDF"
notice the quarantine contract requires — was silently dropped, and a confident
interaction answer could omit exactly the unverified contraindication table it
should have flagged.
If no evidence at all: abstain / no_interaction_evidence, worded as "not found
in each drug's interaction section" and explicitly not as "safe":
Điều này KHÔNG có nghĩa là an toàn khi phối hợp.
Condition → drug path
- Retrieve by indication (keyword, then dense fallback).
- Derive matched drugs from
matched_doc_id({drug_id}__chi_dinh__{n}) — the drugs actually found, neverframe.drugs, which is empty by construction for this turn type. - If the patient context requires a safety review, run stage 2
(
assess_patient_candidates) and abstain if it produces no safety evidence — "Không suy ra thuốc là phù hợp/an toàn." - Generate in
list_mode=Truewith the candidatedrug_idset bound into the prompt and validated after generation.
The docstring is explicit that this is a factual list, not a treatment ranking: no drug is preferred over another, and absence is stated plainly rather than as "no such drug exists".
Request budget — rag/budget.py
max_wall_clock_ms = 40_000 # MAX_WALL_CLOCK_MS
max_llm_calls_per_turn = 8 # MAX_LLM_CALLS_PER_TURN
budget.require() is called immediately before each provider call and raises
RequestBudgetExhausted (a subclass of AnswerGenerationUnavailable, so every
existing fail-closed handler already does the right thing).
Its stated limit: it is checked between calls and cannot cancel a boto3 call
already in flight. That residual gap is bounded separately by
read_timeout=20 with total_max_attempts=2 in
adapters/bedrock_converse.py. The realistic worst case is therefore ~40 s plus
one in-flight call ≈ 60 s — which is why the browser timeout in
ChatPanel.tsx is 65 s.
LLM calls per turn
| Call | When | Fail behaviour |
|---|---|---|
| 1. Understanding | Always (agent path) | Closed |
| 2. Sufficiency | Only on the legacy path — skipped when prechecked=True (i.e. always, on the agent path) or in list_mode, or with <2 evidence blocks |
Open |
| 3. Generation | When evidence is answerable | Closed |
| 3b. Generation retry | Only when the model self-reported evidence_sufficient=false with no clarifying question |
Closed |
| 4. Entailment | After grounding passes | Closed |
| 5–6. Completeness repair + re-verify | Only when entailment reports a grounded omission | Closed |
So a normal answerable agent turn is 3 sequential Bedrock calls; the pathological ceiling is 8 (the budget), of which the repair path is the most likely to exhaust it — observed live on an Isosorbid dinitrat dosage turn at 40.3 s against the 40 s budget.
What ADR 0007 described and this replaced
ADR 0007's Focus/ConversationState/TTL design and its
PLAN/RETRIEVE/ASSESS/REFINE/VERIFY bounded loop, along with
rag/conversation.py and rag/reasoning.py, are gone from the tree. The
LOOP_ROUNDS, LOOP_REFINED, LOOP_REPAIRED and FOLLOWUP_INHERITED metric
names in rag/metrics.py are leftovers of that design and are no longer
incremented anywhere — see 27-technical-debt.md.