Files
duocthu/docs-legacy/10-rag-orchestration.md
T

222 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 10 — RAG orchestration
Implementation: `apps/ai-service/rag/agent.py` (`RagAgent`, 776 lines).
Tests: `tests/test_agent.py`, `tests/test_clinical_condition_flow.py`,
`tests/test_budget.py`.
Decision record: `docs/adr/0008-llm-understanding-one-shot-rag.md` (supersedes
ADR 0007).
## No framework
There is **no** LangChain, LlamaIndex, Haystack, or agent library anywhere in
the dependency set (`apps/ai-service/pyproject.toml` and the `Dockerfile`'s
inline pip list both confirm it). Orchestration is a plain Python class with a
hand-written branch table. `rag/` imports no SDK at all — the LLM arrives as a
`JsonLlm` / `AnswerGenerator` protocol.
## The two operating modes
`bootstrap.py::build_runtime` returns different graphs depending on config:
| `ANSWER_PROVIDER` | `app.state.conversational` | Live path |
|---|---|---|
| `disabled` | `None` | Retrieval-only, single-turn, through `GroundedAnswerService.answer()` + `QueryRoutingService` (fuzzy resolver + keyword section router). Evidence is quoted verbatim. |
| `stub` / `bedrock-converse` / `bedrock-claude` | `RagAgent` | The full understanding-driven path described below |
With `EMBEDDING_PROVIDER=disabled`, `build_runtime` returns `(None, None,
trace_writer, metrics)` and `/ready` answers 503 only if the embedding provider
was *not* disabled — so a disabled deployment reports ready while
`POST /v1/rag/query` returns 503 from the dependency.
## `RagAgent.handle()` — one turn
```mermaid
sequenceDiagram
participant R as routers/rag.py
participant A as RagAgent
participant B as RequestBudget
participant S as PostgresConversationStore
participant U as LlmQueryUnderstander
participant RS as RetrievalService
participant GA as GroundedAnswerService
R->>A: handle(turn, conversation_id)
A->>B: start(40_000 ms, 8 calls)
A->>S: recent(conversation_id, history_turns*2 = 12)
Note over A,S: fail-open — a store outage means this turn has no memory
A->>U: understand(turn, history, budget, prior_frame)
U-->>A: QueryFrame
A->>A: _route(turn, frame, budget)
alt retrieval needed
A->>RS: retrieve_framed / retrieve_by_indication / per-drug interaction
A->>GA: answer_from_result(..., prechecked=True)
end
A->>A: _enforce_clarify_circuit_breaker()
A->>S: append(conversation_id, lines)
A->>A: _last_frame[conversation_id] = frame
A-->>R: AgentReply
```
## The routing table — `_route`
Order matters; the first match wins.
| # | Condition | Outcome |
|---|---|---|
| 1 | `looks_non_human(turn)` | `abstain / out_of_scope` — a deterministic scope guard **before** any conversational clarify, so an out-of-scope request never looks recoverable |
| 2 | `dosing_calc` + drugs + not a section overview, `population is None` | `clarify / missing_population` |
| 3 | same, paediatric and (`age_text` or `weight_kg` missing) | `clarify / missing_pediatric_age_or_weight` |
| 4 | `needs_clarify` + reason, and not `dosing_calc`/`condition_to_drug`/overview | `clarify / needs_more_info`, or `abstain / <system_error>` if the understanding call itself failed |
| 5 | `condition_to_drug` / `symptom_to_drug`, condition ambiguous | `clarify / ambiguous_condition` |
| 6 | same, `condition_relation != INDICATION` | `abstain / unsupported_reverse_relation` |
| 7 | same, condition or indication present | `_condition_to_drug()` |
| 8 | same, neither present | `clarify / no_condition` or `no_indication` |
| 9 | `condition_relation` turn type | `abstain / unsupported_reverse_relation` |
| 10 | `drug_attribute` with drugs but no attribute | `clarify / missing_attribute` |
| 11 | `smalltalk` | `answerable / smalltalk` (fixed greeting) |
| 12 | `out_of_scope` | `abstain / out_of_scope` |
| 13 | no drugs, but `unknown_drugs` | `abstain / drug_not_in_formulary` naming them |
| 14 | no drugs at all | `clarify / no_drug` |
| 15 | `interaction` with ≥2 drugs | `_interaction()` |
| 16 | otherwise | `_single_drug()` |
### Why dosing is a state machine, not a model opinion
The LLM extracts the fields; **code** decides which are required. Live testing
caught the model asking an adult's weight repeatedly after the user had supplied
a route, and previously dumping oral + rectal regimens together.
Paediatric turns require **both** age and weight, because the formulary branches
on both — paracetamol prints an age band (`Trẻ em 4-6 tuổi: 240 mg`) *and* a
weight rule (`10-50 kg: 15 mg/kg`), so answering with only one means picking a
regimen the source does not let you pick.
What changed on 2026-08-11 is the *question*, not the gate:
`_pediatric_clarify_question` now asks only for the missing field and echoes back
the known one (`"Bé nặng 18 kg, vậy bé bao nhiêu tuổi?"`). Reproduced 5/5 before
the fix: `"Bé 18 ký …"`, `"Bé nặng 18 kg …"` and `"Trẻ 5 tuổi …"` all received
the same generic sentence.
**Route is deliberately not a universal required slot.** Retrieval and the answer
contract decide from the actual evidence whether omitting it is harmless (one
applicable route → answer now) or materially ambiguous (several routes → clarify
with model-proposed quick replies). This prevents a chip funnel for a question
that was already precise enough.
### Clarify circuit breaker
```python
MAX_CONSECUTIVE_CLARIFY = 4
```
Found live 2026-08-07: the understanding model could re-ask the same clarifying
question forever — reproduced three times independently, one case never
converging after five real answered turns. `_merge_with_prior_frame` addresses
most of the cause; this is the code-level bound, because nothing otherwise stops
a model that keeps deciding `needs_clarify=true`. Any non-clarify decision resets
the streak. On trip it returns `abstain / clarify_loop_exhausted` with an
instruction to restate the whole question or start a new session.
The streak counter is **in-process only** — see
[02-system-architecture.md](02-system-architecture.md#the-stateful-detail-that-constrains-scaling).
## `_synthesize_query` — the context that reaches generation
`GroundedAnswerService.answer_from_result` has **no conversation history of its
own**; the `query` string it receives *is* the entire context its generation call
sees. `_synthesize_query` folds the resolved frame into one self-contained
question:
```
<turn>. Đối tượng: trẻ em. Tuổi: 5 tuổi. Cân nặng: 18 kg. Đường dùng: uống.
Chỉ định/triệu chứng: …. Bệnh nền: …. Dữ kiện thận: ….
```
Without it, a reply like `"Uống"` three turns into a dose conversation would
reach generation as just `"Uống"` — the two P0s the 2026-08-06 audit named
(population/weight/age/route extracted then discarded downstream) are exactly
this gap. Redundant when the turn is already self-contained; omission is the
failure mode, not repetition.
For a **patient-specific** candidate list, `_patient_generation_query` is used
instead. It deliberately withholds the raw patient values from the prompt: those
values have already done their job (selecting safety sections) and are not Dược
thư evidence, so restating them inside a cited claim would be — correctly —
rejected by the numeric grounding guard.
## Interaction path
For each named drug, retrieve its `tuong_tac_thuoc` section; keep parts whose
decision is `ANSWERABLE` **or** `VERIFY_PDF`; then pass the combined pool through
`RetrievalService.decide()`.
Keeping `VERIFY_PDF` parts is deliberate. Previously only `ANSWERABLE` parts were
kept, so a quarantined drug's evidence — and the "table exists, verify PDF"
notice the quarantine contract requires — was silently dropped, and a confident
interaction answer could omit exactly the unverified contraindication table it
should have flagged.
If no evidence at all: `abstain / no_interaction_evidence`, worded as *"not found
in each drug's interaction section"* and explicitly **not** as "safe":
> Điều này KHÔNG có nghĩa là an toàn khi phối hợp.
## Condition → drug path
1. Retrieve by indication (keyword, then dense fallback).
2. Derive matched drugs from `matched_doc_id` (`{drug_id}__chi_dinh__{n}`) — the
drugs actually found, never `frame.drugs`, which is empty by construction for
this turn type.
3. If the patient context requires a safety review, run stage 2
(`assess_patient_candidates`) and abstain if it produces no safety evidence —
*"Không suy ra thuốc là phù hợp/an toàn."*
4. Generate in `list_mode=True` with the candidate `drug_id` set bound into the
prompt and validated after generation.
The docstring is explicit that this is a factual list, not a treatment ranking:
no drug is preferred over another, and absence is stated plainly rather than as
"no such drug exists".
## Request budget — `rag/budget.py`
```python
max_wall_clock_ms = 40_000 # MAX_WALL_CLOCK_MS
max_llm_calls_per_turn = 8 # MAX_LLM_CALLS_PER_TURN
```
`budget.require()` is called immediately before each provider call and raises
`RequestBudgetExhausted` (a subclass of `AnswerGenerationUnavailable`, so every
existing fail-closed handler already does the right thing).
Its stated limit: it is checked **between** calls and cannot cancel a boto3 call
already in flight. That residual gap is bounded separately by
`read_timeout=20` with `total_max_attempts=2` in
`adapters/bedrock_converse.py`. The realistic worst case is therefore ~40 s plus
one in-flight call ≈ 60 s — which is why the browser timeout in
`ChatPanel.tsx` is 65 s.
## LLM calls per turn
| Call | When | Fail behaviour |
|---|---|---|
| 1. Understanding | Always (agent path) | Closed |
| 2. Sufficiency | Only on the legacy path — skipped when `prechecked=True` (i.e. always, on the agent path) or in `list_mode`, or with <2 evidence blocks | **Open** |
| 3. Generation | When evidence is answerable | Closed |
| 3b. Generation retry | Only when the model self-reported `evidence_sufficient=false` with no clarifying question | Closed |
| 4. Entailment | After grounding passes | Closed |
| 56. Completeness repair + re-verify | Only when entailment reports a *grounded* omission | Closed |
So a normal answerable agent turn is **3** sequential Bedrock calls; the
pathological ceiling is 8 (the budget), of which the repair path is the most
likely to exhaust it — observed live on an Isosorbid dinitrat dosage turn at
40.3 s against the 40 s budget.
## What ADR 0007 described and this replaced
ADR 0007's `Focus`/`ConversationState`/TTL design and its
PLAN/RETRIEVE/ASSESS/REFINE/VERIFY bounded loop, along with
`rag/conversation.py` and `rag/reasoning.py`, are **gone from the tree**. The
`LOOP_ROUNDS`, `LOOP_REFINED`, `LOOP_REPAIRED` and `FOLLOWUP_INHERITED` metric
names in `rag/metrics.py` are leftovers of that design and are no longer
incremented anywhere — see [27-technical-debt.md](27-technical-debt.md).