Files
duocthu/docs/adr/0007-conversational-reasoning-rag.md
T

181 lines
7.7 KiB
Markdown

# ADR 0007: Conversational reasoning RAG — state, bounded loop, and how it is measured
**Status:** accepted, implementation in progress (2026-08-05)
**Supersedes:** nothing. Extends ADR 0005 (segment output contract) and ADR 0006
(quarantined block references) rather than replacing them.
## Context
The service answers one question at a time. `POST /v1/rag/query` carries no
conversation id, `apps/chat-service` holds zero source files, and every request
re-resolves the drug from scratch. Three consequences, all observed in the UI on
2026-08-05:
- `paracetamol` alone is refused rather than asked about.
- `liều dùng paracetamol cho người lớn` returns the identical answer to
`liều dùng paracetamol` — the qualifier is not used at any stage.
- A follow-up such as *"còn trẻ em thì sao?"* cannot work at all, because
nothing carries the drug forward.
The owner's requirement is a **conversational reasoning RAG**: history, an
internal reasoning stage, and a bounded self-improvement loop.
The binding constraint is that this is a drug formulary for clinicians. Every
capability below is designed so that adding it cannot widen what the system is
allowed to assert.
## Decision
### 1. Conversation state
Two stores with different jobs, deliberately not merged.
**`Focus` — structured, drives routing.** This is what makes *"còn trẻ em thì
sao?"* resolvable without an LLM.
| Field | Purpose |
|---|---|
| `drug_id`, `drug_name` | The drug under discussion |
| `section_key` | The attribute last answered |
| `population` | `nguoi_lon` / `tre_em` / `phu_nu_co_thai` / … |
| `verbosity` | `concise` \| `detailed`, set when the user asks |
| `set_at_turn` | Turn index each field was last set |
**`ConversationState` — the whole record.**
```
conversation_id
recent: tuple[Turn, ...] # last K turns, verbatim
summary: str # rolling prose summary of everything older
focus: Focus
turn_count: int
```
A `Turn` carries `role`, `text`, `at`, and — for assistant turns — the
`drug_id`, `section_key` and `evidence_ids` that produced it. Storing the
evidence ids is what lets the planner answer a follow-up **from evidence
already retrieved** instead of retrieving again.
**Carry-over is never silent.** An inherited `drug_id` that is wrong is a
wrong-drug answer, so any answer built on inherited focus must name what it
inherited: *"Về Metformin, ở trẻ em: …"*. This is a hard rule, not a
presentation preference.
**Focus expires.** A field older than `FOCUS_TTL_TURNS` (6) is dropped rather
than inherited. Conversations drift, and a drug from ten turns ago is not
context, it is a hazard.
### 2. Recent history and summary
- `recent` holds the last **K = 6** turns verbatim (three exchanges).
- When a turn falls out of `recent`, it is folded into `summary`.
- `summary` is regenerated at most every **S = 4** turns, capped at **400
tokens**; `recent` is capped at **2000 tokens**, oldest dropped first.
- **The summary records what was discussed, never clinical content.** It may
say *"đã hỏi liều dùng của Metformin cho người lớn"*; it may not carry a dose.
A dose restated from a summary would have no citation and could not be
grounding-verified — the check compares against retrieved evidence, and a
summary is not evidence.
### 3. Reasoning loop
```mermaid
flowchart TD
A[User turn] --> B[UNDERSTAND<br/>resolve against Focus]
B --> C{Clarify signal?}
C -->|ambiguous drug / no attribute /<br/>multi-attribute| Z[ASK — 1 turn, no loop]
C -->|no| D{Simple?}
D -->|drug + section resolved,<br/>no follow-up ambiguity| E[RETRIEVE]
D -->|complex / decomposable| P[PLAN<br/>sub-questions + retrieval set]
P --> E
E --> F[ASSESS sufficiency]
F -->|insufficient AND rounds left| R[REFINE query] --> E
F -->|sufficient OR rounds exhausted| G[GENERATE]
G --> H[VERIFY<br/>grounding + coverage]
H -->|ungrounded / off-target,<br/>repairs left| G
H -->|grounded| Y[RESPOND]
H -->|repairs exhausted| X[FALL BACK<br/>verbatim source]
F -->|exhausted AND still thin| Z
```
**Continue conditions** — a round is spent only when all hold:
1. `retrieval_rounds < MAX_RETRIEVAL_ROUNDS` (2)
2. the assessor named a *specific* missing thing (a section, a population, a
second drug) — "feels incomplete" is not a reason to spend a round
3. the refined query differs from every query already tried this turn
**Stop conditions** — any one ends the loop:
- sufficiency satisfied
- budget exhausted (rounds, LLM calls, wall-clock, tokens)
- a clarify signal fires (these bypass the loop entirely — asking beats guessing)
- grounding verification fails after `MAX_REPAIRS` (1) → extractive fallback
**Fast path.** When the drug resolves and `SectionResolver` returns a section
and no clarify signal fires, the loop is skipped: retrieve → generate → verify.
This is the majority path and it costs one LLM call.
### 4. Budgets
| Limit | Value | Enforced at |
|---|---|---|
| `MAX_RETRIEVAL_ROUNDS` | 2 | loop guard |
| `MAX_REPAIRS` | 1 | loop guard |
| `MAX_LLM_CALLS` per turn | 4 | budget object, checked before each call |
| `MAX_WALL_CLOCK_MS` | 20000 | checked between stages |
| `MAX_EVIDENCE_TOKENS` | 12000 | evidence assembly, oldest-dropped |
| `FOCUS_TTL_TURNS` | 6 | state update |
The budget is a single object threaded through the loop and **decremented
before** each call, so exhaustion degrades to the best answer so far rather
than to an error.
### 5. Integration
New domain modules, no SDK imports:
- `rag/conversation.py``Focus`, `Turn`, `ConversationState`, window and
focus-update rules. Pure; the follow-up resolution in it needs no LLM.
- `rag/reasoning.py` — the loop, its budget, and its stage protocols.
- `rag/ports.py``ConversationStore` (load/save), `Summariser`, `Planner`,
`SufficiencyAssessor`. Each has a deterministic no-LLM default so the whole
loop runs offline.
New adapter: `adapters/postgres.py` gains `PostgresConversationStore`.
Unchanged and still binding: `GroundedAnswerService` remains the single-turn
engine; `grounding.verify` gates every generated answer; `VERIFY_PDF` evidence
is never generated over.
### 6. Measurement
A capability that cannot be shown to help does not ship. Three modes are run
over the same cases — `single-shot`, `+history`, `+reasoning-loop`:
| Metric | Answers |
|---|---|
| follow-up resolution accuracy | does *"còn trẻ em thì sao?"* reach the right drug+section+population |
| on-target rate | does the answer contain the population/attribute actually asked for |
| grounding rejection rate | does reasoning make fabrication more or less likely |
| clarify rate / clarify precision | does it ask when it should, and only then |
| median + p95 latency, LLM calls, tokens per answered turn | what the capability costs |
The evaluation set is a **new multi-turn golden file** — the existing
`golden_e2e_v1.csv` is single-turn by construction and cannot measure any of
this. Counters land in `rag/metrics.py` and on the existing Grafana dashboard.
## Consequences
**Accepted.** More moving parts and more tokens per turn; a stateful service
where there was a stateless one; a summary that must be kept free of clinical
content by rule rather than by mechanism.
**Refused.** An LLM confidence score as the loop's uncertainty signal. The
signals used are the resolver states that already exist — ambiguous drug,
unresolved section, multi-attribute question — because they are deterministic,
testable, and explainable to a reviewer. "The model felt 0.73 sure" is not a
defensible basis for asking or not asking a clinician a question.
**Unchanged.** Nothing here lets the system assert a figure absent from the
retrieved source. Reasoning chooses *what to look up and how to say it*; it is
not a source of facts.