diff --git a/CLAUDE.md b/CLAUDE.md
deleted file mode 100644
index 607bee7..0000000
--- a/CLAUDE.md
+++ /dev/null
@@ -1,141 +0,0 @@
-# Instructions for Claude working in this repo
-
-## Never fabricate, never bluff
-
-Do not state a number, a test result, a "verified" claim, or a capability
-estimate unless it is backed by something you actually ran or actually
-read. If you haven't checked something, say so explicitly ("not verified
-yet", "estimate, not measured") instead of presenting a guess as fact.
-
-**Why:** this project involves parsing a medical reference book into a
-chatbot's knowledge base — false confidence here is not a cosmetic bug, it
-propagates into medical answers. During the ingestion-strategy
-investigation, a size-based heading threshold silently dropped ~15% of real
-monographs before whole-document validation caught it; a monograph-boundary
-scan was initially run on 1405 of 1668 pages before being caught and
-corrected. Confident-sounding claims that turn out wrong cost real rework
-and could cost real answer quality once this is live.
-
-**How to apply:**
-- Prefer "I ran X and got Y" over "X should work" — run the check.
-- When asked something you don't know for certain (throughput estimates,
- whether a tool/library works on this environment, whether a heuristic
- holds at scale), say what's measured vs. estimated, explicitly.
-- Whole-document / whole-scope validation over small-sample claims — if the
- user states a total (e.g. "1668 pages"), any check must cover that literal
- total before being reported as done, not a convenient subset.
-- When a claim turns out wrong after fuller checking, say so plainly and
- show the corrected result — don't quietly smooth over the miss.
-
-See `docs/pdf-parsing-outlier-catalog.md` and
-`docs/adr/0003-pdf-parsing-strategy.md` for the concrete track record this
-rule comes from.
-
-**What "verified" means:** when reporting something as verified, state (1)
-the command/test/script/manual check that was run, (2) the exact input
-scope, (3) the expected invariant or acceptance condition, (4) the observed
-result, and (5) any part of the requested scope that was *not* covered.
-Distinct scopes (unit test, regression fixture, selected-page sample,
-selected monographs, all detected monographs, full 1668-page document) are
-not interchangeable — don't describe one as another. Avoid words like
-"fully verified", "complete", "all", "no data lost", or "production-ready"
-unless the checks actually performed support that literal claim.
-
-## Real code follows Clean Code / Clean Architecture / SoC / DRY / SOLID
-
-Applies to anything meant to be committed as part of the actual system
-(`apps/*`, `ingestion/*`, `packages/*`) — not throwaway investigation
-scripts (e.g. a one-off scan to check a hypothesis), which may stay quick
-and disposable as long as they're never confused for production code and
-get deleted once their finding is written down.
-
-**Why this is a written rule, not just an intention:** intentions from one
-conversation don't carry into the next session, and under time pressure or
-mid-refactor it's easy to let a principle slip without noticing — a written
-checklist is what actually catches that, the same reasoning behind the
-"never fabricate" rule above.
-
-**How to apply, concretely, in this repo:**
-- **SoC**: keep the `ingestion/` pipeline stages (`extract/`, `segment/`,
- `chunk/`, `embed/`, `load/`) genuinely independent — extraction code must
- not know about chunking, chunking must not call OpenAI, etc.
-- **DRY**: shared logic (e.g. the bold-span heading/boundary detector) lives
- in exactly one module that both the real pipeline and any validation
- script import — never re-implemented per script, which is what happened
- during exploratory investigation and is fine there, but must not carry
- into real code.
-- **SOLID**: single-responsibility modules/classes (a detector detects, it
- doesn't also chunk); open/closed section taxonomy (adding a new section
- label — e.g. a field like "Tên thương mại" not in the book's own
- documented list — must not require editing existing matching code, only
- adding an entry); dependency inversion at infrastructure boundaries
- (`ai-service`'s domain/retrieval logic depends on an interface, not a
- hard import of the Qdrant SDK or OpenAI client directly, so it stays
- testable without live services).
-- **Clean Architecture**: domain/business logic (parsing rules, chunking
- rules, retrieval/grounding logic) stays independent of infrastructure
- (OpenAI SDK, Qdrant client, filesystem, NestJS framework details) so it's
- testable in isolation.
-- **Clean Code**: meaningful names, small functions, minimal comments (only
- where the *why* isn't obvious from the code itself) — matches the
- no-comments-unless-non-obvious style already used throughout this
- project's docs and ADRs.
-
-## Preserve provenance
-
-Every extracted or transformed unit must retain enough provenance to trace
-it back to the source document — depending on the data type, this may
-include document id, page number, source block/span id, bounding box,
-reading-order position, table id and row/column coordinates, formula
-source span, monograph id, section path, and extraction method/parser
-version.
-
-**Why:** this is a medical reference book being turned into a chatbot's
-knowledge base — if an answer is wrong, being able to trace a chunk back to
-the exact page/span it came from is how it gets debugged and corrected.
-Normalized text that "looks right" is not the same guarantee as text that
-is traceable.
-
-**How to apply:**
-- Don't discard provenance fields just because the normalized text appears
- correct — a text value that can't be traced back to its source is not a
- fully validated extraction result.
-- When adding a new pipeline stage or record type, carry existing
- provenance fields through rather than dropping them at the boundary.
-
-## Investigation scripts are evidence tools, not production code
-
-One-off investigation scripts (e.g. scanning the corpus to check a
-hypothesis) may optimize for speed, but they must:
-- be clearly named or located as temporary investigation code;
-- state or record the scope they scanned;
-- output enough information to reproduce or inspect the finding;
-- not be imported by production code, and not become the only
- implementation of a parsing rule;
-- not be cited as whole-document evidence unless they actually covered the
- whole document;
-- be deleted after their finding is captured in a regression test, fixture,
- ADR, or the outlier catalog.
-
-When an investigation uncovers a real parsing rule, move that rule into the
-production implementation and validate both the production code and the
-regression fixture — per [[DRY]] above, the rule should end up living in
-exactly one place.
-
-## Definition of done
-
-A task is not complete merely because code was written. Before reporting
-completion:
-- run the most relevant available tests and validation commands, and
- report the exact commands/checks run and whether each passed or failed;
-- state the validation scope (see "What 'verified' means" above);
-- add or update a regression fixture for each parser bug fixed;
-- confirm intended provenance fields remain present;
-- check for silent loss of expected monographs, sections, tables, formulas,
- or source references when the task could affect them;
-- avoid whole-document claims when only sample validation was performed;
-- list anything not tested, not measured, blocked, or still uncertain.
-
-If only part of the task is complete, report the completed and incomplete
-parts separately — don't hide failing tests, unexpected counts, incomplete
-coverage, or contradictory evidence to present a cleaner status.
diff --git a/apps/ai-service/adapters/embedding.py b/apps/ai-service/adapters/embedding.py
index e0630aa..30a7331 100644
--- a/apps/ai-service/adapters/embedding.py
+++ b/apps/ai-service/adapters/embedding.py
@@ -1,7 +1,5 @@
from __future__ import annotations
-import hashlib
-import math
from typing import Any
from rag.ports import QueryEmbeddingUnavailable
@@ -120,59 +118,3 @@ class BedrockCohereQueryEmbedder:
f"expected {self._dimensions}"
)
return values
-
-
-class SectionOnlyQueryEmbedder:
- """Refuses to embed, confining retrieval to the route that is verified.
-
- The section route resolves the question's attribute to a `section_key` and
- filters on it; no vector is involved, and it measured 16/16 on the
- human-written golden questions on 2026-08-04. Similarity measured 0.544 and
- its provider is currently revoked.
-
- Declining locally and immediately is better than the two alternatives it
- replaces: a `cohere-v4` round-trip spends the boto3 retry budget before
- failing, and `LocalHashQueryEmbedder` searches a SHA-256 vector against a
- Cohere collection, which returns confident and meaningless hits.
- """
-
- def __init__(self, dimensions: int) -> None:
- if dimensions <= 0:
- raise ValueError("dimensions must be positive")
- self._dimensions = dimensions
-
- @property
- def dimensions(self) -> int:
- return self._dimensions
-
- def embed_query(self, text: str) -> list[float]: # noqa: ARG002
- # The text is irrelevant: this embedder exists to refuse, not to embed.
- raise QueryEmbeddingUnavailable(
- "no query embedding provider is enabled; set EMBEDDING_PROVIDER to "
- "use the similarity fallback"
- )
-
-
-class LocalHashQueryEmbedder:
- """Deterministic local plumbing probe; not a semantic retrieval model."""
-
- def __init__(self, dimensions: int) -> None:
- if dimensions <= 0:
- raise ValueError("dimensions must be positive")
- self._dimensions = dimensions
-
- @property
- def dimensions(self) -> int:
- return self._dimensions
-
- def embed_query(self, text: str) -> list[float]:
- vector = [0.0] * self._dimensions
- for token in text.casefold().split():
- digest = hashlib.sha256(token.encode("utf-8")).digest()
- index = int.from_bytes(digest[:4], "big") % self._dimensions
- sign = 1.0 if digest[4] & 1 else -1.0
- vector[index] += sign
- norm = math.sqrt(sum(value * value for value in vector))
- if norm == 0:
- return vector
- return [value / norm for value in vector]
diff --git a/apps/ai-service/adapters/postgres.py b/apps/ai-service/adapters/postgres.py
index f94674d..6189c1f 100644
--- a/apps/ai-service/adapters/postgres.py
+++ b/apps/ai-service/adapters/postgres.py
@@ -22,6 +22,19 @@ class RetrievalTrace:
class PostgresTraceRepository:
+ """Opens a new connection per call — no pooling (F-09: a real pool, with
+ startup-time lifecycle, is a further improvement not made here).
+
+ `connect_timeout` matters more than it looks: found live 2026-08-06 that
+ an unreachable Postgres (packets dropped, not actively refused) makes a
+ bare `psycopg.connect()` hang on the OS-level TCP timeout — tens of
+ seconds, not immediate — which defeats a caller's try/except fail-open
+ around `save()` just as effectively as no try/except at all, since the
+ exception it's waiting for never arrives in time. `routers/rag.py`
+ wraps `save()` to keep a trace outage from failing an already-computed
+ answer; this bounds how long that protection can take to kick in.
+ """
+
def __init__(self, dsn: str) -> None:
self._dsn = dsn
@@ -29,7 +42,7 @@ class PostgresTraceRepository:
import psycopg
statement = migration_path.read_text(encoding="utf-8")
- with psycopg.connect(self._dsn) as connection:
+ with psycopg.connect(self._dsn, connect_timeout=5) as connection:
connection.execute(statement)
def save(
@@ -46,7 +59,7 @@ class PostgresTraceRepository:
import psycopg
trace_id = str(uuid.uuid4())
- with psycopg.connect(self._dsn) as connection:
+ with psycopg.connect(self._dsn, connect_timeout=5) as connection:
connection.execute(
"""
INSERT INTO rag_retrieval_trace (
@@ -64,7 +77,7 @@ class PostgresTraceRepository:
def get(self, trace_id: str) -> RetrievalTrace | None:
import psycopg
- with psycopg.connect(self._dsn) as connection:
+ with psycopg.connect(self._dsn, connect_timeout=5) as connection:
row = connection.execute(
"""
SELECT trace_id::text, query_text, subject_scope, query_intent,
diff --git a/apps/ai-service/adapters/prometheus.py b/apps/ai-service/adapters/prometheus.py
index 89a7b5b..c1d311d 100644
--- a/apps/ai-service/adapters/prometheus.py
+++ b/apps/ai-service/adapters/prometheus.py
@@ -31,7 +31,10 @@ _HELP = {
ABSTENTION: "Answers refused, by the reason retrieval gave.",
GENERATION_REJECTED: (
"Generations discarded before reaching the caller. `reason=\"ungrounded_number\"` "
- "counts answers that stated a figure absent from the cited source."
+ "counts answers that stated a figure absent from the cited source; "
+ "`reason=\"uncited_claim\"` counts claims with no valid citation at all; "
+ "`reason=\"unsupported_claim\"` counts claims the entailment pass judged "
+ "not actually stated by the block they cite."
),
GENERATION_SERVED: "Generations that passed grounding verification and were served.",
ANSWER_EXTRACTIVE: "Answers served as verbatim source text.",
diff --git a/apps/ai-service/bootstrap.py b/apps/ai-service/bootstrap.py
index 29ce9e7..21b176e 100644
--- a/apps/ai-service/bootstrap.py
+++ b/apps/ai-service/bootstrap.py
@@ -2,22 +2,51 @@ from __future__ import annotations
from qdrant_client import QdrantClient
-from adapters.embedding import (
- BedrockCohereQueryEmbedder,
- LocalHashQueryEmbedder,
- SectionOnlyQueryEmbedder,
-)
+from adapters.embedding import BedrockCohereQueryEmbedder
from adapters.postgres import PostgresTraceRepository
from adapters.qdrant import QdrantParentStore, QdrantRetriever
from config import Settings
+from rag.agent import RagAgent
from rag.answer import GroundedAnswerService
from rag.artifacts import load_aliases
-from rag.conversation import DeterministicSummariser, InMemoryConversationStore
-from rag.conversational import ConversationalLoopService
+from rag.manifest import MANIFEST_POINT_ID, check_manifest, manifest_collection
from rag.metrics import NullMetrics
from rag.routing import CatalogDrugResolver, QueryRoutingService
from rag.sections import SectionResolver
from rag.service import EvidencePolicy, RetrievalService
+from rag.understanding import LlmQueryUnderstander
+
+# How many aliases to show per candidate drug (F-04 bounds *which* drugs are
+# shown at all, per turn; this bounds how many names each shown one gets).
+_CATALOG_ALIASES_PER_DRUG = 3
+
+
+def _catalog_names(aliases: dict[str, set[str]]) -> dict[str, str]:
+ """The name(s) shown to the LLM for each drug_id.
+
+ Found live 2026-08-06: picking the first 3 aliases *alphabetically* could
+ drop the drug's own canonical/INN name entirely — paracetamol has 191
+ aliases (mostly trade names), and the alphabetically-first 3 were
+ "0Frezefev, ABAB, Ace kid 80", none recognizable. Mid-conversation, once
+ the current turn no longer restates the drug name in raw text, the model
+ has to re-derive it from history + this catalog line alone — with no
+ recognizable name shown, it read the earlier "paracetamol" mention as an
+ unknown drug and answered "not found in the formulary" for a drug that
+ plainly is. Fixed by always showing the drug_id's own name form first
+ (guaranteed present, always recognizable — it's the exact string a user
+ who names a drug is most likely to type), then filling remaining slots
+ with short, ALL-CAPS-preferring aliases (the book's own heading
+ convention, so usually the generic name, not a dosage-suffixed brand
+ like "Ace kid 80").
+ """
+ result: dict[str, str] = {}
+ for drug_id, names in aliases.items():
+ canonical = drug_id.replace("_", " ")
+ ranked = sorted(names, key=lambda n: (not n.isupper(), len(n), n))
+ extra = [n for n in ranked if n.strip().casefold() != canonical]
+ shown = [canonical, *extra][:_CATALOG_ALIASES_PER_DRUG]
+ result[drug_id] = ", ".join(dict.fromkeys(shown))
+ return result
def _build_metrics(settings: Settings):
@@ -73,15 +102,37 @@ def _build_reranker(settings: Settings):
return BedrockCohereReranker(region=settings.aws_region)
+def _verify_corpus_manifest(client, collection: str, embedder, settings: Settings) -> None:
+ """Read the sidecar manifest point the ingestion loader writes
+ (`ingestion/ingestion/load/manifest.py`) and refuse to start on a
+ mismatch. Raises `rag.manifest.ManifestMismatch` (a `RuntimeError`),
+ which crashes startup — a deliberate refusal, not an oversight: this
+ runs at process start (`main.py` calls `build_runtime` at import time),
+ so a mismatch here means the process never comes up and never serves a
+ query from a corpus it wasn't verified against.
+ """
+ sidecar = manifest_collection(collection)
+ payload = None
+ # This qdrant-client version has no `collection_exists`, and
+ # `get_collection` (singular) is a known parse-bug risk in this
+ # environment — list collections and check membership instead.
+ existing = {col.name for col in client.get_collections().collections}
+ if sidecar in existing:
+ points = client.retrieve(sidecar, [MANIFEST_POINT_ID], with_payload=True)
+ if points:
+ payload = points[0].payload
+ check_manifest(payload, collection, embedder.model_id, settings.embedding_dimensions)
+
+
def build_runtime(settings: Settings):
metrics = _build_metrics(settings)
if settings.embedding_provider == "disabled":
return None, None, PostgresTraceRepository(settings.postgres_dsn), metrics
- if settings.embedding_provider not in ("section-only", "local-smoke", "cohere-v4"):
+ if settings.embedding_provider != "cohere-v4":
raise ValueError(
"No production query embedder is configured. Supported values: "
- "EMBEDDING_PROVIDER=section-only (default; section route only), "
- "local-smoke (plumbing only) or cohere-v4"
+ "EMBEDDING_PROVIDER=cohere-v4 (semantic query embedding) or "
+ "disabled. The old local/section-only stubs were removed."
)
client = QdrantClient(
@@ -89,19 +140,21 @@ def build_runtime(settings: Settings):
api_key=settings.qdrant_api_key,
timeout=30,
)
- # A collection built with one model and queried with another returns hits
- # and raises nothing; the results are just meaningless. Keep this in step
- # with `model_id` in the collection's manifest.
- if settings.embedding_provider == "cohere-v4":
- embedder = BedrockCohereQueryEmbedder(
- settings.embedding_dimensions, region=settings.aws_region
- )
- elif settings.embedding_provider == "local-smoke":
- embedder = LocalHashQueryEmbedder(settings.embedding_dimensions)
- else:
- embedder = SectionOnlyQueryEmbedder(settings.embedding_dimensions)
+ embedder = BedrockCohereQueryEmbedder(
+ settings.embedding_dimensions, region=settings.aws_region
+ )
+ # F-05: a collection built with one model and queried with another
+ # returns hits and raises nothing — the results are just meaningless,
+ # with no error to notice. Refuse to start rather than search with
+ # vectors this collection was not built from.
+ _verify_corpus_manifest(client, settings.qdrant_collection, embedder, settings)
section_resolver = SectionResolver()
- resolver = CatalogDrugResolver(load_aliases(settings.entities_path))
+ aliases = load_aliases(settings.entities_path)
+ # Kept only for deterministic input-time autocomplete (`RagAgent.complete`)
+ # and to satisfy `GroundedAnswerService`'s constructor — its fuzzy
+ # `.resolve()` is no longer on the live query path; `RagAgent` resolves
+ # drug identity through `LlmQueryUnderstander` against the same catalog.
+ resolver = CatalogDrugResolver(aliases)
retrieval = RetrievalService(
QdrantRetriever(client, settings.qdrant_collection, embedder),
QdrantParentStore(client, settings.qdrant_collection),
@@ -110,19 +163,22 @@ def build_runtime(settings: Settings):
reranker=_build_reranker(settings),
)
routing = QueryRoutingService(retrieval, resolver)
+ generator = _build_generator(settings)
answers = GroundedAnswerService(
- routing, generator=_build_generator(settings), metrics=metrics or NullMetrics()
+ routing, generator=generator, metrics=metrics or NullMetrics()
)
- # The conversational layer reuses the same resolvers and the safe answer
- # engine, adding only turn understanding, follow-up inheritance and the
- # clarify/refine loop around it. InMemory store for now; a Postgres-backed
- # store is the persistence follow-up.
- conversational = ConversationalLoopService(
+ trace_writer = PostgresTraceRepository(settings.postgres_dsn)
+ if generator is None:
+ # The new front end understands a turn with the same LLM call that
+ # answers it — with no generator configured there is no query
+ # understanding either, so there is no conversational/agent
+ # capability to offer. Answer-only (retrieval-verified, no
+ # generation) still works through `answers` directly.
+ return answers, None, trace_writer, metrics
+ agent = RagAgent(
+ understander=LlmQueryUnderstander(generator, _catalog_names(aliases), resolver),
+ retrieval=retrieval,
answers=answers,
- resolver=resolver,
- section_resolver=section_resolver,
- store=InMemoryConversationStore(),
- summariser=DeterministicSummariser(),
- metrics=metrics or NullMetrics(),
+ autocomplete=resolver,
)
- return answers, conversational, PostgresTraceRepository(settings.postgres_dsn), metrics
+ return answers, agent, trace_writer, metrics
diff --git a/apps/ai-service/config.py b/apps/ai-service/config.py
index 23bcabd..3b58e05 100644
--- a/apps/ai-service/config.py
+++ b/apps/ai-service/config.py
@@ -18,9 +18,10 @@ class Settings(BaseSettings):
default="postgresql://duoc_thu:duoc_thu@localhost:5432/duoc_thu",
repr=False,
)
- # Defaults to the route that is measured and needs no provider. Raising it
- # to `cohere-v4` enables the similarity fallback and requires live Bedrock.
- embedding_provider: str = "section-only"
+ # `cohere-v4` = semantic query embedding in the corpus's own space (requires
+ # live Bedrock). `disabled` skips retrieval entirely. The old local-hash /
+ # section-only stub embedders were removed in the 2026-08-06 rebuild.
+ embedding_provider: str = "cohere-v4"
embedding_dimensions: int = 1024
evidence_minimum_score: float = 0.12
aws_region: str = "us-east-1"
diff --git a/apps/ai-service/rag/agent.py b/apps/ai-service/rag/agent.py
new file mode 100644
index 0000000..5471495
--- /dev/null
+++ b/apps/ai-service/rag/agent.py
@@ -0,0 +1,194 @@
+"""The new RAG orchestrator — LLM understanding in, grounded answer out.
+
+Replaces the old front-end wholesale:
+- `CatalogDrugResolver` (fuzzy) + `SectionResolver` (keyword) -> `understanding.py`
+- `ConversationalLoopService` + `conversation.py` (Focus / Summariser / manual
+ follow-up inheritance) -> the LLM reads a plain turn history and resolves
+ "thuốc đó" / "còn liều thì sao" itself.
+
+What is deliberately KEPT because it is the safety spine, not the brittle part:
+- `RetrievalService.retrieve_framed` (Qdrant section/overview retrieval, whole
+ section, provenance, quarantine `VERIFY_PDF`),
+- `GroundedAnswerService.answer_from_result` (`grounding.verify` + entailment
+ on every generated claim; a configured generator that fails abstains rather
+ than degrading to a raw source dump).
+
+This module owns routing only; it states no medical fact of its own.
+"""
+from __future__ import annotations
+
+from dataclasses import dataclass, field
+from typing import Protocol
+
+from .answer import Citation, GroundedAnswerService
+from .models import EvidenceDecision, RetrievalResult
+from .policy import looks_non_human
+from .service import RetrievalService
+from .understanding import QueryFrame, QueryUnderstander
+
+TUONG_TAC = "tuong_tac_thuoc"
+HISTORY_TURNS = 6
+
+
+class AutocompleteSource(Protocol):
+ """As-you-type suggestion, kept deterministic and independent of the LLM
+ understander — a prefix match needs no model call."""
+
+ def complete(self, prefix: str, k: int) -> list[str]: ...
+
+
+@dataclass(frozen=True)
+class AgentReply:
+ decision: str # answerable | abstain | clarify | verify_pdf
+ reason: str
+ answer: str | None = None
+ clarification: str | None = None
+ citations: tuple[Citation, ...] = ()
+ drugs: tuple[str, ...] = ()
+ turn_type: str = ""
+ generated: bool = False
+
+
+class RagAgent:
+ def __init__(
+ self,
+ understander: QueryUnderstander,
+ retrieval: RetrievalService,
+ answers: GroundedAnswerService,
+ autocomplete: AutocompleteSource | None = None,
+ history_turns: int = HISTORY_TURNS,
+ ) -> None:
+ self._understander = understander
+ self._retrieval = retrieval
+ self._answers = answers
+ self._autocomplete = autocomplete
+ self._history_turns = history_turns
+ self._history: dict[str, list[str]] = {}
+
+ def complete(self, prefix: str, k: int = 8) -> list[str]:
+ """Display names matching a typed prefix, for input autocomplete."""
+ if self._autocomplete is None:
+ return []
+ return [_display_name(drug_id) for drug_id in self._autocomplete.complete(prefix, k)]
+
+ def handle(self, turn: str, conversation_id: str | None = None) -> AgentReply:
+ history = self._history.get(conversation_id, []) if conversation_id else []
+ frame = self._understander.understand(turn, tuple(history))
+ reply = self._route(turn, frame)
+ if conversation_id is not None:
+ self._remember(conversation_id, turn, reply)
+ return reply
+
+ def _route(self, turn: str, frame: QueryFrame) -> AgentReply:
+ tt = frame.turn_type
+
+ if frame.needs_clarify and frame.clarify_reason:
+ return AgentReply("clarify", "needs_more_info",
+ clarification=frame.clarify_reason,
+ drugs=frame.drugs, turn_type=tt)
+
+ if tt == "smalltalk":
+ return AgentReply(
+ "answerable", "smalltalk",
+ answer="Chào anh/chị! Em là trợ lý tra cứu Dược thư Quốc gia Việt Nam "
+ "2018, sẵn sàng hỗ trợ tra liều dùng, chống chỉ định, tương tác "
+ "thuốc... Anh/chị đang cần tra thuốc nào ạ?",
+ turn_type=tt)
+
+ if tt in ("out_of_scope",) or looks_non_human(turn):
+ return AgentReply(
+ "abstain", "out_of_scope",
+ answer="Nội dung này nằm ngoài phần chuyên luận thuốc của Dược thư "
+ "(có thể thuộc phần hướng dẫn chung/phụ lục chưa được đưa vào). "
+ "Tôi chưa có dữ liệu để trả lời chính xác.",
+ turn_type=tt)
+
+ if not frame.drugs:
+ if frame.unknown_drugs:
+ names = ", ".join(frame.unknown_drugs)
+ return AgentReply(
+ "abstain", "drug_not_in_formulary",
+ answer=f"Không tìm thấy \"{names}\" trong Dược thư Quốc gia Việt Nam.",
+ turn_type=tt)
+ if tt == "symptom_to_drug":
+ # Reverse lookup (indication/adverse-effect -> drugs) is a distinct
+ # retrieval mode, not yet wired. Be honest rather than abstain blank.
+ return AgentReply(
+ "clarify", "reverse_lookup_not_ready",
+ clarification="Tra ngược theo triệu chứng/chỉ định đang được bổ "
+ "sung. Anh/chị cho biết tên thuốc cụ thể để tôi tra giúp?",
+ turn_type=tt)
+ return AgentReply(
+ "clarify", "no_drug",
+ clarification="Anh/chị muốn tra thuốc nào?", turn_type=tt)
+
+ if tt == "interaction" and len(frame.drugs) >= 2:
+ return self._interaction(turn, frame)
+
+ # drug_attribute / drug_overview / dosing_calc / fallback: one drug + section
+ return self._single_drug(turn, frame)
+
+ def _single_drug(self, turn: str, frame: QueryFrame) -> AgentReply:
+ result = self._retrieval.retrieve_framed(
+ frame.drugs[0], frame.attribute, turn,
+ is_overview=frame.turn_type == "drug_overview",
+ )
+ return self._grounded(turn, result, frame)
+
+ def _interaction(self, turn: str, frame: QueryFrame) -> AgentReply:
+ """Gather the interaction section of each named drug and synthesise.
+
+ Absence of a match is stated as "not found in each drug's interaction
+ section", never as "safe" — the answer layer's grounding still applies.
+ """
+ evidences = []
+ for drug_id in frame.drugs:
+ part = self._retrieval.retrieve_framed(drug_id, TUONG_TAC, turn)
+ if part.decision == EvidenceDecision.ANSWERABLE:
+ evidences.extend(part.evidence)
+ if not evidences:
+ listed = " và ".join(frame.drugs)
+ return AgentReply(
+ "abstain", "no_interaction_evidence",
+ answer=f"Không tìm thấy mục tương tác thuốc cho {listed} trong Dược "
+ "thư. Điều này KHÔNG có nghĩa là an toàn khi phối hợp.",
+ drugs=frame.drugs, turn_type=frame.turn_type)
+ combined = RetrievalResult(
+ EvidenceDecision.ANSWERABLE, "interaction_evidence",
+ tuple(evidences),
+ )
+ return self._grounded(turn, combined, frame)
+
+ def _grounded(
+ self, turn: str, result: RetrievalResult, frame: QueryFrame
+ ) -> AgentReply:
+ ga = self._answers.answer_from_result(turn, result)
+ decision = ga.result.decision.value
+ if ga.clarification is not None:
+ decision = "clarify"
+ return AgentReply(
+ decision=decision,
+ reason=ga.result.reason,
+ answer=ga.answer,
+ clarification=ga.clarification,
+ citations=ga.citations,
+ drugs=frame.drugs,
+ turn_type=frame.turn_type,
+ generated=ga.generated,
+ )
+
+ def _remember(self, conversation_id: str, turn: str, reply: AgentReply) -> None:
+ history = self._history.setdefault(conversation_id, [])
+ history.append(f"Người dùng: {turn}")
+ spoken = reply.answer or reply.clarification
+ if spoken:
+ history.append(f"Trợ lý: {spoken[:300]}")
+ # Keep only the recent window; the LLM re-reads it every turn.
+ excess = len(history) - self._history_turns * 2
+ if excess > 0:
+ del history[:excess]
+
+
+def _display_name(drug_id: str) -> str:
+ """A readable display name from a drug id ('paracetamol_acetaminophen')."""
+ return drug_id.replace("_", " ").title()
diff --git a/apps/ai-service/rag/answer.py b/apps/ai-service/rag/answer.py
index 20aac9f..e44a076 100644
--- a/apps/ai-service/rag/answer.py
+++ b/apps/ai-service/rag/answer.py
@@ -8,7 +8,7 @@ from . import grounding, metrics as metric_names
from .metrics import Metrics, NullMetrics
from .models import EvidenceDecision, QueryIntent, RetrievalResult, SubjectScope
from .ports import AnswerGenerationUnavailable, AnswerGenerator
-from .prompt import build_request
+from .prompt import build_entailment_request, build_request, build_sufficiency_request
from .routing import QueryRoutingService
@@ -30,17 +30,38 @@ class GroundedAnswer:
answer: str | None
citations: tuple[Citation, ...] = ()
generated: bool = False
+ # Set when the model decided the turn is under-specified and asked back
+ # (e.g. a dose question with no age/weight). The answer field carries the
+ # question; the caller renders it as a clarification, not a final answer.
+ clarification: str | None = None
+
+
+@dataclass(frozen=True)
+class _GenOutcome:
+ answer: str | None = None
+ clarification: str | None = None
class GroundedAnswerService:
"""Retrieval decides what is true; generation only decides how it reads.
- When a generator is configured, its output replaces the extractive text
- **only** if `grounding.verify` confirms every figure and citation in it
- traces back to the retrieved evidence. Anything else — an unsupported
- number, a citation to nothing, a provider outage, malformed output — falls
- back to quoting the source verbatim, which is always available because it
- was computed first.
+ Two operating modes, not to be confused with each other:
+
+ - **No generator configured** (`generator=None`, e.g. `ANSWER_PROVIDER=
+ disabled`) is retrieval-only mode, a deliberate and fully supported
+ way to run this service. It quotes the retrieved source verbatim.
+ - **A generator IS configured.** Its output replaces the extractive text
+ only if it clears two independent checks: `grounding.verify` (every
+ figure and citation traces to the specific evidence block it cites,
+ and every claim carries one) and `_verify_entailment` (a second LLM
+ pass confirming each cited claim's *content* — not just its numbers —
+ is actually stated by that block). If a configured generation fails
+ any check, or the provider itself is unreachable, or its output is
+ malformed, the turn **abstains** (`reason="generation_unavailable"`
+ or the specific `grounding.verify` reason) rather than silently
+ degrading to a raw source dump — this product is a real LLM chatbot,
+ and a citation-stapled paragraph of book text is not an acceptable
+ stand-in for an answer the model was supposed to produce.
"""
def __init__(
@@ -70,6 +91,15 @@ class GroundedAnswerService:
)
else:
result = self._routing.retrieve(query, subject_scope, intent)
+ return self.answer_from_result(query, result)
+
+ def answer_from_result(
+ self, query: str, result: RetrievalResult
+ ) -> GroundedAnswer:
+ """Everything after retrieval — grounding, sufficiency, generation,
+ citations. Split out so the new understanding-driven orchestrator
+ (`rag/agent.py`) reuses the safe answer path without going through the
+ old `QueryRoutingService` text resolution."""
if result.decision == EvidenceDecision.ABSTAIN:
self._metrics.increment(metric_names.ABSTENTION, reason=result.reason)
return GroundedAnswer(result, None)
@@ -102,25 +132,62 @@ class GroundedAnswerService:
f"{text} [{index}]" for index, text in enumerate(evidence_texts, start=1)
)
- generated = self._generate(query, evidence_texts, intro=result.is_drug_overview)
- answer_text = extractive if generated is None else generated
+ # Reasoning step BEFORE answering: if the turn is under-specified (a dose
+ # with several bands and no age/weight/condition), ask instead of dumping.
+ # A separate focused call is more reliable than folding it into generation.
+ clarify_q = self._check_sufficiency(query, evidence_texts, result.is_drug_overview)
+ if clarify_q is not None:
+ return GroundedAnswer(result, clarify_q, (), clarification=clarify_q)
+
+ outcome = self._generate(query, evidence_texts, intro=result.is_drug_overview)
+ if outcome.clarification is not None:
+ # The model judged the turn under-specified (a dose with no
+ # age/weight/renal-function/indication…) and asked back instead of
+ # listing every band. Return the question, not the whole section.
+ return GroundedAnswer(
+ result, outcome.clarification, (), clarification=outcome.clarification
+ )
+
+ if outcome.answer is None:
+ if self._generator is None:
+ # No generator configured at all — retrieval-only mode. A
+ # deliberate operating mode (e.g. ANSWER_PROVIDER=disabled),
+ # not a failure, so the source is quoted verbatim.
+ self._metrics.increment(metric_names.ANSWER_EXTRACTIVE)
+ citations = self._cited_only(indexed, extractive) or all_citations
+ return GroundedAnswer(result, extractive, citations)
+ # A generator WAS configured and this specific generation did not
+ # clear the safety checks (provider outage, malformed output, an
+ # ungrounded/uncited/unsupported claim). This product is a real
+ # LLM chatbot, not the retired offline-extractive build — a raw
+ # source dump is not an acceptable stand-in for a failed
+ # generation, so this abstains instead of silently degrading to
+ # one.
+ self._metrics.increment(
+ metric_names.ABSTENTION, reason="generation_unavailable"
+ )
+ return GroundedAnswer(
+ replace(
+ result,
+ decision=EvidenceDecision.ABSTAIN,
+ reason="generation_unavailable",
+ ),
+ None,
+ )
+
# Show only the sources the answer actually cited, not every chunk that
# was retrieved — a paragraph that cites [4] must not drag 13 citation
# chips onto the screen. Falls back to all when the text cites nothing.
- citations = self._cited_only(indexed, answer_text) or all_citations
- if generated is None:
- self._metrics.increment(metric_names.ANSWER_EXTRACTIVE)
- return GroundedAnswer(result, extractive, citations)
-
+ citations = self._cited_only(indexed, outcome.answer) or all_citations
self._metrics.increment(metric_names.GENERATION_SERVED)
- return GroundedAnswer(result, generated, citations, generated=True)
+ return GroundedAnswer(result, outcome.answer, citations, generated=True)
def _generate(
self, query: str, evidence_texts: tuple[str, ...], intro: bool = False
- ) -> str | None:
- """A verified generation, or None to fall back to the source text."""
+ ) -> "_GenOutcome":
+ """A verified generation, a clarifying question, or empty to fall back."""
if self._generator is None or not evidence_texts:
- return None
+ return _GenOutcome()
request = build_request(query, evidence_texts, intro=intro)
try:
@@ -129,7 +196,7 @@ class GroundedAnswerService:
self._metrics.increment(
metric_names.GENERATION_REJECTED, reason="provider_unavailable"
)
- return None
+ return _GenOutcome()
try:
payload = json.loads(raw)
@@ -139,28 +206,123 @@ class GroundedAnswerService:
self._metrics.increment(
metric_names.GENERATION_REJECTED, reason="malformed_output"
)
- return None
+ return _GenOutcome()
+
+ # The model asked for a missing detail (age/weight/renal function/
+ # indication…) instead of listing every band. A clarify is not a grounded
+ # claim, so it skips the number check — it states no dose.
+ clarify = payload.get("clarifying_question") if isinstance(payload, dict) else None
+ if isinstance(clarify, str) and clarify.strip():
+ return _GenOutcome(clarification=clarify.strip())
if not isinstance(answer, str) or not isinstance(sufficient, bool):
self._metrics.increment(
metric_names.GENERATION_REJECTED, reason="malformed_output"
)
- return None
+ return _GenOutcome()
if not sufficient:
# The model says the evidence does not answer the question. Showing
# the retrieved section verbatim lets the clinician judge that.
self._metrics.increment(
metric_names.GENERATION_REJECTED, reason="evidence_insufficient"
)
- return None
+ return _GenOutcome()
report = grounding.verify(answer, evidence_texts)
if not report.grounded:
self._metrics.increment(
metric_names.GENERATION_REJECTED, reason=report.reason
)
+ return _GenOutcome()
+
+ if not self._verify_entailment(answer, evidence_texts):
+ self._metrics.increment(
+ metric_names.GENERATION_REJECTED, reason="unsupported_claim"
+ )
+ return _GenOutcome()
+ return _GenOutcome(answer=answer)
+
+ def _verify_entailment(self, answer: str, evidence_texts: tuple[str, ...]) -> bool:
+ """A second, adversarial LLM pass over an answer that already passed
+ `grounding.verify`.
+
+ The regex check above only binds numbers and citation indices — it
+ has no notion of meaning, so "Metformin chữa ung thư [1]" citing an
+ evidence block about "điều trị đái tháo đường" sails through it
+ untouched: right drug, right citation shape, fabricated indication.
+ This call is what catches that: each substantive, validly-cited claim
+ is checked against only the evidence block(s) it names, by a model
+ told to compare wording, not to reason about medicine.
+
+ Fails closed on an outage or malformed output. A single rejection is
+ NOT: live probing (2026-08-06) found the judge call itself is noisy
+ — the identical claim/evidence pair, called three times, came back
+ entailed twice and rejected once, discarding a correct, well-cited
+ interaction answer. So a reject triggers one same-claim retry, and
+ only a second, agreeing reject discards the generation; a single
+ provider outage/malformed response still fails closed immediately
+ (that failure mode is reliable, not noisy — no retry needed there).
+ An answer with no claim text at all (nothing between or after its
+ citation markers) is vacuously fine — nothing to verify, no call.
+ """
+ claims = [
+ (claim.text, "\n".join(evidence_texts[i - 1] for i in claim.indices))
+ for claim in grounding.split_claims(answer, len(evidence_texts))
+ if claim.indices and grounding.has_content(claim.text)
+ ]
+ if not claims:
+ return True
+
+ request = build_entailment_request(claims)
+ first = self._run_entailment_check(request)
+ if first is None:
+ return False
+ if first:
+ return True
+ second = self._run_entailment_check(request)
+ return bool(second)
+
+ def _run_entailment_check(self, request) -> bool | None:
+ """One entailment call. `None` = outage/malformed (fails closed by the
+ caller without a retry); `True`/`False` = the judge's verdict."""
+ try:
+ raw = self._generator.generate(request.system, request.user, request.schema)
+ except AnswerGenerationUnavailable:
return None
- return answer
+ try:
+ payload = json.loads(raw)
+ entailed = payload["entailed"]
+ unsupported = payload["unsupported"]
+ except (ValueError, TypeError, KeyError):
+ return None
+ if not isinstance(entailed, bool) or not isinstance(unsupported, list):
+ return None
+ return entailed and not unsupported
+
+ def _check_sufficiency(
+ self, query: str, evidence_texts: tuple[str, ...], intro: bool = False
+ ) -> str | None:
+ """A focused reasoning call: is the turn specific enough to answer, or
+ must we ask? Returns a clarifying question, or None to proceed.
+
+ Skipped without a model, for a bare-name intro (not a dose), or for a
+ single evidence block (nothing to disambiguate)."""
+ if self._generator is None or intro or len(evidence_texts) < 2:
+ return None
+ request = build_sufficiency_request(query, evidence_texts)
+ try:
+ raw = self._generator.generate(request.system, request.user, request.schema)
+ except AnswerGenerationUnavailable:
+ return None
+ try:
+ payload = json.loads(raw)
+ except (ValueError, TypeError):
+ return None
+ if isinstance(payload, dict) and payload.get("sufficient") is False:
+ question = payload.get("clarifying_question")
+ if isinstance(question, str) and question.strip():
+ return question.strip()
+ return None
@staticmethod
def _cited_only(
diff --git a/apps/ai-service/rag/conversation.py b/apps/ai-service/rag/conversation.py
index c6c35b1..83207ce 100644
--- a/apps/ai-service/rag/conversation.py
+++ b/apps/ai-service/rag/conversation.py
@@ -87,23 +87,47 @@ class ConversationState:
summary: str = ""
focus: Focus = field(default_factory=Focus)
turn_count: int = 0
+ # Turns evicted from `recent` since the last time `overflow()` was
+ # consumed and cleared (by the caller passing `pending_overflow=()` to
+ # `replace()` after folding them into the summary). NOT derivable from
+ # `recent` alone — `recent` is already capped at `window`, so comparing
+ # its length against `window` can never find anything (see the bug note
+ # on `overflow` below). Plumbing, not conversation content.
+ pending_overflow: tuple[Turn, ...] = ()
def append(self, turn: Turn, window: int = RECENT_TURNS) -> "ConversationState":
"""Adds a turn and evicts the oldest beyond the window.
- Eviction returns the dropped turns to the caller's summariser via
- `overflow`, rather than discarding them here — this type does not
- decide what a summary says.
+ Eviction accumulates the dropped turns into `pending_overflow` for
+ the caller's summariser to fold via `overflow()`, rather than
+ discarding them here — this type does not decide what a summary
+ says. Accumulates rather than overwrites because one turn commonly
+ triggers two `append()` calls in a row (user, then assistant); each
+ can evict at most one turn, and the second call must not lose the
+ first's.
"""
- recent = (*self.recent, turn)[-window:]
+ combined = (*self.recent, turn)
+ recent = combined[-window:]
+ dropped = combined[:-window] if len(combined) > window else ()
return replace(
self,
recent=recent,
turn_count=self.turn_count + 1,
+ pending_overflow=(*self.pending_overflow, *dropped),
)
- def overflow(self, window: int = RECENT_TURNS) -> tuple[Turn, ...]:
- return self.recent[:-window] if len(self.recent) > window else ()
+ def overflow(self) -> tuple[Turn, ...]:
+ """Turns evicted from `recent` and not yet folded into the summary.
+
+ Bug fixed 2026-08-06 (Codex review, F-06): this used to check
+ `len(self.recent) > window`, but `recent` is already truncated to
+ `window` by every `append()` call, so that comparison could never be
+ true — dropped turns were silently discarded and the summariser
+ never received them, no matter how long a conversation ran. The
+ caller must clear `pending_overflow` (pass `pending_overflow=()` to
+ `replace()`) after folding, or the same turns fold again next time.
+ """
+ return self.pending_overflow
def inherited(self, name: str):
"""A focus value only if it is still fresh; otherwise None."""
diff --git a/apps/ai-service/rag/conversational.py b/apps/ai-service/rag/conversational.py
index ba4fc15..eb71e69 100644
--- a/apps/ai-service/rag/conversational.py
+++ b/apps/ai-service/rag/conversational.py
@@ -343,10 +343,31 @@ class ConversationalLoopService:
query, subject_scope, intent, drug_id=resolved.drug_id
)
+ # The engine asked for a missing detail (age/weight/renal function/
+ # indication…) rather than dumping every dose band — surface it as a
+ # clarification, not a final answer.
+ if grounded is not None and grounded.clarification is not None:
+ self._metrics.increment(
+ metric_names.CLARIFY_ASKED, reason="needs_more_info"
+ )
+ self._persist(state, resolved, None)
+ return ConversationTurnResult(
+ None,
+ Clarification(
+ reason="needs_more_info",
+ question=grounded.clarification,
+ options=(),
+ ),
+ None,
+ False,
+ None,
+ "needs_more_info",
+ )
+
answer = grounded.answer if grounded else None
inherited = resolved.drug_id if resolved.needs_carry_over_notice else None
if answer is not None and inherited is not None:
- answer = f"Về {inherited}: {answer}"
+ answer = f"Về {self._drug_name(inherited)}: {answer}"
if grounded is not None:
grounded = replace(grounded, answer=answer)
@@ -395,5 +416,9 @@ class ConversationalLoopService:
and state.overflow()
):
summary = self._summariser.fold(state.summary, state.overflow())
- state = replace(state, summary=summary)
+ # Clear what was just folded — `replace()` keeps every field not
+ # named here, and `pending_overflow` accumulates across calls
+ # (see `ConversationState.append`), so leaving it would fold the
+ # same already-summarised turns again next time.
+ state = replace(state, summary=summary, pending_overflow=())
self._store.save(state)
diff --git a/apps/ai-service/rag/grounding.py b/apps/ai-service/rag/grounding.py
index 986e105..c838c0a 100644
--- a/apps/ai-service/rag/grounding.py
+++ b/apps/ai-service/rag/grounding.py
@@ -2,10 +2,25 @@
The answer layer may only rephrase retrieved text. This module is what makes
that a checkable property rather than a promise in a prompt: it recomputes,
-from the evidence alone, whether every number and every citation in a
-generated answer can be traced back to the source. A generation that fails is
+from the evidence alone, whether every claim in a generated answer traces
+back to the specific source block it cites. A generation that fails is
discarded, never shown.
+Binding is per-citation, not global. The answer is split at each citation
+marker group (one or more consecutive `[n]`); the text immediately before a
+group is that group's claim, and only the evidence block(s) named in that
+group may support it. A number that is true of evidence block 2 does not
+make a claim citing `[1]` grounded — the old implementation pooled every
+number from every evidence block into one set, which let a number attributed
+to the wrong source pass silently. `evidence_texts` is positional: `[n]`
+refers to `evidence_texts[n - 1]`.
+
+A claim with no valid citation group is rejected outright — a citation
+nobody can follow is not a citation, and an uncited clinical statement is not
+verifiable, numeric or not. This catches a missing-citation defect that the
+old check never looked for at all (it only ever checked numbers already
+carrying a marker).
+
Numbers are compared **character for character**, deliberately. "7,5" and
"7.5" are not treated as equal, and no attempt is made to parse either into a
quantity. Parsing invites the one error that matters most here: `1.500` is
@@ -13,6 +28,13 @@ quantity. Parsing invites the one error that matters most here: `1.500` is
separators maps "7,5" and "75" to the same key — a tenfold dose error scored
as a match. The model is told to copy figures verbatim, so an exact match is
achievable, and every deviation from it is refused rather than interpreted.
+
+What this module still cannot do: confirm that a citation-bearing nonnumeric
+claim is actually *entailed* by the block it cites (e.g. "chữa ung thư [1]"
+where evidence 1 is only about "điều trị đái tháo đường" — same drug name,
+unrelated indication). Regex-level number/citation checking has no notion of
+semantic content. That gap is closed separately by an LLM entailment pass
+(`rag/answer.py`'s post-generation verifier call), not by this module.
"""
from __future__ import annotations
@@ -27,12 +49,23 @@ _NUMBER = re.compile(r"\d+(?:[.,]\d+)*")
# never mistaken for the quantity 2.
_CITATION = re.compile(r"\[(\d+)\]")
+# One or more consecutive markers ("[1]", "[1][2]") count as a single group:
+# the prompt allows citing more than one source for one claim, and each is
+# checked against the union of just those sources, not all evidence.
+_CITATION_GROUP = re.compile(r"(?:\[\d+\])+")
+
+# Any word character — letter (Vietnamese diacritics included) or digit —
+# used to tell "claim with actual content" apart from bare punctuation or
+# whitespace trailing a citation, which needs no citation of its own.
+_LETTER = re.compile(r"\w", re.UNICODE)
+
@dataclass(frozen=True)
class GroundingReport:
grounded: bool
unsupported_numbers: tuple[str, ...]
invalid_citations: tuple[int, ...]
+ uncited_claim: bool
cited_indices: tuple[int, ...]
@property
@@ -41,6 +74,8 @@ class GroundingReport:
return "ungrounded_number"
if self.invalid_citations:
return "invalid_citation"
+ if self.uncited_claim:
+ return "uncited_claim"
return "grounded"
@@ -53,31 +88,90 @@ def citations_in(text: str) -> tuple[int, ...]:
return tuple(int(marker) for marker in _CITATION.findall(text))
-def verify(answer: str, evidence_texts: tuple[str, ...]) -> GroundingReport:
- """Whether `answer` states only figures and sources present in evidence.
+def has_content(text: str) -> bool:
+ """True once `text` carries any letter or digit — i.e. more than
+ punctuation or whitespace left over between/after citation markers."""
+ return _LETTER.search(text) is not None
- `evidence_texts` is positional: citation `[n]` refers to
- `evidence_texts[n - 1]`, so an out-of-range marker is a defect even when
- the prose around it is faithful — a citation nobody can follow is not a
- citation.
+
+@dataclass(frozen=True)
+class Claim:
+ """One citation-bounded segment of an answer: the text before a citation
+ group, and the (in-range) evidence indices that group names.
+
+ `indices` is empty for the trailing segment after the last citation
+ group, or for a claim whose only marker(s) were out of range — in both
+ cases there is no evidence block left to check the claim against.
"""
- source_numbers = set()
- for text in evidence_texts:
- source_numbers.update(numbers_in(text))
- unsupported = tuple(
- token for token in numbers_in(answer) if token not in source_numbers
- )
- invalid = tuple(
- index
- for index in citations_in(answer)
- if not 1 <= index <= len(evidence_texts)
- )
- cited = tuple(sorted({index for index in citations_in(answer)} - set(invalid)))
+ text: str
+ indices: tuple[int, ...]
+
+
+def split_claims(answer: str, evidence_count: int) -> tuple[Claim, ...]:
+ """The claim segmentation `verify` checks numbers against, exposed so a
+ semantic entailment pass can run the same per-claim binding — each claim
+ checked only against the evidence block(s) it actually cites, never the
+ whole evidence set.
+ """
+ claims: list[Claim] = []
+ cursor = 0
+ for group in _CITATION_GROUP.finditer(answer):
+ text = answer[cursor:group.start()]
+ cursor = group.end()
+ indices = tuple(
+ i for i in citations_in(group.group(0)) if 1 <= i <= evidence_count
+ )
+ claims.append(Claim(text, indices))
+ claims.append(Claim(answer[cursor:], ()))
+ return tuple(claims)
+
+
+def verify(answer: str, evidence_texts: tuple[str, ...]) -> GroundingReport:
+ """Whether `answer` states only figures and sources traceable to the
+ specific evidence block(s) cited immediately after each claim.
+
+ See module docstring for the binding rule and its known limit (no
+ semantic entailment check).
+ """
+ unsupported: list[str] = []
+ invalid: list[int] = []
+ uncited = False
+ cited_all: set[int] = set()
+
+ cursor = 0
+ for group in _CITATION_GROUP.finditer(answer):
+ claim = answer[cursor:group.start()]
+ cursor = group.end()
+
+ indices = citations_in(group.group(0))
+ bad = [i for i in indices if not 1 <= i <= len(evidence_texts)]
+ good = [i for i in indices if i not in bad]
+ invalid.extend(bad)
+ cited_all.update(good)
+
+ claim_numbers = numbers_in(claim)
+ if good:
+ source_numbers: set[str] = set()
+ for index in good:
+ source_numbers.update(numbers_in(evidence_texts[index - 1]))
+ unsupported.extend(n for n in claim_numbers if n not in source_numbers)
+ else:
+ # Every marker in this group was out of range: nothing to bind
+ # the claim to, numeric or not.
+ unsupported.extend(claim_numbers)
+ if has_content(claim):
+ uncited = True
+
+ tail = answer[cursor:]
+ unsupported.extend(numbers_in(tail))
+ if has_content(tail):
+ uncited = True
return GroundingReport(
- grounded=not unsupported and not invalid,
- unsupported_numbers=unsupported,
- invalid_citations=invalid,
- cited_indices=cited,
+ grounded=not unsupported and not invalid and not uncited,
+ unsupported_numbers=tuple(unsupported),
+ invalid_citations=tuple(invalid),
+ uncited_claim=uncited,
+ cited_indices=tuple(sorted(cited_all)),
)
diff --git a/apps/ai-service/rag/manifest.py b/apps/ai-service/rag/manifest.py
new file mode 100644
index 0000000..390e78d
--- /dev/null
+++ b/apps/ai-service/rag/manifest.py
@@ -0,0 +1,62 @@
+"""F-05: refuse to become ready on a corpus/model manifest mismatch.
+
+Two different embedding models can produce vectors of the same
+dimensionality; Qdrant returns plausible-looking but meaningless nearest
+neighbours with no error at query time — a stale or wrong collection is
+otherwise invisible until a clinician notices the answers are subtly off.
+The ingestion loader already writes a sidecar manifest recording what a
+collection was built from (`ingestion/ingestion/load/manifest.py`); this is
+the query-time half — compare it against the configured query embedder
+*before* serving anything, not after a bad answer is reported.
+"""
+from __future__ import annotations
+
+MANIFEST_POINT_ID = "00000000-0000-5000-8000-000000000001"
+
+
+class ManifestMismatch(RuntimeError):
+ """The configured query embedder does not match what the collection was
+ built from. Raised at startup so the service refuses to become ready
+ rather than search with mismatched vectors."""
+
+
+def manifest_collection(name: str) -> str:
+ return f"{name}__manifest"
+
+
+def check_manifest(
+ payload: dict | None,
+ collection: str,
+ expected_model_id: str,
+ expected_dimensions: int,
+) -> None:
+ """Raises `ManifestMismatch` unless `payload` (the manifest sidecar
+ point's payload, or `None` if the sidecar/point is missing entirely)
+ matches the configured query embedder.
+
+ A collection with no manifest at all is refused for the same reason a
+ mismatched one is: nothing can be said about what it was built from, and
+ "probably fine" is not a load-bearing claim for a medical formulary.
+ """
+ if payload is None:
+ raise ManifestMismatch(
+ f"{collection!r} has no corpus manifest "
+ f"({manifest_collection(collection)!r}) — refusing to query an "
+ "unattested corpus."
+ )
+ mismatches = []
+ if payload.get("model_id") != expected_model_id:
+ mismatches.append(
+ f"model_id: corpus={payload.get('model_id')!r} "
+ f"query_embedder={expected_model_id!r}"
+ )
+ if payload.get("dimensions") != expected_dimensions:
+ mismatches.append(
+ f"dimensions: corpus={payload.get('dimensions')!r} "
+ f"query_embedder={expected_dimensions!r}"
+ )
+ if mismatches:
+ raise ManifestMismatch(
+ f"{collection!r}'s corpus manifest does not match the configured "
+ "query embedder — " + "; ".join(mismatches)
+ )
diff --git a/apps/ai-service/rag/metrics.py b/apps/ai-service/rag/metrics.py
index 19890cc..15feaf7 100644
--- a/apps/ai-service/rag/metrics.py
+++ b/apps/ai-service/rag/metrics.py
@@ -54,3 +54,9 @@ LOOP_ROUNDS = "duocthu_loop_retrieval_rounds_total"
LOOP_REFINED = "duocthu_loop_refined_total"
LOOP_REPAIRED = "duocthu_loop_repaired_total"
FOLLOWUP_INHERITED = "duocthu_followup_inherited_total"
+
+# Trace persistence is fail-open (F-09): a Postgres outage must not turn an
+# already-computed, safe answer into a 500. This counts how often that
+# degradation actually happens, since a silent fail-open with no counter is
+# indistinguishable from tracing quietly working.
+TRACE_WRITE_FAILED = "duocthu_trace_write_failed_total"
diff --git a/apps/ai-service/rag/policy.py b/apps/ai-service/rag/policy.py
new file mode 100644
index 0000000..769be5a
--- /dev/null
+++ b/apps/ai-service/rag/policy.py
@@ -0,0 +1,71 @@
+"""Server-derived subject scope.
+
+`routing.py`'s `_scope_gate` abstains outright on `NON_HUMAN` scope. Before
+this module, that value came straight from the request body — a caller could
+send `{"subject_scope":"human",...}` regardless of the query text, and the
+shipped web BFF did exactly that, hard-coded on every request without reading
+the message at all (Codex's 2026-08-06 review, F-02). A scope decision must
+not be something the caller gets to assert.
+
+`resolve_subject_scope` derives it from the query text and combines that with
+whatever the caller claimed by taking the more conservative of the two: a
+caller can *narrow* scope (claim `non_human` and have it stick) but can never
+*widen* it — a claim of `human` cannot override a server-detected veterinary
+turn. This is a corpus-coverage check, not a restriction on what a doctor or
+pharmacist is allowed to ask: the book behind this product covers human
+drug monographs only, so a query about dosing a dog is out of scope
+regardless of who is asking. It has nothing to do with, and must never be
+extended into, gatekeeping what kind of *clinical* question a professional
+user is allowed to ask (see `[[feedback_no_recommendation_gate]]`/progress
+log 2026-08-06 for the removed `QueryIntent.RECOMMENDATION` keyword
+detector — this product is for doctors and pharmacists, not lay users, and a
+"nên dùng thuốc gì" question from a clinician is exactly what a formulary
+reference is for, not something to abstain on).
+
+Deliberately not an LLM call: this is the gate every request passes through,
+so it must be cheap, available during a provider outage, and auditable as a
+fixed rule instead of a model judgment. It is a keyword heuristic, which
+means it has real blind spots (an unusual phrasing can still slip past) — the
+same trade-off `rag/agent.py`'s `_looks_non_human` backstop already accepts.
+`rag/agent.py` should consolidate onto this module once the new orchestrator
+is wired (F-03), instead of keeping a second, narrower keyword list.
+"""
+from __future__ import annotations
+
+from .models import SubjectScope
+from .text import normalize_name
+
+# Phrases that name a non-human recipient. Matched on normalized text (casefold,
+# diacritics stripped) so "chó", "CHÓ", "cho chó" all match one entry.
+_NON_HUMAN_PHRASES = (
+ "cho cho", # "cho chó" — normalize_name collapses "chó" -> "cho"
+ "cho meo",
+ "cho ga",
+ "cho vit",
+ "cho lon",
+ "cho heo",
+ "cho bo",
+ "cho ngua",
+ "cho de",
+ "cho cuu",
+ "thu y",
+ "vat nuoi",
+ "gia suc",
+ "gia cam",
+ "dong vat",
+)
+
+
+def looks_non_human(query: str) -> bool:
+ normalized = normalize_name(query)
+ return any(phrase in normalized for phrase in _NON_HUMAN_PHRASES)
+
+
+def resolve_subject_scope(query: str, claimed: SubjectScope) -> SubjectScope:
+ """The scope that actually gates retrieval: the more conservative of what
+ the caller claimed and what the query text itself indicates."""
+ if claimed == SubjectScope.NON_HUMAN or looks_non_human(query):
+ return SubjectScope.NON_HUMAN
+ if claimed == SubjectScope.UNKNOWN:
+ return SubjectScope.UNKNOWN
+ return SubjectScope.HUMAN
diff --git a/apps/ai-service/rag/prompt.py b/apps/ai-service/rag/prompt.py
index 82febdb..30514a0 100644
--- a/apps/ai-service/rag/prompt.py
+++ b/apps/ai-service/rag/prompt.py
@@ -36,6 +36,19 @@ Quy tắc bắt buộc:
trả lời hợp lệ. Không suy diễn để lấp chỗ trống.
6. Giữ nguyên thuật ngữ chuyên môn của sách. Không diễn giải cho người
không chuyên.
+7. HỎI LẠI khi thiếu dữ kiện — ĐÂY LÀ QUY TẮC QUAN TRỌNG NHẤT, ưu tiên hơn việc
+ trả lời. TUYỆT ĐỐI KHÔNG liệt kê nhiều mức liều rồi để người đọc tự chọn.
+ Nếu là câu hỏi LIỀU/CÁCH DÙNG và bằng chứng phân mức theo điều kiện (tuổi,
+ cân nặng, chức năng thận/gan, chỉ định/bệnh, đường dùng, mức độ nặng…) mà
+ người dùng CHƯA nêu đủ (các) điều kiện để chọn ĐÚNG MỘT mức, thì BẮT BUỘC:
+ để `answer`="", `evidence_sufficient`=false, và đặt `clarifying_question`
+ hỏi NGẮN GỌN tất cả dữ kiện còn thiếu.
+ - "trẻ em" hay "cho trẻ" nói chung là CHƯA đủ (liều trẻ em thay đổi theo
+ tuổi/cân nặng) → phải hỏi lại, KHÔNG được liệt kê các nhóm tuổi.
+ - "người lớn" thường là đủ cho liều người lớn tiêu chuẩn → trả lời được.
+ Ví dụ clarifying_question: "Bé mấy tuổi, cân nặng bao nhiêu kg, dùng đường
+ nào (uống/đặt hậu môn/tiêm) và để hạ sốt hay giảm đau?".
+ Nếu đã đủ dữ kiện thì trả lời bình thường, `clarifying_question`=null.
Viết gọn. Trả lời đúng điều được hỏi, không mở rộng phạm vi."""
@@ -54,12 +67,59 @@ ANSWER_SCHEMA = {
"type": "boolean",
"description": "false nếu bằng chứng không đủ để trả lời câu hỏi.",
},
+ "clarifying_question": {
+ "type": ["string", "null"],
+ "description": (
+ "Câu hỏi lại khi người dùng chưa nêu đủ dữ kiện (vd tuổi/cân nặng "
+ "cho câu hỏi liều có nhiều mức). null nếu đã đủ dữ kiện để trả lời."
+ ),
+ },
},
"required": ["answer", "evidence_sufficient"],
"additionalProperties": False,
}
+SUFFICIENCY_SYSTEM = """\
+Bạn là bộ KIỂM TRA ĐỦ DỮ KIỆN cho tra cứu Dược thư, chạy TRƯỚC khi trả lời.
+Cho CÂU HỎI của người dùng và BẰNG CHỨNG, xác định câu hỏi đã đủ dữ kiện để đưa
+ra ĐÚNG MỘT câu trả lời cụ thể hay chưa.
+
+Quy tắc:
+- Nếu là câu hỏi về LIỀU/CÁCH DÙNG và BẰNG CHỨNG có NHIỀU mức khác nhau theo điều
+ kiện (tuổi, cân nặng, chức năng thận/gan, chỉ định/bệnh, đường dùng, mức độ
+ nặng) mà CÂU HỎI chưa nêu đủ (các) điều kiện để chọn đúng MỘT mức → CHƯA đủ.
+- "trẻ em" / "cho trẻ" / "cho bé" nói chung là CHƯA đủ (liều trẻ thay đổi theo
+ tuổi và cân nặng). "người lớn" thường ĐỦ cho liều người lớn tiêu chuẩn.
+- Câu hỏi KHÔNG về liều (chống chỉ định, tương tác, tác dụng phụ, giới thiệu
+ thuốc…) thì thường ĐỦ.
+
+Trả về DUY NHẤT JSON: {"sufficient": bool, "clarifying_question": string|null}.
+Nếu CHƯA đủ: sufficient=false và clarifying_question hỏi NGẮN GỌN tất cả dữ kiện
+còn thiếu (vd: "Bé mấy tuổi, cân nặng bao nhiêu kg, dùng đường nào và để hạ sốt
+hay giảm đau?"). Nếu đủ: sufficient=true, clarifying_question=null."""
+
+SUFFICIENCY_SCHEMA = {
+ "type": "object",
+ "properties": {
+ "sufficient": {"type": "boolean"},
+ "clarifying_question": {"type": ["string", "null"]},
+ },
+ "required": ["sufficient", "clarifying_question"],
+ "additionalProperties": False,
+}
+
+
+def build_sufficiency_request(
+ question: str, evidence_texts: tuple[str, ...]
+) -> "GenerationRequest":
+ blocks = "\n\n".join(
+ f"[{index}] {text}" for index, text in enumerate(evidence_texts, start=1)
+ )
+ user = f"BẰNG CHỨNG:\n\n{blocks}\n\nCÂU HỎI: {question}"
+ return GenerationRequest(system=SUFFICIENCY_SYSTEM, user=user, schema=SUFFICIENCY_SCHEMA)
+
+
@dataclass(frozen=True)
class GenerationRequest:
system: str
@@ -67,6 +127,58 @@ class GenerationRequest:
schema: dict
+ENTAILMENT_SYSTEM = """\
+Bạn là bộ KIỂM TRA ĐỘ CHÍNH XÁC, chạy SAU khi một câu trả lời đã được sinh ra.
+
+Với mỗi CÂU dưới đây, so sánh nó với ĐÚNG đoạn BẰNG CHỨNG ĐƯỢC TRÍCH đi kèm câu
+đó (câu đã được gắn số nguồn [n] trỏ tới đúng đoạn này). Việc DUY NHẤT cần làm:
+nội dung của CÂU có được đoạn BẰNG CHỨNG ĐƯỢC TRÍCH đó — và CHỈ đoạn đó — nói
+tới hay không. KHÔNG dùng kiến thức y khoa của bạn, KHÔNG suy luận thêm, KHÔNG
+tự hỏi liệu câu đó có hợp lý về mặt y khoa hay không.
+
+Một CÂU là KHÔNG được chứng thực nếu nó nêu chỉ định, chống chỉ định, cơ chế,
+tương tác, đối tượng áp dụng, hoặc bất kỳ quan hệ nào mà đoạn BẰNG CHỨNG ĐƯỢC
+TRÍCH của nó KHÔNG nói tới — kể cả khi câu đó đúng về mặt y khoa, và kể cả khi
+đúng tên thuốc nhưng sai ý (ví dụ bằng chứng nói "điều trị đái tháo đường" mà
+câu nói "chữa ung thư").
+
+LƯU Ý QUAN TRỌNG: bằng chứng trong lĩnh vực dược thường liệt kê nhiều tên
+thuốc trong một câu/danh sách dài, phân cách bởi dấu phẩy (ví dụ: "Tác dụng
+của warfarin có thể tăng lên khi dùng với: acetaminophen, allopurinol, ...,
+aspirin, kháng sinh, ..."). Hãy ĐỌC KỸ TOÀN BỘ danh sách trước khi kết luận —
+nếu tên thuốc trong CÂU xuất hiện ở bất kỳ đâu trong danh sách đó với đúng
+quan hệ đang nói (vd "làm tăng tác dụng của X"), đó LÀ được chứng thực, dù
+tên thuốc chỉ là một mục nhỏ giữa danh sách dài.
+
+Trả về DUY NHẤT JSON: {"entailed": bool, "unsupported": [danh sách số thứ tự
+1-based của các CÂU KHÔNG được chứng thực; rỗng nếu tất cả đều được chứng
+thực]}. entailed=true chỉ khi unsupported rỗng."""
+
+ENTAILMENT_SCHEMA = {
+ "type": "object",
+ "properties": {
+ "entailed": {"type": "boolean"},
+ "unsupported": {"type": "array", "items": {"type": "integer"}},
+ },
+ "required": ["entailed", "unsupported"],
+ "additionalProperties": False,
+}
+
+
+def build_entailment_request(claims: list[tuple[str, str]]) -> "GenerationRequest":
+ """`claims` is a list of (claim_text, cited_evidence_text) pairs — already
+ filtered by the caller to the claims worth checking (substantive content,
+ a validly-cited evidence block to check it against)."""
+ if not claims:
+ raise ValueError("cannot build an entailment request with no claims")
+ blocks = "\n\n".join(
+ f"CÂU {index}: {claim}\nBẰNG CHỨNG ĐƯỢC TRÍCH: {evidence}"
+ for index, (claim, evidence) in enumerate(claims, start=1)
+ )
+ user = f"{blocks}\n\nKiểm tra từng CÂU theo đúng BẰNG CHỨNG ĐƯỢC TRÍCH của nó."
+ return GenerationRequest(system=ENTAILMENT_SYSTEM, user=user, schema=ENTAILMENT_SCHEMA)
+
+
def build_request(
question: str, evidence_texts: tuple[str, ...], intro: bool = False
) -> GenerationRequest:
diff --git a/apps/ai-service/rag/service.py b/apps/ai-service/rag/service.py
index 26c718d..22a5ff6 100644
--- a/apps/ai-service/rag/service.py
+++ b/apps/ai-service/rag/service.py
@@ -58,6 +58,59 @@ class RetrievalService:
self._section_resolver = section_resolver
self._reranker = reranker
+ def retrieve_framed(
+ self,
+ drug_id: str,
+ section_key: str | None,
+ query: str,
+ is_overview: bool = False,
+ ) -> RetrievalResult:
+ """Retrieve driven by an already-understood frame, not by parsing text.
+
+ The LLM understanding layer has resolved the drug (against the real
+ catalog) and, when the turn named one, the section. So this skips the
+ fuzzy `CatalogDrugResolver` and the keyword `SectionResolver` entirely:
+ a named `section_key` filters that section whole; without one, this
+ mirrors `retrieve()`'s two remaining cases — `is_overview` (the frame's
+ `turn_type == "drug_overview"`, a bare name) answers from the identity
+ sections only, and a free-form question reranks the full monograph
+ down to `rerank_top_k`.
+
+ Found live 2026-08-06: without the `is_overview` split, a bare drug
+ name sent the ENTIRE ~29-section monograph as evidence for every
+ generation call (retrieval had no notion of "just the intro"), which
+ is both wrong retrieval and, downstream, an answer so long it
+ intermittently failed generation/entailment outright. `query` is only
+ the rerank signal here, never a resolution input.
+ """
+ if not drug_id.strip():
+ return RetrievalResult(EvidenceDecision.ABSTAIN, "missing_query_or_drug")
+
+ if section_key:
+ find_by_section = getattr(self._retriever, "find_by_section", None)
+ if find_by_section is not None:
+ hits = find_by_section(drug_id, section_key)
+ if hits:
+ return self._decide(self._hydrate(hits, limit=None))
+
+ overview_hits = self._drug_overview(drug_id)
+ if overview_hits is None:
+ return RetrievalResult(EvidenceDecision.ABSTAIN, "insufficient_retrieval_score")
+
+ if is_overview:
+ intro = [h for h in overview_hits if h.document.section_key in INTRO_SECTIONS]
+ return self._decide(
+ self._hydrate(intro or overview_hits, limit=None), is_drug_overview=True
+ )
+
+ overview_hits = self._rerank(query, overview_hits)
+ # Capped even when rerank is disabled/unavailable and fails open to
+ # the unfiltered list — an ordering aid must never remove the size
+ # bound too, or the same 29-section explosion returns through here.
+ return self._decide(
+ self._hydrate(overview_hits, limit=self._policy.evidence_limit)
+ )
+
def retrieve(self, query: str, drug_id: str) -> RetrievalResult:
if not query.strip() or not drug_id.strip():
return RetrievalResult(EvidenceDecision.ABSTAIN, "missing_query_or_drug")
diff --git a/apps/ai-service/rag/understanding.py b/apps/ai-service/rag/understanding.py
new file mode 100644
index 0000000..ed34a6a
--- /dev/null
+++ b/apps/ai-service/rag/understanding.py
@@ -0,0 +1,353 @@
+"""LLM-driven query understanding — the front-end of the RAG chatbot.
+
+The old drug-first path resolved a drug with a fuzzy `SequenceMatcher` and routed
+sections with a Vietnamese phrase table. Both are brittle string heuristics: they
+false-matched a made-up name (``aspirinol`` -> aspirin), tied on a correctly
+spelled English INN (``amoxicillin``), and mistook a common word (``uống``) for a
+drug. This module replaces that with the model the system already has.
+
+Division of labour, and why it is safe for a formulary:
+
+- The **LLM** does the messy human-language part — which token is a drug, is this
+ an interaction / a symptom lookup / a weight-based dose, what section is asked,
+ what population/weight. It is good at exactly the fuzziness the heuristics were
+ bad at.
+- The **catalog** stays the authority on drug *identity*. The model may only pick
+ ``drug_id`` values from a list *bounded before the model ever runs* — a
+ deterministic alias/fuzzy pass over the turn and history (`CandidateSource`)
+ decides which real drugs are even plausible candidates, and only those are
+ shown. This closes a gap the catalog-whitelist alone did not (F-04, Codex
+ 2026-08-06 review): validating that an output id is *some* real drug_id does
+ not prove it is the *one the user's text actually named* — an LLM could
+ satisfy that whitelist while mapping an unrelated or invented name to any of
+ the other 683 real drugs. Bounding candidates first removes that degree of
+ freedom: the model can still read ``amoxicillin`` as ``amoxicilin`` (a fuzzy
+ match puts it in the candidate set) but cannot map ``aspirinol`` to aspirin,
+ because nothing about ``aspirinol`` fuzzy-matches anything and the candidate
+ set the model is shown is empty or contains unrelated drugs, not aspirin.
+- Nothing here answers the medical question. It only produces a structured frame;
+ retrieval + ``grounding.verify`` remain the load-bearing safety layer downstream.
+
+`rag/` imports no SDK: the LLM is injected as a ``JsonLlm`` protocol (satisfied by
+`adapters.bedrock_converse.BedrockConverseAnswerGenerator`), and a deterministic
+stub runs the whole path offline in tests.
+"""
+from __future__ import annotations
+
+import json
+from dataclasses import dataclass, field
+from typing import Protocol, Sequence
+
+# The 19 monograph section keys, kept here as the closed vocabulary the model may
+# use for `attribute`. Adding a new section is one entry, not a code change.
+SECTION_KEYS = (
+ "ten_chung_quoc_te",
+ "ten_thuong_mai",
+ "ma_atc",
+ "loai_thuoc",
+ "dang_thuoc_va_ham_luong",
+ "duoc_ly_va_co_che_tac_dung",
+ "chi_dinh",
+ "chong_chi_dinh",
+ "than_trong",
+ "thoi_ky_mang_thai",
+ "thoi_ky_cho_con_bu",
+ "tac_dung_khong_mong_muon",
+ "huong_dan_xu_tri_adr",
+ "lieu_luong_va_cach_dung",
+ "tuong_tac_thuoc",
+ "qua_lieu_va_xu_tri",
+ "do_on_dinh_va_bao_quan",
+ "tuong_ky",
+ "thong_tin_quy_che",
+)
+
+# Short glosses shown to the model alongside each key. Found live 2026-08-06
+# (golden e2e set): a bare key list gives the model nothing to disambiguate
+# "thận trọng" from "chống chỉ định" — 9/9 live calls for "X cần thận trọng
+# gì?" picked chong_chi_dinh, silently answering from the wrong section
+# (and, downstream, dropping the specific safety content the precautions
+# section actually has, e.g. metformin's lactic acidosis warning). The two
+# are genuinely adjacent concepts in Vietnamese medical text; a bare slug
+# name is not enough to tell a model which one a question means.
+SECTION_KEY_HINTS: dict[str, str] = {
+ "ten_chung_quoc_te": "tên chung quốc tế/INN",
+ "ten_thuong_mai": "tên thương mại/biệt dược",
+ "ma_atc": "mã ATC",
+ "loai_thuoc": "phân loại thuốc",
+ "dang_thuoc_va_ham_luong": "dạng bào chế và hàm lượng",
+ "duoc_ly_va_co_che_tac_dung": "dược lý, cơ chế tác dụng",
+ "chi_dinh": "chỉ định — bệnh/triệu chứng thuốc dùng để điều trị",
+ "chong_chi_dinh": (
+ "CHỐNG CHỈ ĐỊNH — trường hợp TUYỆT ĐỐI KHÔNG được dùng thuốc này"
+ ),
+ "than_trong": (
+ "THẬN TRỌNG — KHÁC chống chỉ định: vẫn dùng được nhưng cần cảnh "
+ "giác/theo dõi/chỉnh liều (ví dụ nguy cơ nhiễm toan lactic của "
+ "metformin, độc tính thận/tai của gentamicin). Câu hỏi có chữ "
+ "\"thận trọng\", \"cẩn thận\", \"lưu ý gì\", \"cần chú ý\" → key này, "
+ "KHÔNG PHẢI chong_chi_dinh."
+ ),
+ "thoi_ky_mang_thai": "dùng khi mang thai",
+ "thoi_ky_cho_con_bu": "dùng khi cho con bú",
+ "tac_dung_khong_mong_muon": "tác dụng phụ/ADR",
+ "huong_dan_xu_tri_adr": "cách xử trí khi gặp ADR",
+ "lieu_luong_va_cach_dung": "liều dùng và cách dùng",
+ "tuong_tac_thuoc": "tương tác với thuốc khác",
+ "qua_lieu_va_xu_tri": "quá liều và cách xử trí",
+ "do_on_dinh_va_bao_quan": "độ ổn định, bảo quản",
+ "tuong_ky": "tương kỵ (không pha/trộn được với gì)",
+ "thong_tin_quy_che": "thông tin quy chế/pháp lý",
+}
+
+# What kind of turn this is — the router branches on it. Deliberately explicit so a
+# symptom lookup is never silently treated as a failed drug lookup, and a two-drug
+# interaction never collapses to an "ambiguous drug" abstain.
+TURN_TYPES = (
+ "drug_attribute", # one drug, one/more sections ("liều paracetamol")
+ "drug_overview", # a bare drug name, wants the monograph
+ "interaction", # 2+ drugs, asks about combining them
+ "symptom_to_drug", # a symptom/indication, wants candidate drugs
+ "dosing_calc", # a dose that needs weight/age arithmetic
+ "smalltalk", # greeting / meta, not a medical query
+ "out_of_scope", # not answerable from the Part-2 monographs
+)
+
+
+@dataclass(frozen=True)
+class QueryFrame:
+ """The structured reading of one user turn. No medical content, only intent."""
+
+ turn_type: str
+ drugs: tuple[str, ...] = () # canonical drug_ids, from the catalog only
+ unknown_drugs: tuple[str, ...] = () # mentioned, not in the catalog
+ attribute: str | None = None # a SECTION_KEYS value, or None
+ population: str | None = None # e.g. "tre_em", "nguoi_lon", "suy_than"
+ weight_kg: float | None = None
+ age_text: str | None = None
+ indication: str | None = None # symptom/disease, for symptom_to_drug
+ needs_clarify: bool = False
+ clarify_reason: str | None = None
+ raw: dict = field(default_factory=dict, compare=False)
+
+
+# The JSON contract the model must fill. Stated in the prompt (Converse has no
+# server-side schema) and validated on the way back.
+FRAME_SCHEMA = {
+ "turn_type": "one of: " + " | ".join(TURN_TYPES),
+ "drugs": ["drug_id exactly as it appears in the provided catalog list"],
+ "unknown_drugs": ["a drug name the user mentioned that is NOT in the catalog"],
+ "attribute": "one of the section keys provided, or null",
+ "population": "tre_em | tre_so_sinh | nguoi_lon | nguoi_cao_tuoi | phu_nu_co_thai | phu_nu_cho_con_bu | suy_than | suy_gan | null",
+ "weight_kg": (
+ "number if a body weight is given, else null. Vietnamese casual speech "
+ "states weight as a bare number of 'cân' or 'ký' with no unit word "
+ "('bé 30 cân', 'nặng 30 ký') — both mean kilograms; read the number as "
+ "weight_kg the same as if 'kg' had been written."
+ ),
+ "age_text": "the age exactly as stated (e.g. '3 tuổi', '5 tháng'), else null",
+ "indication": "the symptom or disease if turn_type is symptom_to_drug, else null",
+ "needs_clarify": "true only if the turn cannot be acted on without more info",
+ "clarify_reason": "short Vietnamese question to ask, or null",
+}
+
+_SYSTEM = """Bạn là bộ HIỂU CÂU HỎI cho một chatbot tra cứu Dược thư Quốc gia Việt Nam.
+Nhiệm vụ: đọc câu hỏi (tiếng Việt, có thể sai chính tả, viết tắt, nhiều lượt) và
+XUẤT RA một JSON mô tả ý định. TUYỆT ĐỐI KHÔNG trả lời câu hỏi y khoa, không nêu liều.
+
+Quy tắc bắt buộc:
+- Trường "drugs" CHỈ được chứa các drug_id có trong DANH SÁCH THUỐC được cung cấp.
+ Nếu người dùng nhắc một thuốc KHÔNG có trong danh sách (kể cả tên bịa như
+ "aspirinol"), đưa tên đó vào "unknown_drugs", KHÔNG được gán sang thuốc gần giống.
+- Sai chính tả một thuốc CÓ trong danh sách thì sửa về đúng drug_id của nó
+ (ví dụ "amoxicillin" -> "amoxicilin", "metfomin" -> "metformin").
+- Nếu câu nhắc 2 thuốc trở lên và hỏi về dùng chung/tương tác -> turn_type="interaction".
+- Nếu là triệu chứng/bệnh cần gợi ý thuốc (không nêu tên thuốc) -> "symptom_to_drug",
+ điền "indication".
+- Nếu hỏi liều cần cân nặng/tuổi -> "dosing_calc", điền weight_kg/age_text nếu có.
+ Nói cân nặng kiểu thường ngày ("bé 30 cân", "nặng 30 ký", chỉ 1 số + "cân"/"ký"
+ không kèm đơn vị khác) NGHĨA LÀ 30 kg -> điền weight_kg=30, không bỏ trống.
+- Lượt nối tiếp ("còn liều thì sao", "nó dùng cho trẻ em?") -> dùng LỊCH SỬ để biết
+ thuốc đang nói tới và điền vào "drugs".
+- Chào hỏi/vu vơ -> "smalltalk". Ngoài phạm vi chuyên luận thuốc -> "out_of_scope"."""
+
+
+class JsonLlm(Protocol):
+ """A model that returns a single JSON object as text. Satisfied by the
+ existing Bedrock Converse generator, so this adds no SDK to `rag/`."""
+
+ def generate(self, system: str, user: str, schema: dict) -> str: ...
+
+
+class QueryUnderstander(Protocol):
+ def understand(
+ self, turn: str, history: Sequence[str] = ()
+ ) -> QueryFrame: ...
+
+
+class CandidateSource(Protocol):
+ """Deterministic, no-LLM drug-name matching — what bounds the model's
+ choice before it ever runs (F-04). Satisfied by `routing.CatalogDrugResolver`;
+ kept as a protocol (not an import of it) so this module stays decoupled
+ from the fuzzy-matching implementation, only its shape.
+ """
+
+ def resolve(self, query: str): ...
+ def suggest(
+ self, query: str, k: int = 3, min_score: float = 0.5
+ ) -> list[tuple[str, float]]: ...
+
+
+class LlmQueryUnderstander:
+ """Turns a raw user turn into a `QueryFrame` with one LLM call.
+
+ `catalog` maps drug_id -> a human name (used only to label whichever
+ candidates get shown). `resolver` is what actually decides which real
+ drugs are plausible for this turn, deterministically, before the model
+ runs at all: every drug_id an exact-alias or fuzzy match finds anywhere
+ in the turn or the raw history text. The model then picks only among
+ those — never the full ~684-drug catalog — so it structurally cannot
+ map an invented or unrelated name to some other real drug_id it merely
+ happens to also list correctly (F-04). This also directly answers a
+ separate 2026-08-06 review finding: sending the full catalog on every
+ turn is unbounded token cost; a per-turn candidate shortlist is both
+ safer and cheaper.
+ """
+
+ def __init__(self, llm: JsonLlm, catalog: dict[str, str], resolver: CandidateSource) -> None:
+ self._llm = llm
+ self._catalog = catalog
+ self._resolver = resolver
+
+ def _candidate_ids(self, turn: str, history: Sequence[str]) -> set[str]:
+ """Every drug_id a deterministic pass finds plausible in the turn or
+ the raw history text. Deliberately generous — an exact alias match
+ plus a fuzzy `suggest` well below the resolver's own auto-answer
+ threshold — because the job here is only to rule out drugs nothing
+ in the conversation plausibly refers to, not to pick the right one;
+ that disambiguation is still the model's job, within this bound.
+ """
+ ids: set[str] = set()
+ for line in (turn, *history):
+ if not line.strip():
+ continue
+ resolution = self._resolver.resolve(line)
+ if resolution.status == "resolved" and resolution.drug_id:
+ ids.add(resolution.drug_id)
+ elif resolution.status == "ambiguous":
+ ids.update(resolution.candidate_drug_ids)
+ for drug_id, _score in self._resolver.suggest(line, k=5, min_score=0.55):
+ ids.add(drug_id)
+ return ids
+
+ def understand(self, turn: str, history: Sequence[str] = ()) -> QueryFrame:
+ shown = {
+ drug_id: self._catalog[drug_id]
+ for drug_id in self._candidate_ids(turn, history)
+ if drug_id in self._catalog
+ }
+ catalog_block = (
+ "\n".join(f"{drug_id}\t{name}" for drug_id, name in sorted(shown.items()))
+ if shown
+ else "(không có thuốc nào trong Dược thư khớp với lượt này hoặc lịch sử gần đây)"
+ )
+ history_block = (
+ "LỊCH SỬ HỘI THOẠI (cũ -> mới):\n" + "\n".join(history)
+ if history else "LỊCH SỬ HỘI THOẠI: (chưa có)"
+ )
+ user = (
+ f"DANH SÁCH THUỐC ỨNG VIÊN cho lượt này (drug_id\\ttên) — CHỈ được chọn "
+ f"drug_id từ đây, đây KHÔNG phải toàn bộ Dược thư, chỉ là các thuốc khớp "
+ f"với chữ trong lượt/lịch sử:\n"
+ f"{catalog_block}\n\n"
+ "CÁC SECTION KEY hợp lệ cho 'attribute' (key: ý nghĩa):\n"
+ + "\n".join(f"{key}: {SECTION_KEY_HINTS[key]}" for key in SECTION_KEYS)
+ + "\n\n"
+ f"{history_block}\n\n"
+ f"CÂU HỎI HIỆN TẠI: {turn}"
+ )
+ raw_text = self._llm.generate(_SYSTEM, user, FRAME_SCHEMA)
+ return self._parse(raw_text, shown)
+
+ @staticmethod
+ def _resolve_id(value: str, shown: dict[str, str]) -> str | None:
+ if value in shown:
+ return value
+ # `drug_id` is shown with underscores ("paracetamol_acetaminophen")
+ # but its own canonical display name (bootstrap's `_catalog_names`)
+ # is the same string with spaces — found live 2026-08-06: the two
+ # look near-identical in the "drug_id\tname" table, and the model
+ # sometimes echoes the spaced display form instead of the id. This
+ # is a deterministic, lossless formatting difference (not a fuzzy
+ # match — one specific known substitution), so it's tolerated here
+ # rather than dropping a correctly-identified drug to unknown.
+ spaced = value.strip().casefold()
+ for drug_id in shown:
+ if drug_id.replace("_", " ").casefold() == spaced:
+ return drug_id
+ return None
+
+ def _parse(self, raw_text: str, shown: dict[str, str]) -> QueryFrame:
+ try:
+ data = json.loads(raw_text)
+ except (json.JSONDecodeError, TypeError):
+ # Fail closed to a clarify rather than to a wrong reading.
+ return QueryFrame(
+ turn_type="out_of_scope",
+ needs_clarify=True,
+ clarify_reason="Xin lỗi, tôi chưa hiểu rõ câu hỏi. Anh/chị hỏi lại giúp nhé?",
+ )
+ resolved = [
+ (d, self._resolve_id(d, shown)) for d in _as_list(data.get("drugs"))
+ ]
+ drugs = tuple(dict.fromkeys(rid for _, rid in resolved if rid is not None))
+ # A drug the model named but that resolves to no id among the shown
+ # candidates (exact or underscore/space form) is unknown, not a
+ # silent drop and not a fuzzy substitution to an unrelated drug.
+ unknown = tuple(
+ d for d, rid in resolved if rid is None
+ ) + tuple(_as_list(data.get("unknown_drugs")))
+ attribute = data.get("attribute")
+ if attribute not in SECTION_KEYS:
+ attribute = None
+ turn_type = data.get("turn_type")
+ if turn_type not in TURN_TYPES:
+ turn_type = "drug_attribute" if drugs else "out_of_scope"
+ return QueryFrame(
+ turn_type=turn_type,
+ drugs=drugs,
+ unknown_drugs=tuple(dict.fromkeys(unknown)),
+ attribute=attribute,
+ population=_clean_str(data.get("population")),
+ weight_kg=_clean_float(data.get("weight_kg")),
+ age_text=_clean_str(data.get("age_text")),
+ indication=_clean_str(data.get("indication")),
+ needs_clarify=bool(data.get("needs_clarify")),
+ clarify_reason=_clean_str(data.get("clarify_reason")),
+ raw=data if isinstance(data, dict) else {},
+ )
+
+
+def _as_list(value) -> list[str]:
+ if isinstance(value, str):
+ return [value] if value.strip() else []
+ if isinstance(value, list):
+ return [str(v).strip() for v in value if str(v).strip()]
+ return []
+
+
+def _clean_str(value) -> str | None:
+ if isinstance(value, str) and value.strip() and value.strip().lower() != "null":
+ return value.strip()
+ return None
+
+
+def _clean_float(value) -> float | None:
+ if isinstance(value, (int, float)):
+ return float(value)
+ if isinstance(value, str):
+ try:
+ return float(value.replace(",", ".").split()[0])
+ except (ValueError, IndexError):
+ return None
+ return None
diff --git a/apps/ai-service/routers/rag.py b/apps/ai-service/routers/rag.py
index 1aeac86..ad8f07d 100644
--- a/apps/ai-service/routers/rag.py
+++ b/apps/ai-service/routers/rag.py
@@ -1,12 +1,15 @@
from __future__ import annotations
+import uuid
from typing import Annotated, Any, Protocol
from fastapi import APIRouter, Depends, HTTPException, Request
from pydantic import BaseModel, Field
from rag.answer import GroundedAnswerService
+from rag.metrics import TRACE_WRITE_FAILED, Metrics, NullMetrics
from rag.models import QueryIntent, SubjectScope
+from rag.policy import resolve_subject_scope
class TraceWriter(Protocol):
@@ -57,6 +60,10 @@ def _trace_writer(request: Request) -> TraceWriter:
return writer
+def _metrics(request: Request) -> Metrics:
+ return getattr(request.app.state, "metrics", None) or NullMetrics()
+
+
router = APIRouter(prefix="/v1/rag", tags=["rag"])
@@ -67,10 +74,10 @@ class SuggestResponse(BaseModel):
@router.get("/suggest", response_model=SuggestResponse)
def suggest_drugs(q: str, request: Request) -> SuggestResponse:
"""As-you-type drug-name autocomplete, so a name is picked, not mistyped."""
- conversational = getattr(request.app.state, "conversational", None)
- if conversational is None or not q.strip():
+ agent = getattr(request.app.state, "conversational", None)
+ if agent is None or not q.strip():
return SuggestResponse(suggestions=[])
- return SuggestResponse(suggestions=conversational.complete(q.strip()))
+ return SuggestResponse(suggestions=agent.complete(q.strip()))
def _map_citations(items) -> list[CitationResponse]:
@@ -95,47 +102,72 @@ def query_rag(
request: Request,
answers: Annotated[GroundedAnswerService, Depends(_answer_service)],
traces: Annotated[TraceWriter, Depends(_trace_writer)],
+ metrics: Annotated[Metrics, Depends(_metrics)],
) -> RagQueryResponse:
- conversational = getattr(request.app.state, "conversational", None)
+ agent = getattr(request.app.state, "conversational", None)
- # Single-turn path (no conversation id, or conversational layer disabled):
- # unchanged behaviour so existing callers keep working.
- if payload.conversation_id is None or conversational is None:
- grounded = answers.answer(payload.query, payload.subject_scope, payload.intent)
- decision = grounded.result.decision.value
- reason = grounded.result.reason
- answer = grounded.answer
- resolved_drug_id = grounded.result.resolved_drug_id
- citations = _map_citations(grounded.citations)
+ # `payload.subject_scope`/`payload.intent` are what the CALLER claims —
+ # logged below for audit, but the RagAgent path does not take them as an
+ # input at all. It derives scope from the query text itself (the same
+ # `rag.policy` check used here for the no-agent fallback) as part of its
+ # own LLM understanding call, and does not gate on intent at all (this
+ # product is for doctors/pharmacists; a client label must not be, and
+ # here structurally cannot be, the safety decision — F-02).
+ subject_scope = resolve_subject_scope(payload.query, payload.subject_scope)
+ intent = payload.intent
+
+ if agent is not None:
+ # The live path (F-03): one LLM call understands the turn (drug
+ # identity against the real catalog, turn type, population/weight),
+ # then routes to the safety-verified retrieval + grounded-answer
+ # engine. Replaces the old fuzzy resolver + keyword section router +
+ # manual follow-up inheritance for both single- and multi-turn.
+ reply = agent.handle(payload.query, payload.conversation_id)
+ decision = reply.decision
+ reason = reply.reason
+ answer = reply.clarification if reply.clarification is not None else reply.answer
+ resolved_drug_id = ", ".join(reply.drugs) if reply.drugs else None
+ citations = _map_citations(reply.citations)
else:
- turn = conversational.answer(
- payload.conversation_id,
- payload.query,
- payload.subject_scope,
- payload.intent,
- )
- if turn.clarification is not None:
- decision, reason = "clarify", turn.clarification.reason
- answer, resolved_drug_id, citations = turn.clarification.question, None, []
- elif turn.grounded is not None:
- decision = turn.grounded.result.decision.value
- reason = turn.grounded.result.reason
- answer = turn.answer
- resolved_drug_id = turn.grounded.result.resolved_drug_id
- citations = _map_citations(turn.grounded.citations)
- else: # smalltalk
- decision, reason = "answerable", turn.reason
- answer, resolved_drug_id, citations = turn.answer, None, []
+ # No generator configured (ANSWER_PROVIDER=disabled): there is no LLM
+ # to understand a turn with, so this is retrieval-only, single-turn,
+ # unchanged from before F-03.
+ grounded = answers.answer(payload.query, subject_scope, intent)
+ if grounded.clarification is not None:
+ decision, reason = "clarify", "needs_more_info"
+ answer = grounded.clarification
+ resolved_drug_id = grounded.result.resolved_drug_id
+ citations = []
+ else:
+ decision = grounded.result.decision.value
+ reason = grounded.result.reason
+ answer = grounded.answer
+ resolved_drug_id = grounded.result.resolved_drug_id
+ citations = _map_citations(grounded.citations)
- trace_id = traces.save(
- query=payload.query,
- subject_scope=payload.subject_scope.value,
- intent=payload.intent.value,
- decision=decision,
- reason=reason,
- resolved_drug_id=resolved_drug_id,
- citations=tuple(item.model_dump() for item in citations),
- )
+ # Trace persistence is fail-open (F-09): an already-computed, safe answer
+ # must reach the caller even if Postgres is unreachable. `save()` opens a
+ # fresh connection per call with no retry, so a DB outage previously
+ # turned a good answer into a 500 for a reason that has nothing to do
+ # with whether the answer was safe. `trace_id` degrades to a local,
+ # unpersisted uuid — still a valid response field, just not one `GET
+ # /v1/rag/trace/{id}` (if it existed) could later look up.
+ try:
+ trace_id = traces.save(
+ query=payload.query,
+ # The resolved (server-derived) values, not the caller's claim —
+ # this is what actually gated the answer, so it's what the trace
+ # must show.
+ subject_scope=subject_scope.value,
+ intent=intent.value,
+ decision=decision,
+ reason=reason,
+ resolved_drug_id=resolved_drug_id,
+ citations=tuple(item.model_dump() for item in citations),
+ )
+ except Exception:
+ metrics.increment(TRACE_WRITE_FAILED)
+ trace_id = str(uuid.uuid4())
return RagQueryResponse(
trace_id=trace_id,
decision=decision,
diff --git a/apps/ai-service/tests/test_agent.py b/apps/ai-service/tests/test_agent.py
new file mode 100644
index 0000000..726ca8a
--- /dev/null
+++ b/apps/ai-service/tests/test_agent.py
@@ -0,0 +1,212 @@
+"""`rag/agent.py` — the new orchestrator wired live in F-03.
+
+Codex's 2026-08-06 review (F-03/F-10) flagged that this module had zero test
+coverage despite being built to replace the live front end. This is the
+first coverage: routing branches with fakes, not an exhaustive golden set
+(that is F-10's job — a production-path regression suite against real
+Qdrant/Bedrock fixtures).
+"""
+from __future__ import annotations
+
+from rag.agent import RagAgent
+from rag.answer import GroundedAnswerService
+from rag.models import Evidence, EvidenceDecision, RetrievalResult, SourceRef
+from rag.understanding import QueryFrame
+
+SOURCE = SourceRef(physical_page=100, precision="region", printed_page=100)
+
+
+def _evidence(text: str) -> Evidence:
+ return Evidence(
+ evidence_id="e0", matched_doc_id="e0", kind="prose", text=text, score=1.0,
+ source_refs=(SOURCE,), hydrated_from_parent=False, requires_visual_check=False,
+ )
+
+
+class _FixedUnderstander:
+ def __init__(self, frame: QueryFrame) -> None:
+ self._frame = frame
+
+ def understand(self, turn, history=()):
+ return self._frame
+
+
+class _FixedRetrieval:
+ """Stands in for `RetrievalService.retrieve_framed` — a canned result per
+ drug_id regardless of section/query, so these tests assert routing, not
+ retrieval (that's `test_retrieval_service.py`'s job)."""
+
+ def __init__(self, results: dict[str, RetrievalResult]) -> None:
+ self._results = results
+
+ def retrieve_framed(self, drug_id, section_key, query, is_overview=False):
+ return self._results.get(
+ drug_id, RetrievalResult(EvidenceDecision.ABSTAIN, "not_configured")
+ )
+
+
+def _agent(frame: QueryFrame, results: dict[str, RetrievalResult] | None = None) -> RagAgent:
+ # `routing=None`: `answer_from_result` (the only method this path calls)
+ # never touches it — see `rag/answer.py`.
+ answers = GroundedAnswerService(routing=None)
+ return RagAgent(_FixedUnderstander(frame), _FixedRetrieval(results or {}), answers)
+
+
+def test_smalltalk_does_not_touch_retrieval():
+ agent = _agent(QueryFrame(turn_type="smalltalk"))
+ reply = agent.handle("chào bạn")
+ assert reply.decision == "answerable"
+ assert reply.turn_type == "smalltalk"
+ assert "Dược thư" in reply.answer
+
+
+def test_out_of_scope_turn_type_abstains():
+ agent = _agent(QueryFrame(turn_type="out_of_scope"))
+ reply = agent.handle("cách tiêm truyền tĩnh mạch")
+ assert reply.decision == "abstain"
+ assert reply.reason == "out_of_scope"
+
+
+def test_veterinary_phrase_abstains_even_if_the_model_missed_it():
+ # The model returned an ordinary-looking frame; the keyword backstop
+ # (`rag.policy.looks_non_human`, the same one F-02 wired server-side)
+ # still catches it before any retrieval happens.
+ agent = _agent(QueryFrame(turn_type="drug_attribute", drugs=("metformin",)))
+ reply = agent.handle("liều metformin cho chó bao nhiêu")
+ assert reply.decision == "abstain"
+ assert reply.reason == "out_of_scope"
+
+
+def test_unknown_drug_name_is_reported_not_substituted():
+ agent = _agent(QueryFrame(turn_type="drug_attribute", unknown_drugs=("aspirinol",)))
+ reply = agent.handle("liều aspirinol")
+ assert reply.decision == "abstain"
+ assert reply.reason == "drug_not_in_formulary"
+ assert "aspirinol" in reply.answer
+
+
+def test_no_drug_named_asks_which_one():
+ agent = _agent(QueryFrame(turn_type="drug_attribute"))
+ reply = agent.handle("liều dùng bao nhiêu")
+ assert reply.decision == "clarify"
+ assert reply.reason == "no_drug"
+
+
+def test_needs_clarify_frame_is_surfaced_directly():
+ agent = _agent(QueryFrame(
+ turn_type="dosing_calc", drugs=("paracetamol",),
+ needs_clarify=True, clarify_reason="Bé mấy tuổi, cân nặng bao nhiêu kg?",
+ ))
+ reply = agent.handle("liều paracetamol cho trẻ em")
+ assert reply.decision == "clarify"
+ assert reply.clarification == "Bé mấy tuổi, cân nặng bao nhiêu kg?"
+
+
+def test_single_drug_attribute_retrieves_and_answers():
+ result = RetrievalResult(
+ EvidenceDecision.ANSWERABLE, "grounded_evidence_available",
+ (_evidence("Liều 500 mg mỗi ngày."),), resolved_drug_id="metformin",
+ )
+ agent = _agent(
+ QueryFrame(turn_type="drug_attribute", drugs=("metformin",),
+ attribute="lieu_luong_va_cach_dung"),
+ {"metformin": result},
+ )
+ reply = agent.handle("liều metformin")
+ assert reply.decision == "answerable"
+ assert reply.drugs == ("metformin",)
+ assert "500 mg" in reply.answer
+
+
+def test_interaction_combines_both_drugs_evidence():
+ warfarin = RetrievalResult(
+ EvidenceDecision.ANSWERABLE, "grounded_evidence_available",
+ (_evidence("Tương tác với aspirin làm tăng nguy cơ chảy máu."),),
+ )
+ aspirin = RetrievalResult(
+ EvidenceDecision.ANSWERABLE, "grounded_evidence_available",
+ (_evidence("Tương tác với warfarin làm tăng nguy cơ chảy máu."),),
+ )
+ agent = _agent(
+ QueryFrame(turn_type="interaction", drugs=("warfarin", "aspirin")),
+ {"warfarin": warfarin, "aspirin": aspirin},
+ )
+ reply = agent.handle("warfarin với aspirin có dùng chung được không")
+ assert reply.decision == "answerable"
+ assert reply.turn_type == "interaction"
+ assert "chảy máu" in reply.answer
+
+
+def test_interaction_with_no_evidence_abstains_and_does_not_imply_safety():
+ agent = _agent(
+ QueryFrame(turn_type="interaction", drugs=("drug_a", "drug_b")), {},
+ )
+ reply = agent.handle("drug_a và drug_b dùng chung được không")
+ assert reply.decision == "abstain"
+ assert reply.reason == "no_interaction_evidence"
+ assert "KHÔNG có nghĩa là an toàn" in reply.answer
+
+
+def test_symptom_to_drug_without_a_drug_name_asks_honestly_not_wired_yet():
+ agent = _agent(QueryFrame(turn_type="symptom_to_drug", indication="sốt cao"))
+ reply = agent.handle("sốt cao uống thuốc gì")
+ assert reply.decision == "clarify"
+ assert reply.reason == "reverse_lookup_not_ready"
+
+
+def test_history_is_passed_to_the_understander_on_the_next_turn():
+ received_history: list[tuple[str, ...]] = []
+
+ class _RecordingUnderstander:
+ def understand(self, turn, history=()):
+ received_history.append(tuple(history))
+ return QueryFrame(turn_type="smalltalk")
+
+ answers = GroundedAnswerService(routing=None)
+ agent = RagAgent(_RecordingUnderstander(), _FixedRetrieval({}), answers)
+
+ agent.handle("chào bạn", conversation_id="c1")
+ agent.handle("còn liều thì sao?", conversation_id="c1")
+
+ assert received_history[0] == ()
+ assert any("chào bạn" in line for line in received_history[1])
+
+
+def test_history_is_isolated_per_conversation_id():
+ received_history: list[tuple[str, ...]] = []
+
+ class _RecordingUnderstander:
+ def understand(self, turn, history=()):
+ received_history.append(tuple(history))
+ return QueryFrame(turn_type="smalltalk")
+
+ answers = GroundedAnswerService(routing=None)
+ agent = RagAgent(_RecordingUnderstander(), _FixedRetrieval({}), answers)
+
+ agent.handle("xin chào", conversation_id="a")
+ agent.handle("liều dùng bao nhiêu", conversation_id="b")
+
+ # The second call, on a different conversation id, must not see "a"'s turn.
+ assert received_history[1] == ()
+
+
+def test_autocomplete_delegates_to_the_configured_source():
+ class _Source:
+ def complete(self, prefix, k):
+ return ["metformin_id", "metoprolol_id"]
+
+ answers = GroundedAnswerService(routing=None)
+ agent = RagAgent(
+ _FixedUnderstander(QueryFrame(turn_type="smalltalk")),
+ _FixedRetrieval({}), answers, autocomplete=_Source(),
+ )
+ assert agent.complete("met") == ["Metformin Id", "Metoprolol Id"]
+
+
+def test_autocomplete_with_no_source_configured_returns_empty():
+ answers = GroundedAnswerService(routing=None)
+ agent = RagAgent(
+ _FixedUnderstander(QueryFrame(turn_type="smalltalk")),
+ _FixedRetrieval({}), answers,
+ )
+ assert agent.complete("met") == []
diff --git a/apps/ai-service/tests/test_api.py b/apps/ai-service/tests/test_api.py
index 109eae3..35ec535 100644
--- a/apps/ai-service/tests/test_api.py
+++ b/apps/ai-service/tests/test_api.py
@@ -2,7 +2,9 @@ from fastapi.testclient import TestClient
from config import Settings
from main import create_app
-from rag.answer import GroundedAnswerService
+from rag.agent import AgentReply
+from rag.answer import Citation, GroundedAnswerService
+from rag.metrics import TRACE_WRITE_FAILED, InMemoryMetrics
from rag.models import EvidenceDecision, QueryIntent, RetrievalResult, SubjectScope
@@ -51,3 +53,127 @@ def test_query_requires_structured_scope_and_intent():
)
response = TestClient(app).post("/v1/rag/query", json={"query": "Liều?"})
assert response.status_code == 422
+
+
+# --- F-03: the app.state.conversational slot is the new RagAgent, wired live ---
+
+
+class FakeAgent:
+ """Stands in for `RagAgent` — proves `routers/rag.py` calls `.handle()`
+ and maps `AgentReply` correctly, not that the agent's own routing logic
+ is correct (that's `test_agent.py`)."""
+
+ def __init__(self, reply: AgentReply) -> None:
+ self._reply = reply
+ self.calls: list[tuple[str, str | None]] = []
+
+ def handle(self, turn: str, conversation_id: str | None = None) -> AgentReply:
+ self.calls.append((turn, conversation_id))
+ return self._reply
+
+ def complete(self, prefix: str, k: int = 8) -> list[str]:
+ return ["Metformin"] if prefix else []
+
+
+def _citation() -> Citation:
+ return Citation(
+ chunk_id="metformin::lieu::0", printed_page_start=714, printed_page_end=714,
+ physical_page=812,
+ )
+
+
+def test_query_routes_through_the_agent_when_one_is_configured():
+ agent = FakeAgent(AgentReply(
+ decision="answerable", reason="grounded_evidence_available",
+ answer="Liều 500 mg [1].", citations=(_citation(),),
+ drugs=("metformin",), turn_type="drug_attribute", generated=True,
+ ))
+ app = create_app(
+ settings=Settings(),
+ answer_service=GroundedAnswerService(FixedRouting()),
+ conversational=agent,
+ trace_writer=MemoryTraceWriter(),
+ )
+ response = TestClient(app).post("/v1/rag/query", json={
+ "query": "Liều metformin?", "subject_scope": "human",
+ "intent": "fact_lookup", "conversation_id": "c1",
+ })
+ body = response.json()
+ assert response.status_code == 200
+ assert body["decision"] == "answerable"
+ assert body["answer"] == "Liều 500 mg [1]."
+ assert body["resolved_drug_id"] == "metformin"
+ assert len(body["citations"]) == 1
+ assert agent.calls == [("Liều metformin?", "c1")]
+
+
+def test_query_agent_clarification_is_surfaced_as_the_answer():
+ agent = FakeAgent(AgentReply(
+ decision="clarify", reason="no_drug",
+ clarification="Anh/chị muốn tra thuốc nào?",
+ ))
+ app = create_app(
+ settings=Settings(),
+ answer_service=GroundedAnswerService(FixedRouting()),
+ conversational=agent,
+ trace_writer=MemoryTraceWriter(),
+ )
+ response = TestClient(app).post("/v1/rag/query", json={
+ "query": "liều dùng bao nhiêu", "subject_scope": "human",
+ "intent": "fact_lookup",
+ })
+ body = response.json()
+ assert body["decision"] == "clarify"
+ assert body["answer"] == "Anh/chị muốn tra thuốc nào?"
+ assert body["resolved_drug_id"] is None
+
+
+def test_suggest_delegates_to_the_agent_when_configured():
+ app = create_app(
+ settings=Settings(),
+ conversational=FakeAgent(AgentReply(decision="answerable", reason="x")),
+ trace_writer=MemoryTraceWriter(),
+ )
+ response = TestClient(app).get("/v1/rag/suggest", params={"q": "met"})
+ assert response.json() == {"suggestions": ["Metformin"]}
+
+
+def test_suggest_with_no_agent_configured_returns_empty():
+ app = create_app(settings=Settings(), trace_writer=MemoryTraceWriter())
+ response = TestClient(app).get("/v1/rag/suggest", params={"q": "met"})
+ assert response.json() == {"suggestions": []}
+
+
+# --- F-09: trace persistence is fail-open ------------------------------------
+
+
+class _RaisingTraceWriter:
+ """Stands in for a Postgres outage: `save()` always raises."""
+
+ def save(self, **fields):
+ raise ConnectionError("could not connect to postgres")
+
+
+def test_a_trace_write_failure_does_not_turn_a_good_answer_into_a_500():
+ agent = FakeAgent(AgentReply(
+ decision="answerable", reason="grounded_evidence_available",
+ answer="Liều 500 mg [1].", citations=(_citation(),),
+ drugs=("metformin",), turn_type="drug_attribute", generated=True,
+ ))
+ metrics = InMemoryMetrics()
+ app = create_app(
+ settings=Settings(),
+ answer_service=GroundedAnswerService(FixedRouting()),
+ conversational=agent,
+ trace_writer=_RaisingTraceWriter(),
+ metrics=metrics,
+ )
+ response = TestClient(app).post("/v1/rag/query", json={
+ "query": "Liều metformin?", "subject_scope": "human", "intent": "fact_lookup",
+ })
+ body = response.json()
+
+ assert response.status_code == 200
+ assert body["answer"] == "Liều 500 mg [1]."
+ assert body["trace_id"] # a locally-generated fallback id, still present
+ assert metrics.total(TRACE_WRITE_FAILED) == 1
diff --git a/apps/ai-service/tests/test_bootstrap.py b/apps/ai-service/tests/test_bootstrap.py
new file mode 100644
index 0000000..825e431
--- /dev/null
+++ b/apps/ai-service/tests/test_bootstrap.py
@@ -0,0 +1,38 @@
+"""`_catalog_names` — the drug names shown to `LlmQueryUnderstander`.
+
+Reproduces a live 2026-08-06 bug: picking the first N aliases alphabetically
+could drop a drug's own recognizable name entirely, breaking multi-turn
+follow-ups where the drug is no longer restated in the raw turn text (see
+`docs/progress-log.md` for the exact failure: "Liều paracetamol cho trẻ em"
+-> two clarify rounds -> "Không tìm thấy paracetamol").
+"""
+from bootstrap import _catalog_names
+
+
+def test_alphabetically_early_junk_alias_does_not_bury_the_canonical_name():
+ aliases = {
+ "paracetamol_acetaminophen": {
+ "0Frezefev", "ABAB", "Ace kid 80", "PARACETAMOL", "Acetaminophen",
+ },
+ }
+ shown = _catalog_names(aliases)["paracetamol_acetaminophen"]
+ assert "paracetamol acetaminophen" in shown
+
+
+def test_canonical_name_is_always_first():
+ aliases = {"metformin": {"METFORMIN", "Axiol", "Dybis", "Zzyzx"}}
+ shown = _catalog_names(aliases)["metformin"]
+ assert shown.split(", ")[0] == "metformin"
+
+
+def test_drug_with_no_aliases_still_shows_its_canonical_name():
+ shown = _catalog_names({"some_drug": set()})["some_drug"]
+ assert shown == "some drug"
+
+
+def test_output_is_capped_and_does_not_duplicate_the_canonical_name():
+ aliases = {"drug_x": {f"Brand{i}" for i in range(20)} | {"DRUG X", "drug x"}}
+ shown = _catalog_names(aliases)["drug_x"].split(", ")
+ assert shown[0] == "drug x"
+ assert shown.count("drug x") == 1
+ assert len(shown) <= 3
diff --git a/apps/ai-service/tests/test_citation_and_intro.py b/apps/ai-service/tests/test_citation_and_intro.py
index db4dc96..d8dfd0b 100644
--- a/apps/ai-service/tests/test_citation_and_intro.py
+++ b/apps/ai-service/tests/test_citation_and_intro.py
@@ -41,11 +41,23 @@ class _Routing:
class _Generator:
- def __init__(self, payload: dict) -> None:
+ def __init__(self, payload: dict, entailment_payload: dict | None = None) -> None:
self._payload = payload
+ self._entailment_payload = entailment_payload or {
+ "entailed": True,
+ "unsupported": [],
+ }
def generate(self, system: str, user: str, schema: dict) -> str: # noqa: ARG002
- return json.dumps(self._payload, ensure_ascii=False)
+ # `_generate` also runs a post-generation entailment check; tell the
+ # two request shapes apart by schema so callers here only need to
+ # fake the main answer, not both.
+ payload = (
+ self._entailment_payload
+ if "entailed" in schema.get("properties", {})
+ else self._payload
+ )
+ return json.dumps(payload, ensure_ascii=False)
def _answerable(*evidence: Evidence, is_overview: bool = False) -> RetrievalResult:
@@ -72,17 +84,56 @@ def test_only_cited_sources_are_returned():
assert grounded.citations[0].printed_page_start == 200
-def test_answer_citing_nothing_falls_back_to_all_citations():
+def test_answer_citing_nothing_is_rejected_not_dressed_up_with_borrowed_citations():
result = _answerable(_evidence(0, 100), _evidence(1, 200))
service = GroundedAnswerService(
_Routing(result),
- # no [n] marker at all: rather than show zero provenance, show all.
+ # no [n] marker at all: grounding.verify rejects this outright (an
+ # uncited claim, per F-01). A generator is configured, so the
+ # rejection abstains — it must not attach every retrieved citation
+ # to dress an uncited generation up as sourced (the old behavior),
+ # and it must not silently degrade to a raw extractive quote either
+ # (owner correction, 2026-08-06: no fallback to the retired
+ # offline-extractive shape when a real generator is configured).
_Generator({"answer": "Không có trích dẫn.", "evidence_sufficient": True}),
)
grounded = service.answer("q", SubjectScope.HUMAN, QueryIntent.FACT_LOOKUP)
- assert len(grounded.citations) == 2
+ assert grounded.generated is False
+ assert grounded.answer is None
+ assert grounded.citations == ()
+ assert grounded.result.decision == EvidenceDecision.ABSTAIN
+
+
+def test_underspecified_dose_asks_instead_of_dumping():
+ """The reasoning step: a dose question spanning bands with no age/weight is
+ turned into a clarification, not the whole section."""
+ result = _answerable(_evidence(0, 100), _evidence(1, 200))
+ gen = _Generator(
+ {"sufficient": False,
+ "clarifying_question": "Bé mấy tuổi, cân nặng bao nhiêu kg?"}
+ )
+ service = GroundedAnswerService(_Routing(result), gen)
+
+ g = service.answer("paracetamol cho trẻ em", SubjectScope.HUMAN, QueryIntent.FACT_LOOKUP)
+
+ assert g.clarification is not None
+ assert "tuổi" in g.clarification
+ assert g.answer == g.clarification
+ assert g.generated is False
+
+
+def test_sufficient_query_is_not_turned_into_a_clarification():
+ result = _answerable(_evidence(0, 100), _evidence(1, 200))
+ gen = _Generator({"sufficient": True, "clarifying_question": None})
+ service = GroundedAnswerService(_Routing(result), gen)
+
+ g = service.answer("liều người lớn", SubjectScope.HUMAN, QueryIntent.FACT_LOOKUP)
+
+ # sufficiency passes; generation then runs (its payload lacks answer keys, so
+ # it falls back to the source text) — the point is no clarification fired.
+ assert g.clarification is None
def test_bare_name_builds_an_intro_prompt():
diff --git a/apps/ai-service/tests/test_conversation.py b/apps/ai-service/tests/test_conversation.py
index 57c7695..9a90adf 100644
--- a/apps/ai-service/tests/test_conversation.py
+++ b/apps/ai-service/tests/test_conversation.py
@@ -136,6 +136,49 @@ def test_recent_window_evicts_oldest():
assert state.turn_count == 8
+def test_evicted_turns_reach_overflow_not_silently_dropped():
+ """Bug fixed 2026-08-06 (Codex review, F-06): `overflow()` used to check
+ `len(self.recent) > window`, but `append()` already truncates `recent`
+ to `window`, so that comparison could never be true — evicted turns
+ never reached the summariser no matter how long a conversation ran.
+ Exact repro from the review: 8 turns into a window of 6."""
+ state = ConversationState("c1")
+ for index in range(8):
+ state = state.append(Turn("user", f"q{index}", "2026-08-05"), window=6)
+
+ overflow = state.overflow()
+
+ assert [turn.text for turn in overflow] == ["q0", "q1"]
+
+
+def test_overflow_accumulates_across_the_two_appends_one_turn_makes():
+ """A live turn typically calls `append()` twice in a row (user, then
+ assistant). Each can evict at most one turn; the second call's overflow
+ must not overwrite, and so lose, the first's."""
+ state = ConversationState("c1")
+ for index in range(6):
+ state = state.append(Turn("user", f"q{index}", "2026-08-05"), window=6)
+ assert state.overflow() == () # window exactly full, nothing evicted yet
+
+ state = state.append(Turn("user", "q6", "2026-08-05"), window=6)
+ state = state.append(Turn("assistant", "a6", "2026-08-05"), window=6)
+
+ assert [turn.text for turn in state.overflow()] == ["q0", "q1"]
+
+
+def test_overflow_is_empty_again_after_the_caller_clears_it():
+ from dataclasses import replace
+
+ state = ConversationState("c1")
+ for index in range(8):
+ state = state.append(Turn("user", f"q{index}", "2026-08-05"), window=6)
+ assert state.overflow() != ()
+
+ state = replace(state, pending_overflow=())
+
+ assert state.overflow() == ()
+
+
def test_focus_update_stamps_the_current_turn():
state = _state(turn_count=3)
diff --git a/apps/ai-service/tests/test_conversational_loop.py b/apps/ai-service/tests/test_conversational_loop.py
index f050f7e..c96950d 100644
--- a/apps/ai-service/tests/test_conversational_loop.py
+++ b/apps/ai-service/tests/test_conversational_loop.py
@@ -75,7 +75,7 @@ def test_followup_inherits_drug_and_passes_it_resolved():
svc.answer("c3", "chống chỉ định metformin", SubjectScope.HUMAN, QueryIntent.FACT_LOOKUP)
out = svc.answer("c3", "còn trẻ em thì sao?", SubjectScope.HUMAN, QueryIntent.FACT_LOOKUP)
assert out.inherited_drug == "metformin"
- assert out.answer.startswith("Về metformin:")
+ assert out.answer.startswith("Về Metformin:")
# The inherited drug is passed already-resolved (not re-resolved from the
# rewritten turn text), and the raw follow-up drives section routing.
last_query, last_drug_id = answers.calls[-1]
diff --git a/apps/ai-service/tests/test_grounded_generation.py b/apps/ai-service/tests/test_grounded_generation.py
index c9a32ae..a2470f1 100644
--- a/apps/ai-service/tests/test_grounded_generation.py
+++ b/apps/ai-service/tests/test_grounded_generation.py
@@ -65,23 +65,48 @@ class _FixedRouting:
class _Generator:
- """Returns whatever payload the test wants the model to have produced."""
+ """Returns whatever payload the test wants the model to have produced.
- def __init__(self, payload) -> None:
+ `_generate` now makes up to four calls through this port: the main
+ answer, a sufficiency check (skipped here — one evidence block), and up
+ to two entailment calls (a reject retries once — live probing found the
+ judge noisy on an identical claim/evidence pair). They're told apart by
+ schema, so a test that only cares about the main answer doesn't also
+ have to fake an entailment response by hand; `entailment_payload`
+ overrides it when a test wants the entailment pass to reject. Pass a
+ list of payloads to get a different answer on each successive
+ entailment call (e.g. `[reject, accept]` for the retry-recovers case).
+ """
+
+ def __init__(self, payload, entailment_payload=None) -> None:
self._payload = payload
+ default = {"entailed": True, "unsupported": []}
+ payloads = entailment_payload if entailment_payload is not None else default
+ self._entailment_payloads = (
+ list(payloads) if isinstance(payloads, list) else [payloads]
+ )
+ self._entailment_call = 0
def generate(self, system: str, user: str, schema: dict) -> str:
- if isinstance(self._payload, BaseException):
- raise self._payload
- if isinstance(self._payload, str):
- return self._payload
- return json.dumps(self._payload, ensure_ascii=False)
+ if "entailed" in schema.get("properties", {}):
+ index = min(self._entailment_call, len(self._entailment_payloads) - 1)
+ payload = self._entailment_payloads[index]
+ self._entailment_call += 1
+ else:
+ payload = self._payload
+ if isinstance(payload, BaseException):
+ raise payload
+ if isinstance(payload, str):
+ return payload
+ return json.dumps(payload, ensure_ascii=False)
-def _answer(payload, result: RetrievalResult | None = None):
+def _answer(payload, result: RetrievalResult | None = None, entailment_payload=None):
metrics = InMemoryMetrics()
service = GroundedAnswerService(
- _FixedRouting(result or _result()), _Generator(payload), metrics
+ _FixedRouting(result or _result()),
+ _Generator(payload, entailment_payload),
+ metrics,
)
grounded = service.answer("Liều Metformin?", SubjectScope.HUMAN, QueryIntent.FACT_LOOKUP)
return grounded, metrics
@@ -97,8 +122,12 @@ def test_invented_dose_is_refused_and_never_reaches_the_answer():
)
assert grounded.generated is False
- assert "850" not in grounded.answer
- assert grounded.answer.startswith(EVIDENCE_TEXT)
+ # A generator is configured, so a rejected generation abstains — it does
+ # NOT silently degrade to a raw source dump (owner correction, 2026-08-06:
+ # this is a real LLM chatbot, not the retired offline-extractive build).
+ assert grounded.answer is None
+ assert grounded.result.decision == EvidenceDecision.ABSTAIN
+ assert grounded.result.reason == "generation_unavailable"
assert metrics.total(GENERATION_REJECTED, reason="ungrounded_number") == 1
@@ -119,7 +148,12 @@ def test_citation_pointing_at_nothing_is_refused():
)
assert grounded.generated is False
- assert metrics.total(GENERATION_REJECTED, reason="invalid_citation") == 1
+ # [3] is out of range with one evidence block, so "500" has no valid
+ # citation to bind to — grounding.verify now flags it as unsupported
+ # rather than letting it pass because 500 happens to exist somewhere in
+ # the (single) evidence block anyway. ungrounded_number takes priority
+ # over invalid_citation in GroundingReport.reason; both are present.
+ assert metrics.total(GENERATION_REJECTED, reason="ungrounded_number") == 1
def test_faithful_rewrite_is_served():
@@ -134,6 +168,95 @@ def test_faithful_rewrite_is_served():
assert metrics.total(GENERATION_REJECTED) == 0
+# --- the entailment pass: catches what number/citation checks structurally can't -----
+
+
+def test_fabricated_nonnumeric_claim_with_a_valid_citation_is_rejected():
+ """Reproduces `claim_bia` from the Codex 2026-08-06 review end to end:
+ right drug, syntactically valid citation, fabricated indication.
+ grounding.verify alone cannot see this (no number, citation in range) —
+ the entailment pass, told the model judged evidence 1 does not support
+ it, is what rejects the generation."""
+ grounded, metrics = _answer(
+ {"answer": "Metformin chữa ung thư [1].", "evidence_sufficient": True},
+ entailment_payload={"entailed": False, "unsupported": [1]},
+ )
+
+ assert grounded.generated is False
+ assert grounded.answer is None
+ assert grounded.result.decision == EvidenceDecision.ABSTAIN
+ assert metrics.total(GENERATION_REJECTED, reason="unsupported_claim") == 1
+
+
+def test_entailment_check_running_and_passing_still_serves_the_answer():
+ grounded, metrics = _answer(
+ {"answer": "Metformin dùng điều trị đái tháo đường [1].",
+ "evidence_sufficient": True},
+ entailment_payload={"entailed": True, "unsupported": []},
+ )
+
+ assert grounded.generated is True
+ assert metrics.total(GENERATION_SERVED) == 1
+
+
+def test_entailment_retries_once_after_a_reject_and_a_later_accept_serves():
+ """Reproduces the 2026-08-06 live finding: the same claim/evidence pair,
+ called three times through the real judge, came back entailed twice and
+ rejected once — a single noisy reject must not discard a correct,
+ well-cited answer."""
+ grounded, metrics = _answer(
+ {"answer": "Metformin dùng điều trị đái tháo đường [1].",
+ "evidence_sufficient": True},
+ entailment_payload=[
+ {"entailed": False, "unsupported": [1]},
+ {"entailed": True, "unsupported": []},
+ ],
+ )
+
+ assert grounded.generated is True
+ assert metrics.total(GENERATION_SERVED) == 1
+
+
+def test_entailment_two_agreeing_rejects_still_discard():
+ grounded, metrics = _answer(
+ {"answer": "Metformin chữa ung thư [1].", "evidence_sufficient": True},
+ entailment_payload=[
+ {"entailed": False, "unsupported": [1]},
+ {"entailed": False, "unsupported": [1]},
+ ],
+ )
+
+ assert grounded.generated is False
+ assert grounded.answer is None
+ assert metrics.total(GENERATION_REJECTED, reason="unsupported_claim") == 1
+
+
+def test_entailment_provider_outage_fails_closed_to_abstain():
+ grounded, metrics = _answer(
+ {"answer": "Người lớn: 500 mg, 2 lần/ngày [1].", "evidence_sufficient": True},
+ entailment_payload=AnswerGenerationUnavailable(),
+ )
+
+ assert grounded.generated is False
+ assert grounded.answer is None
+ assert grounded.result.decision == EvidenceDecision.ABSTAIN
+ assert metrics.total(GENERATION_REJECTED, reason="unsupported_claim") == 1
+
+
+def test_entailment_check_is_skipped_when_the_answer_has_no_claim_text():
+ """An answer that is nothing but a citation marker has no claim text for
+ an entailment pass to check against — `_verify_entailment` must not call
+ the provider at all. Proven by making that call raise: if the skip
+ didn't fire, this would reject rather than serve the answer."""
+ grounded, metrics = _answer(
+ {"answer": "[1]", "evidence_sufficient": True},
+ entailment_payload=AnswerGenerationUnavailable(),
+ )
+
+ assert grounded.generated is True
+ assert metrics.total(GENERATION_REJECTED) == 0
+
+
def test_citations_survive_generation():
"""Provenance is the point; a prettier answer must not cost the folio."""
grounded, _ = _answer(
@@ -145,7 +268,7 @@ def test_citations_survive_generation():
assert grounded.citations[0].printed_page_start == 714
-# --- degradation is always to the source, never to an error -------------------
+# --- a configured generator that fails abstains, never a raw source dump -----
@pytest.mark.parametrize(
@@ -158,11 +281,13 @@ def test_citations_survive_generation():
({"answer": "...", "evidence_sufficient": False}, "evidence_insufficient"),
],
)
-def test_every_generation_failure_falls_back_to_the_source_text(payload, reason):
+def test_every_generation_failure_abstains_instead_of_a_raw_source_dump(payload, reason):
grounded, metrics = _answer(payload)
assert grounded.generated is False
- assert grounded.answer.startswith(EVIDENCE_TEXT)
+ assert grounded.answer is None
+ assert grounded.result.decision == EvidenceDecision.ABSTAIN
+ assert grounded.result.reason == "generation_unavailable"
assert metrics.total(GENERATION_REJECTED, reason=reason) == 1
diff --git a/apps/ai-service/tests/test_grounding.py b/apps/ai-service/tests/test_grounding.py
new file mode 100644
index 0000000..dca4f40
--- /dev/null
+++ b/apps/ai-service/tests/test_grounding.py
@@ -0,0 +1,122 @@
+"""Adversarial regression cases for `rag.grounding.verify`.
+
+Each case reproduces a defect found by Codex's 2026-08-06 code review
+(`coordination/CODEX_RAG_CODE_REVIEW_2026-08-06.md`, F-01) against the old
+implementation, which pooled every evidence number into one global set and
+never required a citation at all. `verify` must now bind a claim's numbers
+to only the evidence block(s) its own citation group names, and must reject
+a claim with no citation regardless of whether it contains a number.
+"""
+from rag import grounding
+
+
+def test_number_from_the_wrong_evidence_block_is_rejected():
+ # Reproduces `so_sai_nguon`: 500 mg is real, but only in evidence 2 —
+ # citing [1] for it must fail, not pass because 500 exists *somewhere*.
+ report = grounding.verify(
+ "Liều 500 mg [1].",
+ ("Không dùng khi suy thận.", "Liều 500 mg mỗi ngày."),
+ )
+ assert not report.grounded
+ assert report.unsupported_numbers == ("500",)
+ assert report.reason == "ungrounded_number"
+
+
+def test_citing_the_correct_block_for_the_number_is_grounded():
+ report = grounding.verify(
+ "Liều 500 mg [2].",
+ ("Không dùng khi suy thận.", "Liều 500 mg mỗi ngày."),
+ )
+ assert report.grounded
+ assert report.cited_indices == (2,)
+
+
+def test_answer_with_no_citation_at_all_is_rejected():
+ # Reproduces `khong_citation`: a number that is genuinely in the evidence
+ # still must not pass when the answer never cites anything.
+ report = grounding.verify("Liều 500 mg.", ("Liều 500 mg mỗi ngày.",))
+ assert not report.grounded
+ assert report.uncited_claim
+ assert report.unsupported_numbers == ("500",)
+ assert report.reason == "ungrounded_number"
+
+
+def test_nonnumeric_claim_with_no_citation_is_rejected():
+ report = grounding.verify(
+ "Chống chỉ định với suy gan nặng.", ("Chống chỉ định: suy gan nặng.",)
+ )
+ assert not report.grounded
+ assert report.uncited_claim
+ assert report.reason == "uncited_claim"
+
+
+def test_trailing_text_after_the_last_citation_needs_its_own_citation():
+ report = grounding.verify(
+ "Liều 500 mg [1]. Không dùng khi suy thận.",
+ ("Liều 500 mg mỗi ngày.",),
+ )
+ assert not report.grounded
+ assert report.uncited_claim
+
+
+def test_out_of_range_citation_is_invalid_and_leaves_its_claim_unsupported():
+ report = grounding.verify("Liều 500 mg [3].", ("Liều 500 mg mỗi ngày.",))
+ assert not report.grounded
+ assert report.invalid_citations == (3,)
+ assert report.unsupported_numbers == ("500",)
+ assert report.reason == "ungrounded_number"
+
+
+def test_two_claims_each_binding_correctly_to_their_own_source_is_grounded():
+ report = grounding.verify(
+ "Người lớn 500 mg [1]. Trẻ em 250 mg [2].",
+ ("Liều người lớn 500 mg.", "Liều trẻ em 250 mg."),
+ )
+ assert report.grounded
+ assert report.cited_indices == (1, 2)
+
+
+def test_second_claim_citing_the_first_blocks_source_number_is_rejected():
+ # The child dose (250) is real, but only in evidence 2; citing [1] for it
+ # is the same defect as `so_sai_nguon`, just in a second sentence.
+ report = grounding.verify(
+ "Người lớn 500 mg [1]. Trẻ em 250 mg [1].",
+ ("Liều người lớn 500 mg.", "Liều trẻ em 250 mg."),
+ )
+ assert not report.grounded
+ assert report.unsupported_numbers == ("250",)
+
+
+def test_multiple_markers_on_one_claim_check_against_their_union():
+ report = grounding.verify(
+ "Liều 500 mg [1][2].",
+ ("Liều người lớn 500 mg.", "Liều mỗi ngày."),
+ )
+ assert report.grounded
+ assert report.cited_indices == (1, 2)
+
+
+def test_extractive_quoting_format_is_still_grounded():
+ # The service's extractive fallback formats each block as `text [n]`.
+ report = grounding.verify(
+ "Liều được ghi trong nguồn. [1]", ("Liều được ghi trong nguồn.",)
+ )
+ assert report.grounded
+
+
+def test_decimal_separator_is_compared_verbatim_not_normalised():
+ report = grounding.verify("Liều 7.5 mg [1].", ("Liều 7,5 mg.",))
+ assert not report.grounded
+ assert report.unsupported_numbers == ("7.5",)
+
+
+def test_known_gap_fabricated_nonnumeric_claim_with_a_valid_citation_still_passes():
+ # Documents the residual gap `grounding.verify` cannot close on its own
+ # (see module docstring): a citation-bearing claim whose content the
+ # cited block does not actually support. Closed by a separate LLM
+ # entailment pass in `rag/answer.py`, not by this regex-only check.
+ report = grounding.verify(
+ "Metformin chữa ung thư [1].",
+ ("Metformin dùng điều trị đái tháo đường.",),
+ )
+ assert report.grounded
diff --git a/apps/ai-service/tests/test_live_datastores.py b/apps/ai-service/tests/test_live_datastores.py
index 10547d6..4f6f598 100644
--- a/apps/ai-service/tests/test_live_datastores.py
+++ b/apps/ai-service/tests/test_live_datastores.py
@@ -20,6 +20,35 @@ PDF = ROOT / "ingestion/data/raw/duoc-thu-quoc-gia-viet-nam-2018.pdf"
MIGRATION = Path(__file__).resolve().parents[1] / "migrations/001_rag_retrieval_trace.sql"
+class _PlumbingEmbedder:
+ """Deterministic local vectors for the Qdrant round-trip plumbing tests.
+
+ Not a semantic model — it exists only so an integration test can upsert and
+ query real chunks without a cloud call. Production has exactly one query
+ embedder (`BedrockCohereQueryEmbedder`); the old local/section-only stubs
+ were removed, so this lives with the test that needs it.
+ """
+
+ def __init__(self, dimensions: int) -> None:
+ self._dimensions = dimensions
+
+ @property
+ def dimensions(self) -> int:
+ return self._dimensions
+
+ def embed_query(self, text: str) -> list[float]:
+ import hashlib
+ import math
+
+ vector = [0.0] * self._dimensions
+ for token in text.casefold().split():
+ digest = hashlib.sha256(token.encode("utf-8")).digest()
+ index = int.from_bytes(digest[:4], "big") % self._dimensions
+ vector[index] += 1.0 if digest[4] & 1 else -1.0
+ norm = math.sqrt(sum(value * value for value in vector))
+ return [value / norm for value in vector] if norm else vector
+
+
@lru_cache
def _first_real_chunk() -> dict:
with CHUNKS.open(encoding="utf-8") as handle:
@@ -42,12 +71,11 @@ def test_real_qdrant_round_trip_uses_real_chunk_and_printed_folio():
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, PointStruct, VectorParams
- from adapters.embedding import LocalHashQueryEmbedder
from adapters.qdrant import QdrantRetriever
client = QdrantClient(url="http://localhost:6333")
collection = f"integration_{uuid.uuid4().hex}"
- embedder = LocalHashQueryEmbedder(32)
+ embedder = _PlumbingEmbedder(32)
record = _first_real_chunk()
try:
client.create_collection(
@@ -103,12 +131,144 @@ def test_real_postgres_migration_insert_and_read_back():
assert stored.citations[0]["printed_page_start"] == 101
+class _FakeJsonLlm:
+ """Deterministic stand-in for the Bedrock Converse generator. Satisfies
+ both `JsonLlm` (query understanding) and `AnswerGenerator` (answer +
+ entailment) — both ports are just `generate(system, user, schema) ->
+ str` — and tells the three call shapes apart the same way
+ `tests/test_grounded_generation.py`'s fake does: by schema shape.
+
+ Deliberately not a real Bedrock call: this suite asserts exact
+ drug id / section / citation / grounding outcomes, and this session's
+ own live probing (`docs/progress-log.md`, F-01/F-03 entries) found real
+ generation and entailment calls genuinely non-deterministic — the wrong
+ foundation for a regression assertion. The wiring under test — real
+ `RagAgent`, real `RetrievalService`/`QdrantRetriever` against a real
+ (temporary) Qdrant collection, real `GroundedAnswerService` — is
+ identical to production; only the cloud model call is faked.
+ """
+
+ def __init__(self, frame_payload: dict, answer_payload: dict) -> None:
+ self._frame_payload = frame_payload
+ self._answer_payload = answer_payload
+
+ def generate(self, system: str, user: str, schema: dict) -> str:
+ if "turn_type" in schema: # FRAME_SCHEMA: flat, not JSON-Schema-shaped
+ return json.dumps(self._frame_payload, ensure_ascii=False)
+ if "entailed" in schema.get("properties", {}):
+ return json.dumps({"entailed": True, "unsupported": []})
+ return json.dumps(self._answer_payload, ensure_ascii=False)
+
+
+def test_real_rag_agent_end_to_end_through_the_http_api():
+ """F-10: the *production* orchestrator (`RagAgent`), driven through the
+ real `/v1/rag/query` HTTP endpoint against a real (temporary) Qdrant
+ collection seeded with one real corpus chunk, with a real Postgres trace
+ persisted and read back. Reproduces Codex's exact 2026-08-06 finding —
+ "no current test imports RagAgent, LlmQueryUnderstander, QueryFrame, or
+ retrieve_framed" and "the evaluation runner constructs an in-memory
+ lexical retriever and the old resolver rather than executing the same
+ dependency graph as the live HTTP service" — both false as of this test.
+ """
+ from fastapi.testclient import TestClient
+ from qdrant_client import QdrantClient
+ from qdrant_client.models import Distance, PointStruct, VectorParams
+
+ from adapters.postgres import PostgresTraceRepository
+ from adapters.qdrant import QdrantParentStore, QdrantRetriever
+ from config import Settings
+ from main import create_app
+ from rag.agent import RagAgent
+ from rag.answer import GroundedAnswerService
+ from rag.routing import CatalogDrugResolver, QueryRoutingService
+ from rag.service import EvidencePolicy, RetrievalService
+ from rag.understanding import LlmQueryUnderstander
+
+ qdrant = QdrantClient(url="http://localhost:6333")
+ collection = f"integration_{uuid.uuid4().hex}"
+ embedder = _PlumbingEmbedder(32)
+ record = dict(_first_real_chunk())
+ traces = PostgresTraceRepository(
+ "postgresql://duoc_thu:duoc_thu@localhost:5432/duoc_thu"
+ )
+ traces.migrate(MIGRATION)
+ try:
+ qdrant.create_collection(
+ collection_name=collection,
+ vectors_config=VectorParams(size=32, distance=Distance.COSINE),
+ )
+ qdrant.upsert(
+ collection_name=collection,
+ points=[PointStruct(
+ id=str(uuid.uuid4()),
+ vector=embedder.embed_query(record["text"]),
+ payload=record,
+ )],
+ wait=True,
+ )
+ retrieval = RetrievalService(
+ QdrantRetriever(qdrant, collection, embedder),
+ QdrantParentStore(qdrant, collection),
+ EvidencePolicy(minimum_score=0.01),
+ )
+ # A minimal one-drug catalog, real drug_id — F-04's candidate
+ # bounding runs for real here (`RagAgent`/`LlmQueryUnderstander`
+ # are not mocked), so the query must literally name the drug for
+ # the deterministic resolver to find it as a candidate.
+ resolver = CatalogDrugResolver({record["drug_id"]: {record["drug_name"]}})
+ llm = _FakeJsonLlm(
+ frame_payload={
+ "turn_type": "drug_attribute", "drugs": [record["drug_id"]],
+ "unknown_drugs": [], "attribute": None, "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ },
+ answer_payload={
+ "answer": f"{record['text']} [1].", "evidence_sufficient": True,
+ },
+ )
+ understander = LlmQueryUnderstander(
+ llm, {record["drug_id"]: record["drug_name"]}, resolver,
+ )
+ answers = GroundedAnswerService(
+ QueryRoutingService(retrieval, resolver), generator=llm,
+ )
+ agent = RagAgent(understander, retrieval, answers)
+
+ app = create_app(
+ settings=Settings(), answer_service=answers,
+ conversational=agent, trace_writer=traces,
+ )
+ response = TestClient(app).post("/v1/rag/query", json={
+ "query": record["drug_name"],
+ "subject_scope": "human", "intent": "fact_lookup",
+ })
+
+ assert response.status_code == 200
+ body = response.json()
+ assert body["decision"] == "answerable"
+ assert body["resolved_drug_id"] == record["drug_id"]
+ assert body["citations"][0]["chunk_id"] == record["chunk_id"]
+ assert body["citations"][0]["printed_page_start"] == (
+ record["printed_page_range"][0]
+ )
+ assert record["text"] in body["answer"]
+
+ stored = traces.get(body["trace_id"])
+ assert stored is not None
+ assert stored.decision == "answerable"
+ assert stored.resolved_drug_id == record["drug_id"]
+ assert stored.citations[0]["chunk_id"] == record["chunk_id"]
+ finally:
+ with suppress(Exception):
+ qdrant.delete_collection(collection)
+
+
def test_api_round_trip_uses_qdrant_and_persists_postgres_trace():
from fastapi.testclient import TestClient
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, PointStruct, VectorParams
- from adapters.embedding import LocalHashQueryEmbedder
from adapters.postgres import PostgresTraceRepository
from adapters.qdrant import QdrantParentStore, QdrantRetriever
from config import Settings
@@ -119,7 +279,7 @@ def test_api_round_trip_uses_qdrant_and_persists_postgres_trace():
qdrant = QdrantClient(url="http://localhost:6333")
collection = f"integration_{uuid.uuid4().hex}"
- embedder = LocalHashQueryEmbedder(32)
+ embedder = _PlumbingEmbedder(32)
record = dict(_first_real_chunk())
traces = PostgresTraceRepository(
"postgresql://duoc_thu:duoc_thu@localhost:5432/duoc_thu"
diff --git a/apps/ai-service/tests/test_manifest.py b/apps/ai-service/tests/test_manifest.py
new file mode 100644
index 0000000..f593ecb
--- /dev/null
+++ b/apps/ai-service/tests/test_manifest.py
@@ -0,0 +1,50 @@
+"""F-05: the service must refuse to start on a corpus/model manifest
+mismatch, not silently search with vectors the collection wasn't built from.
+"""
+import pytest
+
+from rag.manifest import ManifestMismatch, check_manifest, manifest_collection
+
+
+def test_manifest_collection_naming():
+ assert manifest_collection("duocthu_v1") == "duocthu_v1__manifest"
+
+
+def test_matching_manifest_passes():
+ check_manifest(
+ {"model_id": "cohere.embed-v4:0", "dimensions": 1024},
+ "duocthu_v1", "cohere.embed-v4:0", 1024,
+ ) # no raise
+
+
+def test_missing_manifest_refuses():
+ with pytest.raises(ManifestMismatch, match="no corpus manifest"):
+ check_manifest(None, "duocthu_v1", "cohere.embed-v4:0", 1024)
+
+
+def test_wrong_model_id_refuses_even_with_matching_dimensions():
+ """The exact scenario the finding names: two unrelated models can both
+ produce 1024-dim vectors."""
+ with pytest.raises(ManifestMismatch, match="model_id"):
+ check_manifest(
+ {"model_id": "amazon.titan-embed-text-v2:0", "dimensions": 1024},
+ "duocthu_v1", "cohere.embed-v4:0", 1024,
+ )
+
+
+def test_wrong_dimensions_refuses():
+ with pytest.raises(ManifestMismatch, match="dimensions"):
+ check_manifest(
+ {"model_id": "cohere.embed-v4:0", "dimensions": 768},
+ "duocthu_v1", "cohere.embed-v4:0", 1024,
+ )
+
+
+def test_both_mismatched_reports_both():
+ with pytest.raises(ManifestMismatch) as excinfo:
+ check_manifest(
+ {"model_id": "other-model", "dimensions": 768},
+ "duocthu_v1", "cohere.embed-v4:0", 1024,
+ )
+ assert "model_id" in str(excinfo.value)
+ assert "dimensions" in str(excinfo.value)
diff --git a/apps/ai-service/tests/test_policy.py b/apps/ai-service/tests/test_policy.py
new file mode 100644
index 0000000..7aa27ac
--- /dev/null
+++ b/apps/ai-service/tests/test_policy.py
@@ -0,0 +1,50 @@
+"""F-02: subject_scope must be server-derived, not client-asserted.
+
+Reproduces the Codex 2026-08-06 finding directly: a caller claiming "human"
+on a veterinary query must not get that claim honored. The claim can only
+make the result MORE conservative, never less.
+
+Scoped to `subject_scope` only. `resolve_query_intent`/`QueryIntent.RECOMMENDATION`
+keyword detection was tried and removed the same day: this product is for
+doctors and pharmacists, and a clinician asking "nên dùng thuốc gì" is a
+normal professional use of a formulary reference, not something to abstain
+on. See `docs/progress-log.md` 2026-08-06.
+"""
+from rag.models import SubjectScope
+from rag.policy import resolve_subject_scope
+
+
+def test_veterinary_query_is_non_human_even_when_client_claims_human():
+ scope = resolve_subject_scope("Liều cho chó bị viêm khớp?", SubjectScope.HUMAN)
+ assert scope == SubjectScope.NON_HUMAN
+
+
+def test_ordinary_dose_question_stays_human():
+ scope = resolve_subject_scope("Liều metformin cho người lớn?", SubjectScope.HUMAN)
+ assert scope == SubjectScope.HUMAN
+
+
+def test_client_cannot_widen_a_server_detected_non_human_scope():
+ # Even an explicit non_human claim from an honest client must stick.
+ scope = resolve_subject_scope("thuốc cho mèo", SubjectScope.NON_HUMAN)
+ assert scope == SubjectScope.NON_HUMAN
+
+
+def test_client_narrowing_to_non_human_is_honored_even_without_a_keyword_hit():
+ # A caller with better information than the keyword list can narrow.
+ scope = resolve_subject_scope("liều dùng", SubjectScope.NON_HUMAN)
+ assert scope == SubjectScope.NON_HUMAN
+
+
+def test_unknown_claim_with_no_server_signal_stays_unknown():
+ scope = resolve_subject_scope("liều dùng", SubjectScope.UNKNOWN)
+ assert scope == SubjectScope.UNKNOWN
+
+
+def test_clinician_asking_which_drug_is_preferred_is_not_flagged_non_human():
+ # A doctor/pharmacist comparing options across monographs is core,
+ # intended use of this product — not a request to gate.
+ scope = resolve_subject_scope(
+ "Bệnh nhân suy thận, nên dùng thuốc hạ áp nào?", SubjectScope.HUMAN
+ )
+ assert scope == SubjectScope.HUMAN
diff --git a/apps/ai-service/tests/test_retrieval_service.py b/apps/ai-service/tests/test_retrieval_service.py
index ad317a3..5d7c8f4 100644
--- a/apps/ai-service/tests/test_retrieval_service.py
+++ b/apps/ai-service/tests/test_retrieval_service.py
@@ -56,6 +56,82 @@ def table_service(*, visual: bool = False) -> RetrievalService:
)
+class _OverviewRetriever:
+ """A fake with `find_by_drug`/`find_by_section` (the Qdrant adapter's
+ shape) — `InMemoryLexicalRetriever` doesn't implement either, so
+ `retrieve_framed`'s overview path is otherwise untestable."""
+
+ def __init__(self, documents: list[RetrievalDocument]) -> None:
+ self._documents = documents
+
+ def find_by_drug(self, drug_id: str) -> list[SearchHit]:
+ return [
+ SearchHit(document=d, score=1.0)
+ for d in self._documents if d.drug_id == drug_id
+ ]
+
+ def find_by_section(self, drug_id: str, section_key: str) -> list[SearchHit]:
+ return [
+ SearchHit(document=d, score=1.0)
+ for d in self._documents
+ if d.drug_id == drug_id and d.section_key == section_key
+ ]
+
+ def search(self, query: str, drug_id: str, limit: int) -> list[SearchHit]:
+ return []
+
+
+_MONOGRAPH_SECTIONS = (
+ "ten_chung_quoc_te", "ma_atc", "loai_thuoc", "dang_thuoc_va_ham_luong",
+ "duoc_ly_va_co_che_tac_dung", "chi_dinh", "chong_chi_dinh", "than_trong",
+ "tac_dung_khong_mong_muon", "lieu_luong_va_cach_dung", "tuong_tac_thuoc",
+ "qua_lieu_va_xu_tri", "do_on_dinh_va_bao_quan", "thong_tin_quy_che",
+)
+
+
+def _monograph_service() -> RetrievalService:
+ documents = [
+ RetrievalDocument(
+ doc_id=f"paracetamol::{section}::0", drug_id="paracetamol",
+ kind="prose", section_key=section,
+ text=f"Nội dung mục {section}.", source_refs=(SOURCE,),
+ )
+ for section in _MONOGRAPH_SECTIONS
+ ]
+ return RetrievalService(
+ _OverviewRetriever(documents), InMemoryParentStore([]),
+ EvidencePolicy(evidence_limit=3),
+ )
+
+
+def test_retrieve_framed_overview_answers_from_intro_sections_only():
+ """Found live 2026-08-06: a bare drug name sent all 14+ sections of the
+ monograph as evidence, producing an answer long enough to intermittently
+ fail generation/entailment. `is_overview=True` must narrow this the same
+ way `retrieve()`'s bare-name branch always has."""
+ result = _monograph_service().retrieve_framed(
+ "paracetamol", None, "paracetamol", is_overview=True
+ )
+ assert result.decision == EvidenceDecision.ANSWERABLE
+ assert result.is_drug_overview is True
+ returned_sections = {e.matched_doc_id.split("::")[1] for e in result.evidence}
+ assert returned_sections <= {
+ "ten_chung_quoc_te", "loai_thuoc", "chi_dinh", "duoc_ly_va_co_che_tac_dung",
+ }
+ assert len(result.evidence) < len(_MONOGRAPH_SECTIONS)
+
+
+def test_retrieve_framed_question_without_section_is_capped_even_without_rerank():
+ # No reranker configured: `_rerank` fails open and returns everything
+ # unfiltered. Hydration must still bound it — an ordering aid failing
+ # open must not also remove the size cap.
+ result = _monograph_service().retrieve_framed(
+ "paracetamol", None, "thuốc này có tác dụng phụ gì", is_overview=False
+ )
+ assert result.decision == EvidenceDecision.ANSWERABLE
+ assert len(result.evidence) <= 3
+
+
def test_row_hit_hydrates_complete_parent_and_keeps_citation():
result = table_service().retrieve("acetylcystein 45 kg bao nhiêu ml", "acetylcystein")
assert result.decision == EvidenceDecision.ANSWERABLE
diff --git a/apps/ai-service/tests/test_understanding.py b/apps/ai-service/tests/test_understanding.py
new file mode 100644
index 0000000..ca9d1e0
--- /dev/null
+++ b/apps/ai-service/tests/test_understanding.py
@@ -0,0 +1,202 @@
+"""`rag/understanding.py::LlmQueryUnderstander` — zero coverage before this
+(Codex's 2026-08-06 review, F-03/F-10), despite being the entry point for
+every live turn once F-03 wired it in.
+
+Also covers F-04 (bounded candidates): the resolver decides which drug_ids
+are even plausible for a turn *before* the model runs, and the model's pick
+is validated against that bound, not the full catalog.
+"""
+from __future__ import annotations
+
+import json
+
+from rag.understanding import SECTION_KEY_HINTS, SECTION_KEYS, LlmQueryUnderstander
+
+CATALOG = {
+ "paracetamol_acetaminophen": "paracetamol acetaminophen, PARACETAMOL",
+ "metformin": "metformin, METFORMIN",
+}
+
+
+class _FixedLlm:
+ def __init__(self, payload) -> None:
+ self._payload = payload
+
+ def generate(self, system: str, user: str, schema: dict) -> str:
+ if isinstance(self._payload, str):
+ return self._payload
+ return json.dumps(self._payload, ensure_ascii=False)
+
+
+class _Resolution:
+ def __init__(self, status="not_found", drug_id=None, candidate_drug_ids=()) -> None:
+ self.status = status
+ self.drug_id = drug_id
+ self.candidate_drug_ids = candidate_drug_ids
+
+
+class _FakeResolver:
+ """Deterministic stand-in for `CatalogDrugResolver`: resolves a line to
+ a drug_id if one of `known`'s substrings appears in it (case-insensitive),
+ with no fuzzy suggestions unless `suggestions` is given."""
+
+ def __init__(self, known: dict[str, str], suggestions: dict[str, str] | None = None) -> None:
+ self._known = known
+ self._suggestions = suggestions or {}
+
+ def resolve(self, query: str) -> _Resolution:
+ low = query.lower()
+ for needle, drug_id in self._known.items():
+ if needle in low:
+ return _Resolution(status="resolved", drug_id=drug_id)
+ return _Resolution()
+
+ def suggest(self, query: str, k: int = 3, min_score: float = 0.5):
+ low = query.lower()
+ return [
+ (drug_id, 0.9) for needle, drug_id in self._suggestions.items() if needle in low
+ ][:k]
+
+
+RESOLVER = _FakeResolver({
+ "metformin": "metformin",
+ "paracetamol": "paracetamol_acetaminophen",
+})
+
+
+def test_drug_id_in_exact_underscore_form_resolves():
+ understander = LlmQueryUnderstander(_FixedLlm({
+ "turn_type": "drug_attribute", "drugs": ["metformin"],
+ "unknown_drugs": [], "attribute": None, "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ }), CATALOG, RESOLVER)
+ frame = understander.understand("liều metformin")
+ assert frame.drugs == ("metformin",)
+ assert frame.unknown_drugs == ()
+
+
+def test_drug_id_echoed_with_spaces_instead_of_underscores_still_resolves():
+ """Reproduces the live 2026-08-06 bug on a genuine multi-turn shape: the
+ drug is named in an earlier turn (in history), the current turn is just
+ "30 cân", and the model echoed the spaced display name instead of the
+ underscored id. The old strict-equality check demoted a correctly
+ identified drug to unknown_drugs — producing "Không tìm thấy
+ paracetamol trong Dược thư" for a drug that plainly is in it."""
+ understander = LlmQueryUnderstander(_FixedLlm({
+ "turn_type": "dosing_calc", "drugs": ["paracetamol acetaminophen"],
+ "unknown_drugs": [], "attribute": None, "population": "tre_em",
+ "weight_kg": 30, "age_text": "7 tuổi", "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ }), CATALOG, RESOLVER)
+ frame = understander.understand(
+ "30 cân",
+ history=("Người dùng: Liều paracetamol cho trẻ em", "Trợ lý: Bé mấy tuổi?"),
+ )
+ assert frame.drugs == ("paracetamol_acetaminophen",)
+ assert frame.unknown_drugs == ()
+
+
+def test_a_genuinely_invented_name_is_unknown_not_substituted():
+ understander = LlmQueryUnderstander(_FixedLlm({
+ "turn_type": "drug_attribute", "drugs": [],
+ "unknown_drugs": ["aspirinol"], "attribute": None, "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ }), CATALOG, RESOLVER)
+ frame = understander.understand("liều aspirinol")
+ assert frame.drugs == ()
+ assert frame.unknown_drugs == ("aspirinol",)
+
+
+def test_a_name_with_no_deterministic_candidate_is_unknown_even_if_the_model_names_a_real_id():
+ """F-04's actual guarantee: the model naming a *real* catalog id is not
+ enough — that id must also be among the turn's deterministic candidates.
+ Nothing in "liều aspirinol" fuzzy/exact-matches any real drug (per
+ RESOLVER), so even if the model output a real id here, it must be
+ rejected: the candidate bound, not just catalog membership, is what's
+ trusted."""
+ understander = LlmQueryUnderstander(_FixedLlm({
+ "turn_type": "drug_attribute", "drugs": ["metformin"],
+ "unknown_drugs": [], "attribute": None, "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ }), CATALOG, RESOLVER)
+ frame = understander.understand("liều aspirinol")
+ assert frame.drugs == ()
+ assert "metformin" in frame.unknown_drugs
+
+
+def test_fuzzy_suggestion_bounds_a_typo_into_the_candidate_set():
+ resolver = _FakeResolver({}, suggestions={"metfomin": "metformin"})
+ understander = LlmQueryUnderstander(_FixedLlm({
+ "turn_type": "drug_attribute", "drugs": ["metformin"],
+ "unknown_drugs": [], "attribute": None, "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ }), CATALOG, resolver)
+ frame = understander.understand("liều metfomin")
+ assert frame.drugs == ("metformin",)
+
+
+def test_malformed_json_fails_closed_to_a_clarify():
+ understander = LlmQueryUnderstander(_FixedLlm("not json"), CATALOG, RESOLVER)
+ frame = understander.understand("gì đó")
+ assert frame.turn_type == "out_of_scope"
+ assert frame.needs_clarify is True
+
+
+def test_unrecognised_turn_type_falls_back_based_on_whether_a_drug_resolved():
+ understander = LlmQueryUnderstander(_FixedLlm({
+ "turn_type": "not_a_real_type", "drugs": ["metformin"],
+ "unknown_drugs": [], "attribute": None, "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ }), CATALOG, RESOLVER)
+ frame = understander.understand("metformin")
+ assert frame.turn_type == "drug_attribute"
+
+
+def test_every_section_key_has_a_hint():
+ # A key with no gloss shown to the model is exactly the bug this fixed —
+ # never let a new SECTION_KEYS entry silently ship without one.
+ assert set(SECTION_KEYS) == set(SECTION_KEY_HINTS)
+
+
+def test_prompt_disambiguates_than_trong_from_chong_chi_dinh():
+ """Reproduces the live 2026-08-06 golden-eval finding: a bare section
+ key list gave the model nothing to tell "thận trọng" (precautions) apart
+ from "chống chỉ định" (contraindications) — 9/9 live calls for "X cần
+ thận trọng gì?" picked chong_chi_dinh, silently answering from the wrong
+ section and dropping safety content the precautions section actually
+ has (metformin's lactic-acidosis warning, gentamicin's oto/nephrotoxicity).
+ Fixed with an inline gloss; this pins the gloss's presence in the actual
+ request sent, not just its existence in the hints dict."""
+ captured = {}
+
+ class _CapturingLlm:
+ def generate(self, system, user, schema):
+ captured["user"] = user
+ return json.dumps({
+ "turn_type": "drug_attribute", "drugs": ["metformin"],
+ "unknown_drugs": [], "attribute": "than_trong", "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ })
+
+ understander = LlmQueryUnderstander(_CapturingLlm(), CATALOG, RESOLVER)
+ understander.understand("metformin cần thận trọng gì?")
+
+ assert "KHÁC chống chỉ định" in captured["user"]
+ assert "nhiễm toan lactic" in captured["user"]
+
+
+def test_invalid_attribute_is_dropped_not_passed_through():
+ understander = LlmQueryUnderstander(_FixedLlm({
+ "turn_type": "drug_attribute", "drugs": ["metformin"],
+ "unknown_drugs": [], "attribute": "not_a_real_section", "population": None,
+ "weight_kg": None, "age_text": None, "indication": None,
+ "needs_clarify": False, "clarify_reason": None,
+ }), CATALOG, RESOLVER)
+ frame = understander.understand("metformin")
+ assert frame.attribute is None
diff --git a/apps/web/app/_components/ChatPanel.tsx b/apps/web/app/_components/ChatPanel.tsx
index 5c8e016..2709172 100644
--- a/apps/web/app/_components/ChatPanel.tsx
+++ b/apps/web/app/_components/ChatPanel.tsx
@@ -1,121 +1,343 @@
"use client";
-import { useState } from "react";
-import { Pill, Send } from "lucide-react";
-import type { ChatMessage, Citation } from "@duoc-thu/shared-types";
-import { ChatBubble, CitationCard, Card, Input, Button, cn } from "@duoc-thu/ui";
-import { sendChatMessage } from "@duoc-thu/api-client";
+import React, { useState, useEffect, useRef } from "react";
+import type { ChatMessage, Citation, SendMessageResponse } from "@duoc-thu/shared-types";
+import { ChatBubble, CitationBeamOverlay, useTheme } from "@duoc-thu/ui";
+import { Composer } from "./Composer";
+import {
+ Sparkles,
+ Pill,
+ ShieldCheck,
+ BookOpen,
+ Activity,
+ Zap,
+ Info,
+ AlertCircle,
+ Stethoscope,
+} from "lucide-react";
+import { cn } from "@duoc-thu/ui";
-function TypingIndicator() {
- return (
-
-
-
-
-
- );
-}
-
-export interface ChatPanelProps {
- onCitationClick?: (citation: Citation) => void;
+interface ChatPanelProps {
+ sessionId: string;
+ initialQuery?: string;
+ onCitationClick?: (citation: Citation, index: number) => void;
+ onCitationsLoaded?: (citations: Citation[]) => void;
+ activeCitationIndex?: number | null;
className?: string;
}
-const ERROR_MESSAGE =
- "Hệ thống tạm thời không phản hồi. Vui lòng thử lại — nếu vẫn lỗi, có thể dịch vụ tra cứu đang tạm ngưng.";
+const STARTER_QUESTIONS = [
+ {
+ category: "Liều Dùng Lâm Sàng",
+ query: "Liều dùng Paracetamol người lớn và trẻ em theo cân nặng là bao nhiêu?",
+ icon: Pill,
+ },
+ {
+ category: "Chống Chỉ Định",
+ query: "Chống chỉ định tuyệt đối và tương đối của Amoxicillin là gì?",
+ icon: Stethoscope,
+ },
+ {
+ category: "Tương Tác Thuốc",
+ query: "Tương tác giữa Metformin và thuốc cản quang chứa iốt xử trí thế nào?",
+ icon: Activity,
+ },
+ {
+ category: "Thận Trọng & ADR",
+ query: "Thận trọng khi dùng Aspirin cho bệnh nhân có tiền sử loét dạ dày?",
+ icon: Zap,
+ },
+];
-export function ChatPanel({ onCitationClick, className }: ChatPanelProps) {
+export function ChatPanel({
+ sessionId,
+ initialQuery,
+ onCitationClick,
+ onCitationsLoaded,
+ activeCitationIndex = null,
+ className,
+}: ChatPanelProps) {
+ const { resolvedTheme } = useTheme();
const [messages, setMessages] = useState([]);
- const [input, setInput] = useState("");
- const [isSending, setIsSending] = useState(false);
- // One id per chat session, so follow-ups ("còn trẻ em thì sao?") resolve
- // against the same conversation on the backend.
- const [conversationId] = useState(() =>
- typeof crypto !== "undefined" && crypto.randomUUID
- ? crypto.randomUUID()
- : `conv-${Date.now()}`
- );
+ const [isLoading, setIsLoading] = useState(false);
+ const [error, setError] = useState(null);
+ const messagesEndRef = useRef(null);
+ const abortControllerRef = useRef(null);
- async function handleSubmit(event: React.FormEvent) {
- event.preventDefault();
- const content = input.trim();
- if (!content || isSending) return;
+ const scrollToBottom = () => {
+ messagesEndRef.current?.scrollIntoView({ behavior: "smooth" });
+ };
- const userMessage: ChatMessage = {
- id: `local-${messages.length}`,
+ useEffect(() => {
+ scrollToBottom();
+ }, [messages, isLoading]);
+
+ const handleSendMessage = async (userText: string) => {
+ if (!userText.trim() || isLoading) return;
+
+ setError(null);
+
+ const userMsg: ChatMessage = {
+ id: `user-${Date.now()}`,
role: "user",
- content,
+ content: userText,
createdAt: new Date().toISOString(),
};
- setMessages((prev) => [...prev, userMessage]);
- setInput("");
- setIsSending(true);
+
+ setMessages((prev) => [...prev, userMsg]);
+ setIsLoading(true);
+
+ abortControllerRef.current = new AbortController();
+
try {
- const response = await sendChatMessage(content, conversationId);
- setMessages((prev) => [...prev, response.message]);
- } catch {
- // Never leave the user staring at their own message with no reply: an
- // error is surfaced as a labelled bubble, not swallowed silently.
- setMessages((prev) => [
- ...prev,
- {
- id: `error-${messages.length}`,
- role: "assistant",
- content: ERROR_MESSAGE,
- createdAt: new Date().toISOString(),
- },
- ]);
+ const res = await fetch("/api/chat", {
+ method: "POST",
+ headers: { "Content-Type": "application/json" },
+ body: JSON.stringify({
+ content: userText,
+ conversationId: sessionId,
+ }),
+ signal: abortControllerRef.current.signal,
+ });
+
+ if (!res.ok) {
+ throw new Error(`Upstream returned status ${res.status}`);
+ }
+
+ const data: SendMessageResponse = await res.json();
+ const assistantMsg = data.message;
+
+ setMessages((prev) => [...prev, assistantMsg]);
+
+ if (assistantMsg.citations && assistantMsg.citations.length > 0) {
+ onCitationsLoaded?.(assistantMsg.citations);
+ }
+ } catch (err: any) {
+ if (err.name === "AbortError") {
+ return;
+ }
+ setError("Không thể kết nối đến máy chủ AI Service. Vui lòng kiểm tra lại dịch vụ backend.");
} finally {
- setIsSending(false);
+ setIsLoading(false);
+ abortControllerRef.current = null;
}
- }
+ };
+
+ const handleStop = () => {
+ if (abortControllerRef.current) {
+ abortControllerRef.current.abort();
+ setIsLoading(false);
+ abortControllerRef.current = null;
+ }
+ };
+
+ useEffect(() => {
+ if (initialQuery) {
+ handleSendMessage(initialQuery);
+ }
+ // eslint-disable-next-line react-hooks/exhaustive-deps
+ }, [initialQuery]);
+
+ // Empty state renderer per theme
+ const renderEmptyState = () => {
+ if (resolvedTheme === "light") {
+ return (
+
+ Hệ thống AI y tế tra cứu chính xác theo 684 chuyên luận chính thức. Mọi thông tin đều được xác thực suy luận (Entailment Verification) kèm trích dẫn trang in PDF.
+