# Architecture — Dược Thư RAG Medical Chatbot ## Overview A medical chatbot grounded in the Vietnamese National Drug Formulary (Dược thư quốc gia Việt Nam 2018), built as a microservices monorepo. Users ask drug-related questions through a web chat UI; answers are generated via retrieval-augmented generation (RAG) over the formulary content, always citing the source drug monograph/section, and always carrying a medical disclaimer. ## Service responsibilities & communication | Service | Owns | Talks to | |---|---|---| | **api-gateway** (NestJS) | Single public entry point; request routing, JWT validation, rate limiting | Routes to auth-service, user-service, chat-service, ai-service over internal REST | | **auth-service** (NestJS) | Signup/login, password hashing, JWT issuance/refresh | Postgres (users); no dependency on other services | | **user-service** (NestJS) | Profile data, preferences, account settings | Postgres (profiles), called by gateway | | **chat-service** (NestJS) | Chat session lifecycle, message history persistence | Postgres (chat_sessions, chat_messages); calls ai-service per user message, persists both turns | | **ai-service** (Python/FastAPI) | RAG orchestration: embed query → vector search in Qdrant → build grounded prompt → call OpenAI → return answer + citations | Qdrant (vector search), OpenAI API; stateless, does not own chat history | | **ingestion** (Python, offline batch) | One-time/periodic job: parse PDF → monographs → chunks → embeddings → upsert to Qdrant | Qdrant (write), OpenAI embeddings API; runs as CLI/CI/k8s Job, never in the live request path | | **web** (Next.js) | Chat UI, auth UI, citation/disclaimer rendering, session list | Calls api-gateway only | **Sync vs async**: the live chat path (web → gateway → chat-service → ai-service → Qdrant + OpenAI → back) is synchronous request/response. Ingestion is fully decoupled, offline, batch — it populates Qdrant ahead of time and is never triggered by a chat request, since parsing the 37MB PDF and embedding thousands of chunks takes minutes. Internal protocol is REST/JSON for v1; a future gRPC migration is a documented option (see ADRs), not needed now. ## Data stores - **Vector DB: Qdrant.** Chosen over pgvector because retrieval quality here depends on metadata-filtered ANN search (filter by drug name / section type combined with vector similarity) over a highly structured corpus — Qdrant makes that a first-class, single query. It also scales independently from the transactional Postgres and has a mature Helm chart for the production k8s target. See `docs/adr/0001-vector-db-qdrant.md`. - **Relational DB: PostgreSQL.** One instance, logically separated per service (users/credentials, profiles, chat sessions+messages). - **Redis.** Session/refresh-token cache, rate-limit counters, and reserved as the future job-queue backend (BullMQ/Celery) if async admin-triggered re-ingestion or background jobs are added later. ## RAG ingestion pipeline (PDF-specific) The formulary is a structured per-drug reference, not free prose — the pipeline exploits that structure instead of naive fixed-size chunking. This section reflects an actual empirical investigation of the real PDF (not assumptions) — see `docs/adr/0003-pdf-parsing-strategy.md` for the full methodology, cross-tool comparison, and validation numbers. 1. **Extraction**: PyMuPDF (`fitz`) as primary extractor. This document has **no bookmark/outline** (`doc.get_toc()` returns 0 entries — confirmed, do not rely on it) and is a **tagged PDF with only a shallow, unusable structure tree** (~29 generic H1/P elements covering a fraction of 1668 pages — also confirmed dead-end, not a data source). PyMuPDF's reading order was cross-validated against `pdfplumber` and `opendataloader-pdf` on real sample pages: pdfplumber's default text order is **unreliable** for this layout (scrambles paragraph order, leaks marked-content artifacts) — use it only for its dedicated table-extraction API, never for body text. Raw per-page extraction is persisted to `ingestion/data/interim/` so re-segmentation doesn't require re-running the expensive extraction step. 2. **Segmentation**: drug-entry boundaries are detected via **bold-font spans** (PyMuPDF span `font` containing `"Bold"`), not font-size alone — font size for title/heading spans varies between monographs (confirmed: 10.0pt and 9.5pt both occur for genuine drug-title headings), so bold is the reliable signal, all-caps + short length narrows it to monograph titles specifically. Section headings inside a monograph are also bold spans, cross-checked against a canonical taxonomy (`chi_dinh`, `chong_chi_dinh`, `lieu_dung`, `tac_dung_phu`, `tuong_tac_thuoc`, plus real observed extras like `ten_thuong_mai` "Tên thương mại" not in the book's own documented 19-field list — treat the taxonomy as open/ extensible, not a fixed enum). Multi-line wrapped titles/headings (long Vietnamese names/vaccine names) must be merged across consecutive bold+all-caps lines before matching — this was the single largest source of missed detections in validation. Output: `{drug_id, drug_name, source_page_range, sections: {...}}` per drug, persisted to `ingestion/data/processed/monographs.jsonl` and validated both automatically (see ADR 0003) and via manual spot-check in `ingestion/notebooks/`. 3. **Chunking** (monograph range only, pp. 99-1496 — see `docs/adr/0004-chunking-strategy.md` for the full measured rationale): each `(drug_id, section_key)` pair is the chunk unit; a section stays one chunk if it's under an **800-token ceiling** (chars/4 estimate — a validated line, not a guess: whole-corpus measurement across 682 monographs shows ~16 of 18 section types clear it comfortably at their p90). Two sections routinely exceed it — `dược lý và cơ chế tác dụng` (35.7% of monographs that have it) and `liều lượng và cách dùng` (29.6%) — sub-chunking is the **routine** path for those two, not a rare edge case. Oversized sections are split with a **sentence-boundary-aware sliding window** (~600-700 tokens/sub-chunk, ~1 sentence/50-80 token overlap), never a blind character/line window — PDF line-wrap points are not safe cut points, and a mid-sentence split risks separating an adult/child dosing instruction (a measured, common pattern — outlier catalog item 17) into two chunks. Every chunk carries `chunk_id`, `drug_id`, `drug_name`, `section_key`, `section_display_name`, `atc_codes`, `source_page_range`, `part_index`/`part_count` as Qdrant payload — this is what makes citations possible. **Known open gaps** (see ADR 0004): sub-compound tagging inside class-level/multi-ATC monographs (25.5% of the corpus) is not yet solved; `source_page_range` is monograph-level, not sub-chunk-exact; chunking for general chapters/ appendices is a separate, not-yet-designed task; a confirmed header/footer-boilerplate leak into section text (98.4% of monographs affected) must be fixed upstream before this design runs against real data. 4. **Embedding + load**: OpenAI `text-embedding-3-small` in batches, upserted into a versioned Qdrant collection (`drug_monographs_v1`) keyed by `chunk_id` for idempotent re-runs; collection aliasing allows re-ingesting with a changed chunking strategy without downtime. 5. **Batch job, not synchronous**: runs as a CLI command locally, and as a Kubernetes `Job`/`CronJob` in production — never inside the ai-service request path. ## Safety / guardrails - **System prompt** instructs the model to answer only from retrieved context, never state a dosage/contraindication/interaction not present in it, always append a disclaimer, and say "not found in the formulary" rather than guess when retrieval is irrelevant. - **Retrieval-confidence gate**: below a similarity threshold, skip the LLM call entirely and return a canned "consult a professional" response. - **Citations from metadata, not LLM prose**: the `citations` list is built directly from retrieved-chunk metadata, independent of what the LLM says, so the frontend can always show verifiable sources. - **Disclaimer enforced at multiple layers**: system prompt + a non-LLM-generated static string always appended to the API response + a persistent, non-dismissible UI banner. - **Scoped refusal**: out-of-scope questions (e.g. general symptom diagnosis) get a scoped refusal directing to a professional, not an ungrounded general-knowledge answer. ## Build roadmap 1. **Ingestion pipeline + populated, queryable vector DB.** Done when a CLI run populates Qdrant and a test script retrieves the correct drug/section chunk for a sample query — no API, no LLM call yet. 2. **ai-service (FastAPI) wrapping RAG + OpenAI.** Done when a `curl` to `/query` returns a grounded answer with a traceable citation and an always-present disclaimer. 3. **auth/user/chat services + api-gateway.** Done when register → login → chat message flows end-to-end through the gateway only, persisted in Postgres. 4. **Next.js frontend chat UI.** Done when a browser user can log in, ask a question, and see a grounded answer with citation + disclaimer banner. 5. **Containerize + docker-compose local.** Done when `docker compose up` from a clean checkout brings up the full stack and the Phase 4 flow works. 6. **Kubernetes/Helm + Terraform + CI + ArgoCD (GitOps) deployment.** Done when CI builds/tests/pushes an image and bumps the target environment's Helm values file, the team's ArgoCD instance (see `infra/argocd/`, `docs/adr/0002-argocd-gitops.md`) picks up the change and syncs the cluster, and the Phase 4 flow works against the k8s-hosted stack. CI never runs `kubectl`/`helm` directly against a cluster. Cloud provider choice (AWS/GCP/Azure) only affects the Terraform module implementations, not this repo's structure. See `docs/adr/` for architecture decision records and `docs/runbooks/` for operational runbooks (added as they're needed).