Initial monorepo scaffold for Duoc Thu RAG medical chatbot

This commit is contained in:
2026-07-30 20:38:34 +07:00
commit 9bad1f61ea
91 changed files with 740 additions and 0 deletions
+45
View File
@@ -0,0 +1,45 @@
# ADR 0001: Use Qdrant as the vector database
## Status
Accepted
## Context
The RAG pipeline needs a vector store for drug-monograph chunks. The main
alternative considered was **pgvector** (a Postgres extension), which would
let us reuse the Postgres instance already needed for users/chat history —
one fewer moving part to operate.
The corpus is not free-flowing prose: it's a structured per-drug reference
with rich per-chunk metadata (drug name, section type, page range). The
common retrieval pattern this domain calls for is "vector similarity search,
filtered by metadata" — e.g. "search only within chỉ định sections" or
"filter to a specific drug the user named" combined with the semantic query.
## Decision
Use **Qdrant** as a dedicated vector database, separate from Postgres.
## Rationale
- Qdrant gives first-class combined payload-filtering + ANN search in a
single query, which is exactly the retrieval pattern this structured
corpus needs — pgvector supports filtering too, but it's a less natural
fit layered on top of a general-purpose relational engine.
- Vector search becomes its own independent scaling axis, separate from the
transactional Postgres workload (users/chat) — re-indexing or re-ingesting
the formulary doesn't contend with transactional traffic.
- Mature standalone Docker image for local dev, a well-supported Python
client, and a Helm chart for the production Kubernetes deployment target.
- Corpus size (tens of thousands of chunks) is trivial for Qdrant's HNSW
indexing.
## Consequences
- One additional service to operate/deploy/monitor compared to pgvector
(which would ride on the existing Postgres).
- Revisit if operational overhead becomes a real burden at our actual scale,
or if we want tighter transactional consistency between chat data and
retrieval — pgvector remains a viable fallback documented here for that
case.
+48
View File
@@ -0,0 +1,48 @@
# ADR 0002: Use the team's existing ArgoCD instance for deployment (GitOps)
## Status
Accepted
## Context
Phase 6 of the build roadmap needs a way to actually deploy the Helm chart to
Kubernetes across dev/staging/prod. The original scaffold (`infra/ci/github-actions/deploy-cd.yml`)
assumed a push-based CI deploy step (CI runs `helm upgrade`/`kubectl apply`
directly against the cluster). The team already runs an ArgoCD instance used
by other projects.
## Decision
Deploy via **GitOps through the team's existing ArgoCD instance** instead of
building a custom push-based CD pipeline. ArgoCD Applications
(`infra/argocd/applications/{dev,staging,prod}/app.yaml`) point at
`infra/helm/medical-chatbot` in this repo; ArgoCD watches the repo and
reconciles the cluster to match.
## Rationale
- Reuses infrastructure the team already operates and trusts, instead of
standing up a parallel deploy mechanism.
- GitOps gives an auditable history of every deploy (it's just git commits
changing values files/image tags) and a built-in rollback path (revert the
commit).
- Removes the need for CI to hold cluster credentials — CI's job shrinks to
"build, test, push image, bump tag," which is a smaller security surface
than "CI can directly mutate the production cluster."
- Prod uses a non-automated `syncPolicy` (manual approval in ArgoCD) while
dev/staging auto-sync, matching normal caution around production changes.
## Consequences
- CI workflows (`infra/ci/github-actions/*.yml`) build/test/push images and
bump the relevant `values-<env>.yaml` image tag + push that commit; they do
**not** call `kubectl`/`helm` against any cluster directly.
- Actual deploy execution and health/sync status live in the team's ArgoCD
UI/CLI, outside this repo — runbooks in `docs/runbooks/` should document how
to check sync status and roll back once the team's ArgoCD instance details
(cluster/server, project, repo URL) are confirmed (see TODOs in
`infra/argocd/README.md`).
- If the team's ArgoCD instance becomes unavailable or this project needs to
fully own its own deploy tooling later, the push-based `deploy-cd.yml`
approach remains a documented fallback.
+121
View File
@@ -0,0 +1,121 @@
# Architecture — Dược Thư RAG Medical Chatbot
## Overview
A medical chatbot grounded in the Vietnamese National Drug Formulary (Dược
thư quốc gia Việt Nam 2018), built as a microservices monorepo. Users ask
drug-related questions through a web chat UI; answers are generated via
retrieval-augmented generation (RAG) over the formulary content, always
citing the source drug monograph/section, and always carrying a medical
disclaimer.
## Service responsibilities & communication
| Service | Owns | Talks to |
|---|---|---|
| **api-gateway** (NestJS) | Single public entry point; request routing, JWT validation, rate limiting | Routes to auth-service, user-service, chat-service, ai-service over internal REST |
| **auth-service** (NestJS) | Signup/login, password hashing, JWT issuance/refresh | Postgres (users); no dependency on other services |
| **user-service** (NestJS) | Profile data, preferences, account settings | Postgres (profiles), called by gateway |
| **chat-service** (NestJS) | Chat session lifecycle, message history persistence | Postgres (chat_sessions, chat_messages); calls ai-service per user message, persists both turns |
| **ai-service** (Python/FastAPI) | RAG orchestration: embed query → vector search in Qdrant → build grounded prompt → call OpenAI → return answer + citations | Qdrant (vector search), OpenAI API; stateless, does not own chat history |
| **ingestion** (Python, offline batch) | One-time/periodic job: parse PDF → monographs → chunks → embeddings → upsert to Qdrant | Qdrant (write), OpenAI embeddings API; runs as CLI/CI/k8s Job, never in the live request path |
| **web** (Next.js) | Chat UI, auth UI, citation/disclaimer rendering, session list | Calls api-gateway only |
**Sync vs async**: the live chat path (web → gateway → chat-service →
ai-service → Qdrant + OpenAI → back) is synchronous request/response.
Ingestion is fully decoupled, offline, batch — it populates Qdrant ahead of
time and is never triggered by a chat request, since parsing the 37MB PDF and
embedding thousands of chunks takes minutes. Internal protocol is REST/JSON
for v1; a future gRPC migration is a documented option (see ADRs), not
needed now.
## Data stores
- **Vector DB: Qdrant.** Chosen over pgvector because retrieval quality here
depends on metadata-filtered ANN search (filter by drug name / section type
combined with vector similarity) over a highly structured corpus — Qdrant
makes that a first-class, single query. It also scales independently from
the transactional Postgres and has a mature Helm chart for the production
k8s target. See `docs/adr/0001-vector-db-qdrant.md`.
- **Relational DB: PostgreSQL.** One instance, logically separated per
service (users/credentials, profiles, chat sessions+messages).
- **Redis.** Session/refresh-token cache, rate-limit counters, and reserved
as the future job-queue backend (BullMQ/Celery) if async admin-triggered
re-ingestion or background jobs are added later.
## RAG ingestion pipeline (PDF-specific)
The formulary is a structured per-drug reference, not free prose — the
pipeline exploits that structure instead of naive fixed-size chunking:
1. **Extraction**: PyMuPDF (`fitz`) as primary extractor (font size/style/
position metadata enables heading detection); pdfplumber as a fallback
specifically for tabular content (dosing/interaction tables). Raw
per-page extraction is persisted to `ingestion/data/interim/` so
re-segmentation doesn't require re-running the expensive extraction step.
2. **Segmentation**: detect drug-entry boundaries (prefer the PDF's
bookmark/outline via `doc.get_toc()` when present, else font-size/style
heuristics), then classify each heading against a canonical section
taxonomy (`chi_dinh`, `chong_chi_dinh`, `lieu_dung`, `tac_dung_phu`,
`tuong_tac_thuoc`, etc., Vietnamese diacritic-insensitive matching).
Output: `{drug_id, drug_name, source_page_range, sections: {...}}` per
drug, persisted to `ingestion/data/processed/monographs.jsonl` and
manually spot-checked via `ingestion/notebooks/`.
3. **Chunking**: each `(drug, section)` pair is the natural chunk unit;
never split a section unless it exceeds a token budget (~500-800 tokens),
in which case sub-chunk with a sliding window (400 tokens, 50 overlap),
tagging the same drug+section metadata plus `part_index`. Every chunk
carries `drug_name`, `section_type`, `source_page_range`, `chunk_id` as
Qdrant payload — this is what makes citations possible.
4. **Embedding + load**: OpenAI `text-embedding-3-small` in batches, upserted
into a versioned Qdrant collection (`drug_monographs_v1`) keyed by
`chunk_id` for idempotent re-runs; collection aliasing allows re-ingesting
with a changed chunking strategy without downtime.
5. **Batch job, not synchronous**: runs as a CLI command locally, and as a
Kubernetes `Job`/`CronJob` in production — never inside the ai-service
request path.
## Safety / guardrails
- **System prompt** instructs the model to answer only from retrieved
context, never state a dosage/contraindication/interaction not present in
it, always append a disclaimer, and say "not found in the formulary"
rather than guess when retrieval is irrelevant.
- **Retrieval-confidence gate**: below a similarity threshold, skip the LLM
call entirely and return a canned "consult a professional" response.
- **Citations from metadata, not LLM prose**: the `citations` list is built
directly from retrieved-chunk metadata, independent of what the LLM says,
so the frontend can always show verifiable sources.
- **Disclaimer enforced at multiple layers**: system prompt + a
non-LLM-generated static string always appended to the API response + a
persistent, non-dismissible UI banner.
- **Scoped refusal**: out-of-scope questions (e.g. general symptom
diagnosis) get a scoped refusal directing to a professional, not an
ungrounded general-knowledge answer.
## Build roadmap
1. **Ingestion pipeline + populated, queryable vector DB.** Done when a CLI
run populates Qdrant and a test script retrieves the correct
drug/section chunk for a sample query — no API, no LLM call yet.
2. **ai-service (FastAPI) wrapping RAG + OpenAI.** Done when a `curl` to
`/query` returns a grounded answer with a traceable citation and an
always-present disclaimer.
3. **auth/user/chat services + api-gateway.** Done when register → login →
chat message flows end-to-end through the gateway only, persisted in
Postgres.
4. **Next.js frontend chat UI.** Done when a browser user can log in, ask a
question, and see a grounded answer with citation + disclaimer banner.
5. **Containerize + docker-compose local.** Done when `docker compose up`
from a clean checkout brings up the full stack and the Phase 4 flow works.
6. **Kubernetes/Helm + Terraform + CI + ArgoCD (GitOps) deployment.** Done
when CI builds/tests/pushes an image and bumps the target environment's
Helm values file, the team's ArgoCD instance (see `infra/argocd/`,
`docs/adr/0002-argocd-gitops.md`) picks up the change and syncs the
cluster, and the Phase 4 flow works against the k8s-hosted stack. CI
never runs `kubectl`/`helm` directly against a cluster. Cloud provider
choice (AWS/GCP/Azure) only affects the Terraform module implementations,
not this repo's structure.
See `docs/adr/` for architecture decision records and `docs/runbooks/` for
operational runbooks (added as they're needed).
+43
View File
@@ -0,0 +1,43 @@
# Progress Log
Chronological record of work done on this project, newest entry on top. The
goal is continuity across sessions: if a work session ends unexpectedly
(context/token limit, interruption), whoever picks this up next — human or
Claude — should be able to read the latest entry and know exactly what's
done and what's next, without having to reconstruct it from git history.
**Convention**: add a new entry at the top before ending a session whenever
meaningful progress was made, and proactively the moment it looks like the
session might run out of context/tokens mid-task — don't wait until the very
end if that risk is showing.
---
## 2026-07-30 — Initial monorepo scaffold
**Done:**
- Designed the microservices architecture (see `docs/architecture.md`):
Python/FastAPI `ai-service` for RAG, NestJS for `api-gateway`/`auth-service`/
`user-service`/`chat-service`, Next.js `web`, Qdrant for vectors, Postgres
for relational data, Redis reserved for caching/queues.
- Scaffolded the full monorepo directory tree (`apps/`, `packages/`,
`ingestion/`, `infra/`, `docs/`) with baseline config (package.json/
pyproject.toml stubs, pnpm workspace, docker-compose topology stub).
- Moved `duoc-thu-quoc-gia-viet-nam-2018.pdf` into `ingestion/data/raw/`.
- Decided vector DB: **Qdrant** over pgvector (`docs/adr/0001-vector-db-qdrant.md`).
- Decided deployment: GitOps via the **team's existing ArgoCD instance**,
not a custom push-based CD pipeline (`docs/adr/0002-argocd-gitops.md`,
`infra/argocd/`). CI's job is build/test/push image + bump the Helm values
image tag; ArgoCD does the actual sync.
- `git init` + initial commit (this scaffold).
- Created a private GitHub repo and pushed the initial commit.
**Not done yet / next up (Phase 1 of the build roadmap in `docs/architecture.md`):**
- No business logic exists yet anywhere — this was scaffold only.
- Phase 1: build the `ingestion/` pipeline for real (PDF extraction via
PyMuPDF, monograph/section segmentation, section-aware chunking, OpenAI
embeddings, Qdrant upsert) and validate retrieval quality via the
`ingestion/notebooks/` QA step.
- Still pending/TBD: which cloud provider (AWS/GCP/Azure) for Terraform
(`infra/terraform/README.md`), and the team's ArgoCD instance's actual
cluster/server + project details (`infra/argocd/README.md` TODOs).
View File