Dược Thư RAG — Medical Chatbot Platform

Medical chatbot grounded in the Vietnamese National Drug Formulary (Dược thư quốc gia Việt Nam 2018), built as a microservices monorepo.

See docs/architecture.md for the full design (service responsibilities, data stores, RAG ingestion strategy, safety guardrails), docs/adr for architecture decision records, docs/pdf-parsing-outlier-catalog.md for a reusable checklist of confirmed PDF-parsing risks (useful for this book and any similarly-structured PDF), and docs/progress-log.md for a running log of what's been done and what's next.

Dated planning and audit documents (docs/v1-delivery-plan.md, docs/rag-rebuild-plan.md, docs/current-rag-pipeline-audit.md, docs/answer-experience-implementation-plan.md) record what was known on their date and are kept for their reasoning rather than as current status — this README and git log are the better reference for where things stand today.

Status (2026-08-11): live in production at realvuxbaro.me — a real RAG chatbot over the whole formulary, not a scaffold. What exists and what does not:

Part State
ingestion/ Done — 15,100 chunks embedded and loaded into Qdrant duocthu_v1
apps/ai-service/ Done — live grounded RAG (retrieval, generation, grounding, abstention, citations, traces)
apps/web/ Done — chat UI with citation/evidence panel
apps/api-gateway, auth-service, user-service, chat-service Not builtREADME.md + package.json only
apps/mobile/ Not built — reserved
infra/docker/ Done — production runs Compose, including the Prometheus/Grafana/Tempo observability overlay
infra/helm/medical-chatbot/ Built and validated as an offline migration kit; not deployed to Docker Desktop, k3s or ArgoCD
infra/k8s, terraform, argocd Not built yet — still the target (ADR 0002), not abandoned: the plan is the team's self-hosted Gitea + ArgoCD; the current EC2/Compose setup is an interim stopgap

Because the gateway and auth services do not exist, apps/web talks directly to apps/ai-service; there is no authentication layer. See the build roadmap in docs/architecture.md.

Directory map

apps/
  web/            Next.js frontend (also hosts the BFF route the browser calls)
  ai-service/     Python FastAPI — RAG orchestration + AWS Bedrock calls
  api-gateway/    NestJS — public entry point, routes to internal services
  auth-service/   NestJS — signup/login/JWT
  user-service/   NestJS — profile/preferences
  chat-service/   NestJS — chat session + message history
  mobile/         reserved for a future mobile app
packages/
  shared-types/   TS DTOs shared across Node services + web
  api-client/     typed HTTP client for web
  ui/             shared React components
  config/         shared eslint/tsconfig presets
ingestion/        offline batch pipeline: PDF -> monographs -> chunks -> embeddings -> Qdrant
infra/            docker-compose, k8s/Helm, Terraform, CI
docs/             architecture docs and ADRs

Prerequisites

  • Node.js + pnpm (JS workspace: apps/web, packages/*; the NestJS service directories are unbuilt placeholders)
  • Python 3.11+ (apps/ai-service, ingestion)
  • Docker (local Postgres + Qdrant via infra/docker/docker-compose.yml)
  • AWS credentials with Bedrock invoke permission, for anything that generates an answer. Without them ai-service still starts, but every answer abstains rather than falling back to raw source text.

Running it locally

docker compose -f infra\docker\docker-compose.yml up -d postgres qdrant
cd apps\ai-service
python -m migrate
python -m uvicorn main:app --port 8079     # NOT --reload, see below
pnpm install
pnpm --filter web dev                      # http://localhost:3000

ai-service needs a populated Qdrant collection to serve answers: it verifies a duocthu_v1__manifest sidecar at startup and refuses to run against a corpus whose sha/model/dimensions do not match. A fresh machine either restores a Qdrant snapshot or re-runs ingestion/ (the latter costs real Bedrock spend).

Prefer a plain restart over uvicorn --reload on Windows here. The reloader has been observed serving the previous code after an edit on this project, which makes it hard to tell whether a change took effect.

Tests: cd apps/ai-service && python -m pytest -q — 230 pass. test_api.py and test_live_datastores.py need Postgres and Qdrant actually running; skip them with --ignore when the stack is down. apps/web has no test setup at all, so a green suite says nothing about the frontend — drive it in a browser.

Observability: Prometheus, Grafana and Tempo

The observability stack is provisioned in the repository and has been deployed to the production EC2 instance since 2026-08-11.

  • Prometheus scrapes /metrics from ai-service. It records request rate and latency, latency for each RAG stage, routing decisions and reasons, provider failures, trace-write failures and the existing domain counters.
  • Grafana is the user interface for dashboards and metric queries. Its datasource and the Dược Thư dashboard are provisioned automatically.
  • Tempo stores OpenTelemetry traces. A trace contains the receive, understanding, routing, retrieval, rerank/evidence, generation, grounding/entailment, persistence and response stages. Correlation and trace IDs follow the request from the Next.js BFF into FastAPI.
  • OpenTelemetry Collector receives spans from ai-service and exports them to Tempo. Grafana exemplars link aggregate latency metrics to an individual Tempo trace.

For answer lineage, use the three views together:

  1. The web citation/evidence panel shows which source chunks, pages and exact evidence text were selected for the answer.
  2. Grafana -> Explore -> Tempo shows which pipeline stages ran, their nesting and timing, the final decision/reason, provider failures and the persisted trace ID.
  3. PostgreSQL table rag_retrieval_trace is the durable audit record. It stores the query, resolved drug, decision/reason, selected citations/evidence, correlation ID and OpenTelemetry trace ID, so a returned trace_id can be joined to its Tempo trace.

This is provenance and execution tracing, not model chain-of-thought logging. Full prompts/responses, hidden reasoning, every rejected retrieval candidate and every ranking score are deliberately not stored today. If deeper debugging is needed, add bounded audit fields rather than putting sensitive prompt or patient content into metric labels or span names.

Start the local stack from the repository root:

docker compose -f infra\docker\docker-compose.yml up -d prometheus tempo otel-collector grafana

Local endpoints:

Service Address Use
Grafana http://localhost:3002 Dashboards and Explore
Prometheus http://localhost:9090 Raw targets, PromQL and metrics
Tempo http://localhost:3200 Trace backend; normally queried through Grafana
ai-service metrics http://localhost:8079/metrics Raw OpenMetrics output when ai-service runs on port 8079

For the existing EC2 Compose deployment, the optional overlay is infra/docker/docker-compose.observability.yml. It leaves docker-compose.prod.yml unchanged. A deployment, when explicitly approved, uses both files:

docker compose `
  -f infra/docker/docker-compose.prod.yml `
  -f infra/docker/docker-compose.observability.yml `
  up -d

Grafana is available directly through Caddy and the existing production TLS certificate at https://realvuxbaro.me/grafana/. Anonymous access is disabled; sign in with the Grafana admin account. A dedicated grafana.realvuxbaro.me hostname can replace this path after its Namecheap A record exists.

Grafana and Prometheus are also bound to EC2 loopback only. This keeps both SSH fallbacks available without exposing their native ports to the Internet:

ssh `
  -L 3002:127.0.0.1:3002 `
  -L 9090:127.0.0.1:9090 `
  <ssh-user>@52.0.158.61

Keep that session open and use http://localhost:3002/grafana/ for Grafana or http://localhost:9090 for the raw Prometheus UI. The same Grafana account is used through both the public HTTPS path and the SSH tunnel. The production password lives in the GitHub Actions secret GRAFANA_ADMIN_PASSWORD; do not use the Compose fallback password in production.

Prometheus intentionally has no public URL. Normally use Grafana -> Explore -> Prometheus; use its SSH tunnel only for low-level target or PromQL diagnostics. To investigate a slow request, open the request-latency panel, follow its exemplar/trace link, or paste the returned X-Trace-ID into Explore -> Tempo.

Production

Live at realvuxbaro.me: a single EC2 t3.large running infra/docker/docker-compose.prod.yml (postgres, qdrant, ai-service, web, Caddy for automatic Let's Encrypt TLS). Bedrock is reached through an IAM instance role — there are no long-lived AWS keys on the box or in any env file.

Pushing to master deploys: .github/workflows/deploy.yml SSHes in, resets to the pushed commit, rebuilds ai-service/web, reconciles the observability containers, reloads Caddy, runs migrations and verifies health, metrics and an exact request trace. Postgres and Qdrant data survive deploys because they live in named volumes rather than the containers.

This is interim infrastructure, not the end state. The intended target is still the team's self-hosted Gitea (company domain) plus their ArgoCD instance, per docs/adr/0002-argocd-gitops.md — that work is not started, not cancelled. Until it is deliberately started, the project stays on private GitHub, and the team's existing git.vinmec.tech/ai-team/gitops repository is reference-only: never push this project into it.

S
Description
Duoc Thu RAG medical chatbot (personal mirror)
Readme 27 MiB
Languages
Python 83.5%
TypeScript 15.5%
CSS 0.4%
Dockerfile 0.3%
JavaScript 0.2%
Other 0.1%