Files
duocthu/docs/24-production-operations.md
T

8.3 KiB
Raw Blame History

24 — Production operations

Production is one EC2 host running Docker Compose. Every command below assumes an SSH session on that host in ~/app/infra/docker, matching what .github/workflows/deploy.yml does.

The Compose project name is docker (the directory name), so containers are named docker-<service>-1.

Compose invocation

Both overlay files are always used together:

cd ~/app/infra/docker
COMPOSE="sudo docker compose -f docker-compose.prod.yml -f docker-compose.observability.yml"

The observability overlay is what sets OTEL_ENABLED=true on ai-service, so omitting it silently disables tracing.

Start / stop / restart

$COMPOSE ps
$COMPOSE up -d ai-service web caddy                  # start/refresh app tier
$COMPOSE restart ai-service                          # restart one service
$COMPOSE stop ai-service
$COMPOSE logs -f --tail 200 ai-service

ai-service builds its whole runtime at import time, so a restart re-runs the corpus-manifest check. If that check fails the container exits immediately and keeps restarting — check the logs for ManifestMismatch before assuming a crash loop is resource-related.

Deploy

Normal path: push to master. The workflow SSHes in, resets the checkout, rebuilds, reloads Caddy, migrates and runs ~18 assertions (22-ci-cd.md).

Manual equivalent:

cd ~/app && git fetch origin master && git reset --hard origin/master
cd infra/docker
export GRAFANA_ADMIN_PASSWORD='<value>'
sudo -E docker compose -f docker-compose.prod.yml -f docker-compose.observability.yml \
  up -d --build ai-service web prometheus tempo otel-collector grafana caddy
sudo docker exec docker-caddy-1 caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile
sudo docker exec docker-caddy-1 caddy reload   --config /etc/caddy/Caddyfile --adapter caddyfile
sudo docker exec docker-ai-service-1 python -m migrate

Note postgres and qdrant are deliberately absent from that list — a code deploy never restarts the stateful services.

Rollback

There is no image to roll back to (images are built on the host, untagged). The procedure is:

cd ~/app
git reset --hard <last-good-sha>      # or push a revert to master and let CI deploy
cd infra/docker && sudo -E docker compose -f docker-compose.prod.yml \
  -f docker-compose.observability.yml up -d --build ai-service web

A rollback that crosses a migration is not covered — migrations are forward-only with no down scripts.

Health checks

NET=docker_default
sudo docker run --rm --network $NET curlimages/curl -sf http://ai-service:8000/health
sudo docker run --rm --network $NET curlimages/curl -sf http://ai-service:8000/ready
sudo docker run --rm --network $NET curlimages/curl -sf -o /dev/null http://web:3000
sudo docker run --rm --network $NET curlimages/curl -sf http://prometheus:9090/-/ready
sudo docker run --rm --network $NET curlimages/curl -sf http://tempo:3200/ready
sudo docker run --rm --network $NET curlimages/curl -sf http://grafana:3000/api/health

ai-service publishes no host port, so every check goes through a throwaway container on the Compose network — the same technique the deploy workflow uses.

Smoke test a real answer

sudo docker run --rm --network docker_default curlimages/curl -sf \
  -X POST http://ai-service:8000/v1/rag/query \
  -H 'Content-Type: application/json' \
  --data '{"query":"Đợt gout cấp có thuốc nào được Dược thư ghi chỉ định?",
           "subject_scope":"human","intent":"fact_lookup",
           "conversation_id":"ops-smoke"}'

Expect "decision":"answerable" and at least one citation with "section_key":"chi_dinh" — the same two assertions the deploy makes.

Database operations

# psql
sudo docker exec -it docker-postgres-1 psql -U duoc_thu -d duoc_thu

# apply migrations
sudo docker exec docker-ai-service-1 python -m migrate

Useful queries:

-- recent decisions
SELECT created_at, decision, reason, resolved_drug_id
FROM rag_retrieval_trace ORDER BY created_at DESC LIMIT 50;

-- abstain reasons over the last day
SELECT reason, count(*) FROM rag_retrieval_trace
WHERE decision = 'abstain' AND created_at > now() - interval '1 day'
GROUP BY reason ORDER BY 2 DESC;

-- find a support request by either correlation id
SELECT * FROM rag_retrieval_trace WHERE correlation_id = '<id>';
SELECT * FROM rag_retrieval_trace WHERE otel_trace_id  = '<32-hex>';

-- negative feedback with the question that caused it
SELECT f.created_at, f.rating, f.comment, t.query_text, t.decision, t.reason
FROM rag_answer_feedback f JOIN rag_retrieval_trace t USING (trace_id)
WHERE f.rating = 'not_helpful' ORDER BY f.created_at DESC LIMIT 50;

Backup and restore

No backup automation exists in this repository. What the code supports:

# PostgreSQL logical dump
sudo docker exec docker-postgres-1 pg_dump -U duoc_thu duoc_thu > duoc_thu_$(date +%F).sql

# Qdrant snapshot (HTTP API, from inside the network)
sudo docker run --rm --network docker_default curlimages/curl -s -X POST \
  http://qdrant:6333/collections/duocthu_v1/snapshots

Both are manual. Whether EBS snapshots are configured on the instance cannot be determined from the repository.

Re-indexing / re-ingestion

Two situations, with very different costs:

Corpus content unchanged, moving or restoring it — snapshot and restore the Qdrant collection. Free and exact; ingestion/README.md recommends it explicitly.

Corpus content changed — the full pipeline must re-run and the embed step costs real AWS Bedrock spend on a personal account. ingestion/README.md requires explicit approval for any specific run. Order:

  1. python -m ingestion.cli run → new monographs.jsonl
  2. python -m ingestion.cli chunk → new chunks.jsonl
  3. python -m ingestion.cli chunk-readymust exit 0
  4. python -m ingestion.load.run --provider cohere-v4 --collection duocthu_v2 …

Use a new collection name. The loader refuses to write a different corpus_sha256 into an existing collection (CorpusMismatch), which is the intended behaviour, not an obstacle to work around. Then point QDRANT_COLLECTION at the new collection and restart ai-service; the startup manifest check verifies the binding. Keep the old collection until the new one is confirmed — that is the rollback.

The embedding cache in ingestion/data/processed/embeddings/ is keyed by content hash, so unchanged chunks are not re-paid for.

Grafana

Reachable at https://realvuxbaro.me/grafana/ with the admin credentials from GRAFANA_ADMIN_PASSWORD. Locally on the host: http://127.0.0.1:3002. Provisioned datasources prometheus and tempo; dashboard uid duocthu-observability.

Following one request end to end

  1. Take X-Correlation-ID or X-Trace-ID from the user's response headers (the UI surfaces traceId on each message).
  2. SELECT * FROM rag_retrieval_trace WHERE correlation_id = … → the resolved scope, decision, reason and citations.
  3. Open the trace id in Grafana → Tempo → per-stage spans (rag.stage.understanding, retrieval, generation, entailment) with duocthu.* attributes.
  4. Cross-check duocthu_generation_rejected_total{reason=…} and duocthu_abstention_total{reason=…} in Prometheus for the same window.

Cost control

Every chat turn makes 38 Bedrock calls on a personal AWS account. The only guard is the in-memory rate limiter in apps/web/middleware.ts (12/min, 120/hour per IP for /api/chat). There is no budget alarm, no per-day cap and no authentication in the repository. Scaling web past one replica multiplies the effective allowance.

Incident quick reference

Symptom First check
Every answer is an abstain duocthu_abstention_total{reason} — a single dominant reason points at a provider or corpus problem
ai-service restart loop docker logs docker-ai-service-1 for ManifestMismatch
503 from /v1/rag/query EMBEDDING_PROVIDER in .env.prod, and whether the manifest check passed
Answers take ~60 s then fail duocthu_generation_rejected_total{reason="request_budget_exhausted"}
429s Rate limiter; X-RateLimit-* headers on the response
No traces in Grafana Was the observability overlay included in the last up?

Full table in 25-troubleshooting.md.