4.9 KiB
22 — CI/CD
What exists
Exactly one workflow: .github/workflows/deploy.yml.
flowchart LR
C[push to master] --> D["job: deploy<br/>ubuntu-latest"]
D --> S["appleboy/ssh-action@v1.0.3<br/>ssh to EC2_HOST as ubuntu"]
S --> G["git fetch + reset --hard origin/master"]
G --> B["docker compose up -d --build<br/>(prod + observability overlays)"]
B --> R["caddy validate + reload"]
R --> M["python -m migrate"]
M --> V["verification block — 15+ assertions"]
V -->|any fails| F["job fails; ai-service logs dumped"]
V -->|all pass| OK[done]
There is no CI in the usual sense — the pipeline stops at the first box of the conventional diagram and jumps straight to deploy:
commit → [ lint ✗ ] → [ tests ✗ ] → [ build ✓ on prod host ] →
[ registry ✗ ] → [ manifest update ✗ ] → [ ArgoCD ✗ ] → rollout ✓
Concretely, none of the following runs anywhere in CI:
ruff(configured inapps/ai-service/pyproject.toml, never invoked)pytestfor either suite (555 tests)tsc/next lint/turbo run lint/turbo run buildhelm lintorhelm template- Any dependency or image vulnerability scan
A commit that breaks every test deploys to production.
Triggers
on:
push:
branches: [master]
workflow_dispatch:
No pull_request trigger, so a PR receives no automated feedback at all. No
environment protection rule, no required approval.
Secrets used
| Secret | Use |
|---|---|
EC2_HOST |
SSH target |
EC2_SSH_KEY |
SSH private key |
GRAFANA_ADMIN_PASSWORD |
Passed through envs:; the script test -ns it and exports it for Compose |
No AWS credentials are needed — Bedrock is reached through the instance role.
The verification block is the real quality gate
Everything after docker compose up is assertion, and set -e makes each one
fatal. In order:
| # | Assertion |
|---|---|
| 1 | caddy validate --config /etc/caddy/Caddyfile then caddy reload |
| 2 | python -m migrate inside the ai-service container |
| 3 | GET ai-service:8000/health |
| 4 | GET ai-service:8000/ready |
| 5 | GET web:3000 |
| 6 | POST /v1/rag/query with a real condition→drug question; on failure, dump the last 200 ai-service log lines |
| 7 | Response contains "decision":"answerable" |
| 8 | Response contains "section_key":"chi_dinh" |
| 9 | GET prometheus:9090/-/ready |
| 10 | GET tempo:3200/ready, retried 12 × 5 s, dumping tempo logs on final failure |
| 11 | GET grafana:3000/api/health |
| 12 | Grafana datasource prometheus exists (admin-authenticated) |
| 13 | Grafana datasource tempo exists |
| 14 | Grafana dashboard duocthu-observability exists |
| 15 | GET https://realvuxbaro.me/grafana/login — through the public edge |
| 16 | A second POST /v1/rag/query with a generated correlation id; the X-Trace-ID response header must match ^[0-9a-f]{32}$ |
| 17 | After 20 s, duocthu_requests_total is queryable in Prometheus |
| 18 | That exact trace id is retrievable from tempo:3200/api/traces/<id>, retried 12 × 5 s |
Assertions 6–8 and 16–18 are unusually strong for a deploy script: one verifies a real grounded answer from the real corpus, the other verifies that a specific request's trace actually landed in Tempo.
Consequences of the current design
| Property | Effect |
|---|---|
| Build happens on the production host | A build failure occurs after git reset --hard, so the checkout has already moved even if the new image never starts |
| No image tags | No artifact to roll back to; recovery is a revert commit plus a full rebuild |
| No test gate | Regressions are caught by the deploy smoke test (one behaviour) or by users |
| No PR feedback | Review is unassisted |
| Deploy is in-place | Brief downtime per service while it rebuilds and restarts |
postgres/qdrant are not in the up list |
Stateful services are never restarted by a deploy — good for uptime, but changes to their compose definitions silently do not apply |
What infra/ci/github-actions/README.md promises
Five workflows, described as "not yet functional — filled in during Phase 6":
ai-service-ci.yml, node-services-ci.yml, web-ci.yml, ingestion-ci.yml,
bump-image-tag.yml. None of them exists. bump-image-tag.yml is the
linchpin of the GitOps flow described in
21-kubernetes-and-argocd.md, so that flow cannot
run.
Lowest-effort improvements, in order
- Add a
pull_request+pushworkflow that runs both pytest suites — the commands are two lines and already work (18-testing.md), andEMBEDDING_PROVIDER=disabledis the only setup needed. - Add
ruff checkforapps/ai-service(config already present) andturbo run lint buildfor the JS workspace. - Make
deploydepend on those jobs. - Build and tag images in CI, push to a registry, and have the host pull a tag — which also makes rollback possible.