121 lines
4.9 KiB
Markdown
121 lines
4.9 KiB
Markdown
# 22 — CI/CD
|
||
|
||
## What exists
|
||
|
||
Exactly one workflow: `.github/workflows/deploy.yml`.
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
C[push to master] --> D["job: deploy<br/>ubuntu-latest"]
|
||
D --> S["appleboy/ssh-action@v1.0.3<br/>ssh to EC2_HOST as ubuntu"]
|
||
S --> G["git fetch + reset --hard origin/master"]
|
||
G --> B["docker compose up -d --build<br/>(prod + observability overlays)"]
|
||
B --> R["caddy validate + reload"]
|
||
R --> M["python -m migrate"]
|
||
M --> V["verification block — 15+ assertions"]
|
||
V -->|any fails| F["job fails; ai-service logs dumped"]
|
||
V -->|all pass| OK[done]
|
||
```
|
||
|
||
There is **no CI** in the usual sense — the pipeline stops at the *first* box of
|
||
the conventional diagram and jumps straight to deploy:
|
||
|
||
```
|
||
commit → [ lint ✗ ] → [ tests ✗ ] → [ build ✓ on prod host ] →
|
||
[ registry ✗ ] → [ manifest update ✗ ] → [ ArgoCD ✗ ] → rollout ✓
|
||
```
|
||
|
||
Concretely, none of the following runs anywhere in CI:
|
||
|
||
- `ruff` (configured in `apps/ai-service/pyproject.toml`, never invoked)
|
||
- `pytest` for either suite (555 tests)
|
||
- `tsc` / `next lint` / `turbo run lint` / `turbo run build`
|
||
- `helm lint` or `helm template`
|
||
- Any dependency or image vulnerability scan
|
||
|
||
A commit that breaks every test deploys to production.
|
||
|
||
## Triggers
|
||
|
||
```yaml
|
||
on:
|
||
push:
|
||
branches: [master]
|
||
workflow_dispatch:
|
||
```
|
||
|
||
No `pull_request` trigger, so a PR receives no automated feedback at all. No
|
||
environment protection rule, no required approval.
|
||
|
||
## Secrets used
|
||
|
||
| Secret | Use |
|
||
|---|---|
|
||
| `EC2_HOST` | SSH target |
|
||
| `EC2_SSH_KEY` | SSH private key |
|
||
| `GRAFANA_ADMIN_PASSWORD` | Passed through `envs:`; the script `test -n`s it and exports it for Compose |
|
||
|
||
No AWS credentials are needed — Bedrock is reached through the instance role.
|
||
|
||
## The verification block is the real quality gate
|
||
|
||
Everything after `docker compose up` is assertion, and `set -e` makes each one
|
||
fatal. In order:
|
||
|
||
| # | Assertion |
|
||
|---|---|
|
||
| 1 | `caddy validate --config /etc/caddy/Caddyfile` then `caddy reload` |
|
||
| 2 | `python -m migrate` inside the ai-service container |
|
||
| 3 | `GET ai-service:8000/health` |
|
||
| 4 | `GET ai-service:8000/ready` |
|
||
| 5 | `GET web:3000` |
|
||
| 6 | `POST /v1/rag/query` with a real condition→drug question; on failure, dump the last 200 ai-service log lines |
|
||
| 7 | Response contains `"decision":"answerable"` |
|
||
| 8 | Response contains `"section_key":"chi_dinh"` |
|
||
| 9 | `GET prometheus:9090/-/ready` |
|
||
| 10 | `GET tempo:3200/ready`, retried 12 × 5 s, dumping tempo logs on final failure |
|
||
| 11 | `GET grafana:3000/api/health` |
|
||
| 12 | Grafana datasource `prometheus` exists (admin-authenticated) |
|
||
| 13 | Grafana datasource `tempo` exists |
|
||
| 14 | Grafana dashboard `duocthu-observability` exists |
|
||
| 15 | `GET https://realvuxbaro.me/grafana/login` — through the public edge |
|
||
| 16 | A second `POST /v1/rag/query` with a generated correlation id; the `X-Trace-ID` response header must match `^[0-9a-f]{32}$` |
|
||
| 17 | After 20 s, `duocthu_requests_total` is queryable in Prometheus |
|
||
| 18 | That exact trace id is retrievable from `tempo:3200/api/traces/<id>`, retried 12 × 5 s |
|
||
|
||
Assertions 6–8 and 16–18 are unusually strong for a deploy script: one verifies
|
||
a real grounded answer from the real corpus, the other verifies that a specific
|
||
request's trace actually landed in Tempo.
|
||
|
||
## Consequences of the current design
|
||
|
||
| Property | Effect |
|
||
|---|---|
|
||
| Build happens on the production host | A build failure occurs *after* `git reset --hard`, so the checkout has already moved even if the new image never starts |
|
||
| No image tags | No artifact to roll back to; recovery is a revert commit plus a full rebuild |
|
||
| No test gate | Regressions are caught by the deploy smoke test (one behaviour) or by users |
|
||
| No PR feedback | Review is unassisted |
|
||
| Deploy is in-place | Brief downtime per service while it rebuilds and restarts |
|
||
| `postgres`/`qdrant` are not in the `up` list | Stateful services are never restarted by a deploy — good for uptime, but changes to their compose definitions silently do not apply |
|
||
|
||
## What `infra/ci/github-actions/README.md` promises
|
||
|
||
Five workflows, described as "not yet functional — filled in during Phase 6":
|
||
`ai-service-ci.yml`, `node-services-ci.yml`, `web-ci.yml`, `ingestion-ci.yml`,
|
||
`bump-image-tag.yml`. **None of them exists.** `bump-image-tag.yml` is the
|
||
linchpin of the GitOps flow described in
|
||
[21-kubernetes-and-argocd.md](21-kubernetes-and-argocd.md), so that flow cannot
|
||
run.
|
||
|
||
## Lowest-effort improvements, in order
|
||
|
||
1. Add a `pull_request` + `push` workflow that runs both pytest suites — the
|
||
commands are two lines and already work
|
||
([18-testing.md](18-testing.md)), and `EMBEDDING_PROVIDER=disabled` is the
|
||
only setup needed.
|
||
2. Add `ruff check` for `apps/ai-service` (config already present) and
|
||
`turbo run lint build` for the JS workspace.
|
||
3. Make `deploy` depend on those jobs.
|
||
4. Build and tag images in CI, push to a registry, and have the host pull a tag
|
||
— which also makes rollback possible.
|