150 lines
6.3 KiB
Markdown
150 lines
6.3 KiB
Markdown
# 21 — Kubernetes and ArgoCD
|
|
|
|
**Status: written, complete enough to render, never applied.** Nothing in this
|
|
document describes a running system. The live deployment is Docker Compose —
|
|
see [20-deployment.md](20-deployment.md).
|
|
|
|
Evidence that it is unapplied: three `TODO` placeholders in each ArgoCD
|
|
`Application`, `infra/k8s/base/*` and `infra/k8s/overlays/*` containing only
|
|
`.gitkeep`, no image registry anywhere in the repository, and no CI job that
|
|
renders, lints or applies the chart.
|
|
|
|
## Helm chart — `infra/helm/medical-chatbot/`
|
|
|
|
`Chart.yaml`: `medical-chatbot`, version `0.1.0`, appVersion `0.1.0`, type
|
|
`application`, no dependencies (everything is templated in-chart, not
|
|
sub-charted).
|
|
|
|
### Templates
|
|
|
|
| Template | Renders |
|
|
|---|---|
|
|
| `ai-service.yaml` | ConfigMap (all env vars), Deployment (+ optional `migrate` initContainer), Service |
|
|
| `web.yaml` | Deployment + Service |
|
|
| `data-services.yaml` | PostgreSQL and Qdrant workloads with PVCs |
|
|
| `observability-config.yaml` | Prometheus / Tempo / collector / Grafana configuration |
|
|
| `observability-workloads.yaml` | Their Deployments/StatefulSets, PVCs and Services |
|
|
| `ingress.yaml` | Ingress (disabled by default) |
|
|
| `secret.yaml` | `postgres-dsn`, Grafana admin password |
|
|
| `serviceaccount.yaml` | ServiceAccount (no RBAC bound) |
|
|
| `servicemonitor.yaml` | Prometheus-Operator `ServiceMonitor` (disabled by default) |
|
|
| `_helpers.tpl` | Name/label helpers |
|
|
|
|
### Rendered topology
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
ING["Ingress<br/>enabled: false by default<br/>class nginx, host duocthu.local"]
|
|
WEBS["Service web :3000"]
|
|
WEBD["Deployment web<br/>replicas 1"]
|
|
AIS["Service ai-service :8000"]
|
|
AID["Deployment ai-service<br/>replicas 1<br/>initContainer: python migrate.py"]
|
|
CM["ConfigMap ai-service<br/>QDRANT_URL, EMBEDDING_PROVIDER,<br/>ANSWER_PROVIDER, OTEL_*, MAX_*"]
|
|
SEC["Secret<br/>postgres-dsn, grafana admin"]
|
|
PGD[("postgres + PVC 5Gi")]
|
|
QDD[("qdrant + PVC 10Gi")]
|
|
OBS["prometheus 5Gi/7d · tempo 5Gi/24h<br/>otel-collector · grafana 2Gi"]
|
|
SM["ServiceMonitor<br/>enabled: false by default"]
|
|
|
|
ING --> WEBS --> WEBD --> AIS --> AID
|
|
CM --> AID
|
|
SEC --> AID
|
|
AID --> PGD
|
|
AID --> QDD
|
|
AID --> OBS
|
|
SM -.-> AIS
|
|
```
|
|
|
|
### Probes (the one thing genuinely production-shaped)
|
|
|
|
```yaml
|
|
readinessProbe: { httpGet: /ready, initialDelaySeconds: 5, periodSeconds: 10 }
|
|
livenessProbe: { httpGet: /health, initialDelaySeconds: 15, periodSeconds: 20 }
|
|
startupProbe: { httpGet: /health, failureThreshold: 30, periodSeconds: 5 }
|
|
```
|
|
|
|
The startup probe allows 150 s, which matters because `ai-service` builds its
|
|
whole runtime — including the Qdrant manifest check — at import time.
|
|
|
|
Pod annotations also set `prometheus.io/scrape`, `path` and `port`, so a
|
|
scrape-annotation-based Prometheus works even with `serviceMonitor.enabled=false`.
|
|
|
|
### Chart defaults that would break a naive install
|
|
|
|
| Value | Default | Consequence |
|
|
|---|---|---|
|
|
| `aiService.config.embeddingProvider` | `disabled` | `/v1/rag/query` returns 503 |
|
|
| `aiService.config.answerProvider` | `disabled` | No understanding, no generation, no multi-turn |
|
|
| `secret.postgresPassword` | `duoc_thu` | Default credential |
|
|
| `secret.grafanaAdminPassword` | `change-me` | Default credential |
|
|
| `ingress.enabled` | `false` | Nothing is reachable from outside the cluster |
|
|
| `qdrant.url` | `""` → in-cluster Service | A fresh Qdrant has **no corpus**, so the manifest check fails and the pod crash-loops |
|
|
|
|
That last one is the important one: the chart provisions an empty Qdrant, and
|
|
`ai-service` refuses to start against a collection with no manifest. A working
|
|
Kubernetes deployment needs a corpus load or a snapshot restore as a prerequisite
|
|
step that the chart does not model.
|
|
|
|
### Missing from the chart
|
|
|
|
No HPA, no PodDisruptionBudget, no `securityContext`/`runAsNonRoot`, no
|
|
`NetworkPolicy`, no anti-affinity, no `resources` on the initContainer, no
|
|
`imagePullSecrets` values beyond an empty list, and no init/job for corpus
|
|
loading.
|
|
|
|
## ArgoCD — `infra/argocd/applications/{dev,staging,prod}/app.yaml`
|
|
|
|
One `Application` per environment, each pointing at
|
|
`path: infra/helm/medical-chatbot` with `values.yaml` + `values-<env>.yaml`.
|
|
|
|
```yaml
|
|
spec:
|
|
project: default # TODO: confirm the team's ArgoCD project/RBAC scope
|
|
source:
|
|
repoURL: https://github.com/BaoVu2k4/vsf-duocthu.git # TODO: confirm once repo is created
|
|
targetRevision: master
|
|
destination:
|
|
server: https://kubernetes.default.svc # TODO: point at the team's target cluster
|
|
namespace: medical-chatbot-prod
|
|
syncPolicy: {} # intentionally NOT automated — prod sync requires manual approval
|
|
```
|
|
|
|
The three `TODO`s are present in all three files. `syncPolicy: {}` on prod is a
|
|
deliberate choice, not an omission — the comment says prod sync requires manual
|
|
approval in the ArgoCD UI/CLI.
|
|
|
|
## Intended GitOps flow (from `infra/argocd/README.md`)
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
DEV[merge to master] --> CI["CI builds + pushes an image per app"]
|
|
CI --> BUMP["CI bumps the image tag in<br/>values-<env>.yaml and pushes that commit"]
|
|
BUMP --> ARGO["ArgoCD (team-managed) detects the change"]
|
|
ARGO --> SYNC["sync — dev/staging auto, prod manual"]
|
|
SYNC --> K8S[cluster converges]
|
|
```
|
|
|
|
CI is explicitly forbidden from running `kubectl apply` or `helm upgrade`;
|
|
ArgoCD owns the deploy step, and promotion between environments is a Git
|
|
operation.
|
|
|
|
**None of that pipeline exists.** The five CI workflows
|
|
`infra/ci/github-actions/README.md` describes — including `bump-image-tag.yml`,
|
|
the linchpin of the flow — are named as "planned" and no workflow file exists
|
|
for any of them.
|
|
|
|
## Gap between the target and reality
|
|
|
|
| Element | Target | Actual |
|
|
|---|---|---|
|
|
| Runtime | Kubernetes | Docker Compose on one EC2 host |
|
|
| Deploy trigger | ArgoCD sync on a values-file commit | `appleboy/ssh-action` running `docker compose up --build` |
|
|
| Image source | Registry, tagged | Built on the production host, untagged |
|
|
| Environments | dev / staging / prod | prod only |
|
|
| Prod approval | Manual ArgoCD sync | Automatic on push to `master` |
|
|
| Secrets | Kubernetes Secret | `.env.prod` on the host + one GitHub secret |
|
|
| Config | ConfigMap from Helm values | `.env.prod` on the host |
|
|
|
|
ADR 0002 remains accepted and un-superseded; the interim Compose deployment was
|
|
a pragmatic step, not a decision reversal.
|