11 KiB
Decision update, same day, later session: owner confirmed intent to run one EC2 only (k3s). This supersedes PR A step 2 below (reduce to
workflow_dispatch) —deploy.yml,rollback.yml, and the never-appliedinfra/argocd/applications/**scaffold (wrong Application names, deadvalues-prod.yamlreference, placeholder team-repo TODOs) were deleted outright on branchagent/retire-compose-cicd, not disabled. D1/D2/D3/D4 are closed by deletion rather than by fixing the probe.docs/operations.mdDeploy/Rollback sections rewritten to describe the actual k3s/ArgoCD path. PR #23 (both commits above) merged tomasterase804817.Update, same day, further into the session: owner gave explicit go-ahead and Compose EC2
i-039fc8f6102467a54was stopped (not terminated) viaaws ec2 stop-instances. Confirmed transitionrunning→stopping. Root EBS still carriesDeleteOnTermination=true, so it is intact and restartable, but it is no longer a live rollback target —docs/operations.mdRollback section now documents this as the current state, including the manualstart-instances+ compatibility-check steps needed before ever trusting it as a DNS fallback again. Nothing else about the box (AMI, EBS, tags, security group) was touched.PR B (Helm hygiene: drop the redundant
AWS_REGIONenv block, add a baseline render-diff tohelm-chart.yml) and the remaining half of PR C (an actual one-command k3s rollback script/workflow, not just the manual runbook now indocs/operations.md) remain open.
Plan — make the CI/CD path safe after the k3s cutover (2026-08-18)
Written by Claude (Opus) for execution by another agent. Every claim below was verified on 2026-08-18 by the command shown next to it. Re-verify before acting — runtime state can change after this snapshot.
1. The topology changed and two workflows never noticed
This is the root cause of everything in this plan. The 2026-08-17 cutover moved
realvuxbaro.me from the Compose EC2 to the k3s/ArgoCD cluster, but the
Compose-era workflows still describe and probe the old world.
| Fact | Verified by |
|---|---|
realvuxbaro.me → 44.206.194.195 (k3s) |
nslookup realvuxbaro.me 8.8.8.8 |
readytochat.realvuxbaro.me → 44.206.194.195 — same cluster, same release |
same |
Compose EC2 52.0.158.61 is off DNS; rollback target only |
values-production.yaml ingress.host + DNS above |
Compose is still at df57e6b |
git log df57e6b..52e8828 -- <deploy.yml paths> returns empty |
App code at df57e6b == app code at master |
the intervening commits touch only docs/CI/coordination |
So: the Compose rollback is currently valid, and the window to fix this is now, before the first post-cutover app change lands.
values-production.yaml sets ingress.host: realvuxbaro.me and is rendered by
ArgoCD Application medical-chatbot-app — the same Application
.github/scripts/sync_practice_argocd.py repoints (APP_NAME = "medical-chatbot-app"). "Practice" and production are the same release.
2. Confirmed defects, most dangerous first
D1 — both Compose workflows verify the wrong machine (P0)
deploy.yml:88 and rollback.yml:57 both end with:
docker run --rm --network docker_default curlimages/curl -sf \
-o /dev/null https://realvuxbaro.me/grafana/login
That hostname now resolves to k3s, not the box the workflow just rebuilt. The check passes by hitting a completely different server.
This is worst in rollback.yml, whose whole purpose is to be trustworthy in an
emergency: it prints Rollback to <sha> verified healthy on the strength of a
probe that never touched the rolled-back box. A broken rollback would report
success.
D2 — one bad commit poisons production and the rollback (P0)
deploy.yml and build-practice-images.yml trigger on the same four paths:
apps/ai-service/**, apps/web/**, packages/**,
ingestion/data/verified/drug_entities.json.
A single push to master therefore rolls production forward on k3s and
rebuilds the Compose box that is supposed to be the known-good fallback. The
"proven rollback" only holds while Compose stays on a good commit.
D3 — build-practice-images.yml header comment is now false (P1)
"Does not touch deploy.yml or the production EC2/Compose stack — production never pulls a GHCR image and isn't ArgoCD-managed at all, so this workflow has no path to affect it."
Production is ArgoCD-managed and does pull GHCR images. This workflow is the production deploy pipeline. The comment invites exactly the push that breaks production, and its name reinforces the error.
D4 — a dead ArgoCD manifest sits in Git (P1)
infra/argocd/applications/prod/app.yaml lists
valueFiles: [values.yaml, values-prod.yaml], but values-prod.yaml was
deleted in 6468b16. It also declares syncPolicy: {} with a comment claiming
prod sync needs manual approval — the live Application is automated. Applying
this file would fail or deploy something wrong. dev/ and staging/ under the
same directory have not been checked and may share the defect.
D5 — the WIP Helm change shadows an existing variable (P2)
Uncommitted in the main worktree. The chart already emits AWS_REGION in
the ConfigMap from aiService.config.awsRegion
(templates/ai-service.yaml:15, asserted in helm-chart.yml). The WIP diff
adds a second AWS_REGION as a container env: entry — and in Kubernetes an
explicit env: overrides envFrom, so enabling it would silently shadow
the ConfigMap. Inert today only because aws.region defaults to "".
The static-credentials half is a genuine gap and worth keeping.
D6 — helm-chart.yml asserts invariants but never diffs (P2)
It renders both live releases and checks a strong list (Qwen, rerank, TLS
secrets, Grafana role, refute volumeClaimTemplates). But nothing compares the
render against the previous commit, so a chart change that alters anything
outside that list reaches production silently.
D7 — production has no rollback workflow at all (P2)
rollback.yml targets secrets.EC2_HOST — the Compose box. Nothing rolls back
k3s. Real production rollback today is a manual Namecheap A-record revert
(~60s TTL) or an ArgoCD revision/tag revert, neither written down.
3. Execution plan
Ground rules
- Never push to
master. Every change goes through a PR. ArgoCD auto-syncsmasterwithselfHeal+prune; a merge touchinginfra/helm/**applies to production with no human gate. ci.ymlruns on every PR (ruff, pytest, ingestion tests, web lint+build).helm-chart.ymlruns on PRs touchinginfra/helm/**. Both must be green.- Do not touch
git.vinmec.tech, team ArgoCD/k3s, or team repos. - Do not start, stop, or terminate any EC2 instance without an explicit go.
- Do not edit files in the other worktrees (
D:\VSF-DUOCTHU-codex-*,D:\VSF-DUOCTHU-claude-gitops); checkcoordination/for active claims and file your own claim before starting.
PR A — workflow safety (D1, D2, D3, D4)
Touches only .github/** and infra/argocd/**. ArgoCD renders
infra/helm/medical-chatbot, so this PR cannot alter production manifests.
Confirm that rather than assume it.
- Fix the misdirected probes (D1). In
deploy.ymlandrollback.yml, make the Caddy/Grafana check target the box being deployed instead of a public DNS name that now points elsewhere — e.g. resolve the hostname to the local Caddy container so TLS and routing are still exercised. Verification gate: prove the fixed check fails when Caddy is broken. A probe that cannot fail is the defect being fixed, not a fix. (See theset -e/! grepsilent-pass traps already documented inhelm-chart.yml.) - Stop
deploy.ymlfiring on push (D2). Reduce it toworkflow_dispatchonly. Compose then stays pinned atdf57e6b— a stable rollback rather than one that tracksmaster. State the trade-off in the commit message: the fallback stops drifting, but also stops receiving fixes, so it goes stale as production moves. That is acceptable for a time-boxed acceptance window and is the subject of PR D. - Tell the truth in
build-practice-images.yml(D3). Replace the false header comment. Renaming the workflow to name it as the production deploy path is preferable — trap: it self-references in its ownpaths:filter and the sync script path, so both must be updated together or the workflow silently stops triggering. - Remove or correct
infra/argocd/applications/prod/app.yaml(D4). Checkdev/andstaging/in the same directory for the same rot. Deleting is fine if nothing applies them; verify that first.
PR B — Helm hygiene (D5, D6)
Touches infra/helm/**, so merging does reach production. Highest care.
- Drop the redundant
aws.regionblock from the WIP diff (D5); keep the static-credentials support, still defaulted off. Guard against a nilawskey so a values file that omits it cannot break the render — a template error here means ArgoCD cannot sync production at all. - Add a baseline render diff to
helm-chart.yml(D6). Rendervalues-production.yamlandvalues-production-data.yamlat the PR base and at HEAD, then surface the diff in the job summary. The goal is that no chart change ever reaches production without a human having seen exactly what it does to the manifests. Verification gate: the diff for PR B itself must be empty — the static-credentials change is defaults-off and must render byte-identically. If it is not empty, stop and explain why before merging.
Note: helm is not installed on this workstation (helm: command not found
in both bash and PowerShell). Either install it or rely on the CI render — but
do not claim the render is unchanged without one of the two actually running it.
PR C — document the real rollback (D7)
- Write the production rollback runbook: revert the ArgoCD Application to the
previous image tag, and/or revert the
realvuxbaro.meA record to52.0.158.61. Note that the Namecheap edit is a manual owner step — a harness permission classifier has blocked agent form input on that page before, so the runbook must not assume an agent can do it.
PR D — decide the Compose lifecycle (owner call, not an agent call)
- Compose is a second
t3.largerunning purely as a fallback. Once the acceptance window closes, ArgoCD's own revision history covers rollback and the instance is redundant. Surface the choice and the monthly cost; do not act on it without an explicit go.
Out of scope
Feature-List-AI-Duoc-thu-V1.md, presentation/, and .claude/skills/ are
untracked and match no workflow trigger path. Committing them is inert and can
be a separate trivial commit — keep it out of PRs A–D.
4. Order and why
D1 first: an untrustworthy rollback is worse than no rollback, because it fails silently at the moment of maximum pressure. D2 next: it is the defect that would consume the rollback. Everything after is hardening.