Files
duocthu/docs-legacy/16-security.md
T

9.4 KiB
Raw Blame History

16 — Security

Findings from reading the code, not a penetration test. Each row states what exists and what does not; nothing here should be read as an assurance.

Summary

Control State
Authentication Not implemented — no login, no token, anywhere
Authorization Not implemented — no roles, no per-user scoping
TLS in transit (public edge) Implemented — Caddy with automatic ACME
TLS inside the Compose network Not implemented — plain HTTP between containers
Input validation Partially implemented
Prompt-injection handling Implemented (input fencing + output verification)
Rate limiting Partially implemented — frontend only, in-memory, two routes
Secrets management Partially implemented — IAM role for AWS; a default DB credential is committed
Metrics endpoint auth Configured but unset
Container hardening Not implemented — root user, no read-only FS, no capability drop
K8s security context / NetworkPolicy Not found in the Helm chart
Data retention / redaction Not implemented
Audit logging Partially implemented — every answer is traced; no auth identity to attach
Dependency scanning Not found — no Dependabot, no pip-audit, no npm audit in CI

Authentication and authorization

There is none. POST /api/chat accepts an unauthenticated request from anyone who can reach https://realvuxbaro.me, and each turn spends AWS Bedrock credit on a personal account. middleware.ts states this plainly:

It is a cost and abuse guard, not a security control. It does not authenticate anyone and must not be described as if it does.

apps/auth-service and apps/user-service contain no source. The pre-existing docs/architecture.md assigns JWT validation to api-gateway, which does not exist.

Conversation isolation

conversation_id is an arbitrary client-supplied string, at most 128 characters, with no ownership check. PostgresConversationStore.recent() returns the last lines for whatever id is sent, and those lines are placed into the understanding prompt. Anyone who knows or guesses another session's id can read its conversation history into their own turn's model context. There is no entropy requirement on the id.

Transport

  • Caddy terminates TLS for realvuxbaro.me with automatic certificates and proxies /grafana/*grafana:3000 and everything else → web:3000.
  • ai-service, postgres, qdrant, prometheus, tempo and otel-collector publish no host ports in the production files; Prometheus and Grafana bind to 127.0.0.1 only in the observability overlay. Reaching ai-service therefore requires being on the Compose network.
  • No HSTS, CSP, X-Frame-Options or other security headers are set — the Caddyfile has no header directive and next.config.js defines no headers().
  • CORS is not configured on the FastAPI app: no CORSMiddleware is added, so the browser's default same-origin policy is what protects it. That is adequate only because the browser never talks to ai-service directly.

Input validation

Surface Validation
POST /v1/rag/query Pydantic: query 14000 chars, conversation_id ≤128, subject_scope/intent enum-constrained
POST /v1/rag/feedback trace_id must parse as a UUID, rating literal-constrained, comment ≤2000
GET /v1/rag/suggest q is a bare string with no max length
POST /api/chat (web) content non-empty and ≤4000, conversationId ≤128
X-Correlation-ID Regex-validated, regenerated when malformed — on both sides
Model output _parse_claims, _sanitize_quick_replies, _clean_enum, _clean_float (0 < kg ≤ 500) — every field validated, fail-closed

SQL access uses parameterised psycopg queries throughout; no string interpolation into SQL was found.

Prompt injection

Two layers, both described in 11:

  • Input — the user's text is fenced in markers stripped from the input first, and all three system prompts carry _UNTRUSTED_RULE telling the model the fenced text is data.
  • Output — a fabricated figure cannot survive grounding.verify, citations are assembled from retrieved metadata rather than from model prose, and a claim the entailment judge does not confirm is discarded.

Tested by tests/test_prompt_untrusted_input.py.

Residual exposure: the understanding prompt embeds raw conversation history, and the drug_id labels in _prompt_evidence_texts come from corpus payloads (trusted). A user cannot inject into the evidence section.

Rate limiting

apps/web/middleware.ts, in-process, keyed by the left-most X-Forwarded-For entry:

Route prefix Rules
/api/chat 12 per minute and 120 per hour
/api/suggest 120 per minute
everything else under /api/* no limit — including /api/pdf (37 MB per request) and /api/feedback

Stated limitations, from the source comments: counters are per process (a second web replica doubles the allowance), the key is an IP so a shared NAT is throttled as one caller, and the correct home is Redis or the unbuilt gateway. A rejected request is deliberately not recorded, so a hammering client cannot extend its own lockout.

An unknown IP falls back to the shared key "unknown" rather than to unlimited — the comment notes that mattering.

Secrets

See the inventory in 15-configuration.md.

The concrete issue: PostgreSQL credentials duoc_thu / duoc_thu are committed in infra/docker/docker-compose.prod.yml (as POSTGRES_USER/POSTGRES_PASSWORD) and as the Helm default secret.postgresPassword. Exposure today is bounded because PostgreSQL publishes no host port in production, so the credential is only usable from inside the Compose network — but it is a default credential in version control, and the Helm path would carry it into a cluster where the blast radius is larger.

infra/helm/.../values.yaml also ships grafanaAdminPassword: change-me. The deploy workflow requires a real GRAFANA_ADMIN_PASSWORD and fails fast if it is empty (test -n "${GRAFANA_ADMIN_PASSWORD:-}").

AWS access is via the EC2 instance role — no keys in any file. The two policy documents under infra/aws/iam/ scope Bedrock invocation.

Metrics endpoint

GET /metrics supports an optional bearer token compared with hmac.compare_digest (constant time — a == on a shared secret leaks its prefix through timing). METRICS_TOKEN defaults to empty, i.e. no auth. main.py explains the trade: the endpoint is unreachable from the internet today because Caddy proxies only web and ai-service publishes no host port, and it "stops being safe the moment the service is exposed through an Ingress, which the Helm chart now makes possible". Metrics carry query volumes, provider failure counts and abstain reasons.

Grafana exposure

Grafana is internet-reachable at https://realvuxbaro.me/grafana/. The overlay sets GF_AUTH_ANONYMOUS_ENABLED=false and a real admin password from the environment, with GF_SERVER_ROOT_URL and GF_SERVER_SERVE_FROM_SUB_PATH for the subpath. The local-dev Compose file enables anonymous admin access, with a comment forbidding carrying that into a deployed stack.

Container and cluster hardening

apps/ai-service/Dockerfile:

  • runs as root (no USER directive);
  • installs gcc into the runtime image rather than using a build stage;
  • pins dependency ranges inline instead of installing from pyproject.toml, so the image's dependency set can drift from the project's;
  • has no HEALTHCHECK.

apps/web/Dockerfile runs as root and copies the entire /repo (source and node_modules) into the runtime stage rather than using Next's standalone output.

In infra/helm/medical-chatbot/: no securityContext, no runAsNonRoot, no readOnlyRootFilesystem, no NetworkPolicy, no PodDisruptionBudget. A ServiceAccount is created but no RBAC is bound to it. Probes are configured (/ready, /health, plus a startup probe).

Data privacy

The product invites clinicians to type patient context — age, weight, comorbidities, allergies, previous ADRs, current medications, eGFR/CrCl/CKD stage, Child-Pugh, pregnancy status, lab values (rag/clinical.py).

Consequences, all currently unaddressed:

  • rag_retrieval_trace.query_text and rag_conversation_turn.line store that text verbatim, forever. No retention, no deletion path, no redaction.
  • The same text is sent to AWS Bedrock on every turn.
  • agent.py logs turn timings at WARNING level; understanding.py logs the model's raw output on a parse failure (logger.warning("… returned unparseable JSON: %r", raw_text)) and answer.py logs claims and repair verdicts — so fragments of user and model text can reach container logs.
  • There is no consent flow, no DPA, no anonymisation, and no access control on the database.

Dependency and supply-chain risk

  • No Dependabot, no pip-audit, no npm audit, no SBOM, no image scanning anywhere in .github/.
  • pnpm-lock.yaml is committed; there is no Python lockfile — the Dockerfile installs unpinned ranges ("fastapi>=0.115,<1", "boto3" with no bound at all), so two builds of the same commit can differ.
  • qdrant/qdrant:latest is unpinned.
  • CI runs no tests before deploying (see 22-ci-cd.md).