8.6 KiB
02 — System architecture
Architectural style
As built: a two-service application (web + ai-service) plus an offline
batch pipeline, deployed as Docker Compose services on one host. Communication
is synchronous HTTP/JSON. There is no message broker, no queue, no async
worker, and no service mesh.
As designed on paper: a seven-service microservices platform
(api-gateway, auth-service, user-service, chat-service, ai-service,
web, ingestion), described in the pre-existing docs/architecture.md. Four
of those seven do not exist — their directories hold a README.md and a
four-line package.json with no dependencies and no source files. The
monorepo scaffolding (pnpm workspace entries, infra/k8s/base/<service>/
directories) still reserves their names.
Both facts matter: the second explains why apps/, pnpm-workspace.yaml and
the Helm chart look bigger than the running system.
Component diagram — what actually runs
flowchart TB
subgraph browser["Browser"]
UI["Chat UI<br/>ChatPanel.tsx · 65s abort"]
end
subgraph ec2["EC2 host — docker compose"]
CADDY["caddy:2-alpine<br/>:80 :443 · ACME TLS"]
subgraph webc["web (Next.js 14, :3000)"]
MW["middleware.ts<br/>in-memory IP rate limit"]
BFF["/api/chat · /api/suggest<br/>/api/feedback · /api/pdf"]
end
subgraph aic["ai-service (FastAPI, :8000)"]
HTTP["routers/rag.py"]
AGENT["RagAgent"]
RET["RetrievalService"]
ANS["GroundedAnswerService"]
end
PG[("postgres:16-alpine")]
QD[("qdrant/qdrant")]
subgraph obs["observability overlay"]
OTELC["otel-collector"]
TEMPO["tempo"]
PROM["prometheus"]
GRAF["grafana :3002 (127.0.0.1)"]
end
end
BEDROCK["AWS Bedrock<br/>embed-v4 · rerank-v3.5 · Converse"]
UI -->|HTTPS| CADDY
CADDY -->|"/*"| MW --> BFF
CADDY -->|"/grafana/*"| GRAF
BFF -->|"POST /v1/rag/query"| HTTP
HTTP --> AGENT
AGENT --> RET
AGENT --> ANS
RET --> QD
ANS -->|generation + entailment| BEDROCK
AGENT -->|understanding| BEDROCK
RET -->|embed + rerank| BEDROCK
HTTP --> PG
AGENT --> PG
aic -->|OTLP/HTTP| OTELC --> TEMPO
PROM -->|scrape /metrics| aic
GRAF --> PROM
GRAF --> TEMPO
ai-service publishes no host port in docker-compose.prod.yml; it is
reachable only on the Compose network. Caddy proxies web and grafana only.
Service boundaries
| Service | Owns | Depends on | Stateless? |
|---|---|---|---|
web |
Rendering, reason-code → Vietnamese message mapping, citation grouping, rate limiting | ai-service over HTTP; the source PDF on a bind-mounted path |
No — rate-limit counters are per-process in memory |
ai-service |
Understanding, retrieval, generation, grounding, citations, traces | Qdrant, PostgreSQL, Bedrock | Mostly — RagAgent keeps two in-process dicts (_last_frame, _clarify_streak); conversation text is in PostgreSQL |
ingestion |
Turning the PDF into chunks.jsonl and Qdrant points |
Qdrant, Bedrock, local disk | N/A — batch |
The stateful detail that constrains scaling
RagAgent (rag/agent.py) holds three dicts:
self._history: dict[str, list[str]] # unused when a ConversationStore is configured
self._last_frame: dict[str, QueryFrame] # ALWAYS in-process
self._clarify_streak: dict[str, int] # ALWAYS in-process
PostgresConversationStore replaces _history only. _last_frame carries the
structured merge that stops the model re-asking an answered clarify question,
and _clarify_streak drives the clarify circuit breaker. Both are lost on
restart and not shared between replicas. Running more than one ai-service
replica therefore degrades multi-turn quality in a way nothing detects. The Helm
chart's aiService.replicaCount defaults to 1; nothing enforces it.
Module boundaries inside ai-service
Ports-and-adapters, enforced by import discipline rather than by tooling:
rag/ports.pydeclaresRetriever,SectionRetriever,ParentStore,AnswerGenerator,Reranker, plus three exception types (QueryEmbeddingUnavailable,AnswerGenerationUnavailable,RerankUnavailable) that adapters raise and the domain catches.adapters/is the only placeqdrant_client,boto3,psycopgandprometheus_clientare imported — and always lazily, inside a method, so the domain imports cleanly on a machine with none of them installed.bootstrap.pyis the composition root. Nothing else constructs an adapter.
One boundary is looser than the protocol suggests: RetrievalService reaches
optional retriever capabilities with getattr(self._retriever, "find_by_section", None) rather than through a declared protocol. find_by_indication,
search_indication, search_lexical and find_by_drug are all discovered this
way and none of them appear in ports.py. A retriever missing one silently
disables a whole route instead of failing a type check.
Synchronous communication
Every hop is a blocking HTTP or SDK call. One answerable turn issues, in sequence:
POST /api/chat(browser → web)POST /v1/rag/query(web → ai-service)- Bedrock Converse — understanding
- Qdrant
scroll/query_points— retrieval (1–N calls) - Bedrock Converse — generation (plus one retry on
evidence_sufficient=false) - Bedrock Converse — entailment
- optional Bedrock Converse ×2 — completeness repair + its re-verification
- PostgreSQL insert — trace
Measured production latencies recorded in ChatPanel.tsx (n=8, 2026-08-11):
6.2 / 6.4 / 8.4 / 10.9 / 12.4 / 21.7 / 25.1 / 40.3 seconds.
Asynchronous communication
Not found. No broker, no queue, no background worker, no SSE, no
WebSocket, no streaming response. web's /api/chat awaits the full upstream
response before replying.
Failure boundaries
| Boundary | Policy | Implemented in |
|---|---|---|
| Corpus/model manifest mismatch at startup | Fail closed, crash the process | bootstrap.py::_verify_corpus_manifest → rag/manifest.py |
| Query embedder unreachable | Fail closed — abstain, never a 500 | rag/service.py, rag/ports.py |
| Generator unreachable / malformed / budget exhausted | Fail closed — abstain with a specific reason code | rag/answer.py::_generate |
| Understanding call fails | Fail closed — abstain, tagged system_error |
rag/understanding.py::understand |
| Grounding or entailment rejects | Fail closed — abstain | rag/answer.py |
| Reranker unreachable | Fail open — keep original order | rag/service.py::_rerank |
| Sufficiency check unreachable | Fail open — proceed to generate | rag/answer.py::_check_sufficiency |
| PostgreSQL trace write fails | Fail open — answer returned, TRACE_WRITE_FAILED counter |
routers/rag.py |
| Conversation store read/write fails | Fail open — this turn has no memory | rag/agent.py::_get_history / _remember |
| Metrics package missing | Degrade — NullMetrics |
bootstrap.py::_build_metrics |
OpenTelemetry packages missing / OTEL_ENABLED=false |
Degrade — no-op tracer | rag/telemetry.py |
The asymmetry is deliberate and documented in-code: anything that could change what is stated fails closed; anything that only affects quality or observability fails open.
Deployment units
| Unit | Image | Built by |
|---|---|---|
ai-service |
apps/ai-service/Dockerfile — python:3.12-slim, deps pinned inline (not from pyproject.toml) |
docker compose up --build on the EC2 host |
web |
apps/web/Dockerfile — 3-stage node:20-slim, pnpm --filter @duoc-thu/web build |
same |
postgres, qdrant, caddy, prometheus, tempo, grafana, otel-collector |
Upstream images | pulled |
There is no container registry. Images are built on the production host at
deploy time. The Helm chart assumes registry images
(duocthu-ai-service:<tag>) that nothing currently produces.
Scaling implications
ai-serviceis CPU-light and latency-bound on Bedrock. Horizontal scaling is blocked by the in-process_last_frame/_clarify_streakstate above.web's rate limiter is per-process; a second replica doubles the effective allowance.middleware.tssays so explicitly.- Qdrant and PostgreSQL are single containers with named Docker volumes on one EBS-backed host. No replication, no backup job in the repository.
- The
evidence_limit/max_context_tokenspolicy (rag/service.py,EvidencePolicy) bounds prompt size; nothing bounds concurrent Bedrock calls beyond the per-request budget.