Files
duocthu/docs-legacy/02-system-architecture.md
T

186 lines
8.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 02 — System architecture
## Architectural style
**As built:** a two-service application (`web` + `ai-service`) plus an offline
batch pipeline, deployed as Docker Compose services on one host. Communication
is synchronous HTTP/JSON. There is no message broker, no queue, no async
worker, and no service mesh.
**As designed on paper:** a seven-service microservices platform
(`api-gateway`, `auth-service`, `user-service`, `chat-service`, `ai-service`,
`web`, `ingestion`), described in the pre-existing `docs/architecture.md`. Four
of those seven do not exist — their directories hold a `README.md` and a
four-line `package.json` with no `dependencies` and no source files. The
monorepo scaffolding (pnpm workspace entries, `infra/k8s/base/<service>/`
directories) still reserves their names.
Both facts matter: the second explains why `apps/`, `pnpm-workspace.yaml` and
the Helm chart look bigger than the running system.
## Component diagram — what actually runs
```mermaid
flowchart TB
subgraph browser["Browser"]
UI["Chat UI<br/>ChatPanel.tsx · 65s abort"]
end
subgraph ec2["EC2 host — docker compose"]
CADDY["caddy:2-alpine<br/>:80 :443 · ACME TLS"]
subgraph webc["web (Next.js 14, :3000)"]
MW["middleware.ts<br/>in-memory IP rate limit"]
BFF["/api/chat · /api/suggest<br/>/api/feedback · /api/pdf"]
end
subgraph aic["ai-service (FastAPI, :8000)"]
HTTP["routers/rag.py"]
AGENT["RagAgent"]
RET["RetrievalService"]
ANS["GroundedAnswerService"]
end
PG[("postgres:16-alpine")]
QD[("qdrant/qdrant")]
subgraph obs["observability overlay"]
OTELC["otel-collector"]
TEMPO["tempo"]
PROM["prometheus"]
GRAF["grafana :3002 (127.0.0.1)"]
end
end
BEDROCK["AWS Bedrock<br/>embed-v4 · rerank-v3.5 · Converse"]
UI -->|HTTPS| CADDY
CADDY -->|"/*"| MW --> BFF
CADDY -->|"/grafana/*"| GRAF
BFF -->|"POST /v1/rag/query"| HTTP
HTTP --> AGENT
AGENT --> RET
AGENT --> ANS
RET --> QD
ANS -->|generation + entailment| BEDROCK
AGENT -->|understanding| BEDROCK
RET -->|embed + rerank| BEDROCK
HTTP --> PG
AGENT --> PG
aic -->|OTLP/HTTP| OTELC --> TEMPO
PROM -->|scrape /metrics| aic
GRAF --> PROM
GRAF --> TEMPO
```
`ai-service` publishes **no host port** in `docker-compose.prod.yml`; it is
reachable only on the Compose network. Caddy proxies `web` and `grafana` only.
## Service boundaries
| Service | Owns | Depends on | Stateless? |
|---|---|---|---|
| `web` | Rendering, reason-code → Vietnamese message mapping, citation grouping, rate limiting | `ai-service` over HTTP; the source PDF on a bind-mounted path | **No** — rate-limit counters are per-process in memory |
| `ai-service` | Understanding, retrieval, generation, grounding, citations, traces | Qdrant, PostgreSQL, Bedrock | **Mostly**`RagAgent` keeps two in-process dicts (`_last_frame`, `_clarify_streak`); conversation *text* is in PostgreSQL |
| `ingestion` | Turning the PDF into `chunks.jsonl` and Qdrant points | Qdrant, Bedrock, local disk | N/A — batch |
### The stateful detail that constrains scaling
`RagAgent` (`rag/agent.py`) holds three dicts:
```python
self._history: dict[str, list[str]] # unused when a ConversationStore is configured
self._last_frame: dict[str, QueryFrame] # ALWAYS in-process
self._clarify_streak: dict[str, int] # ALWAYS in-process
```
`PostgresConversationStore` replaces `_history` only. `_last_frame` carries the
structured merge that stops the model re-asking an answered clarify question,
and `_clarify_streak` drives the clarify circuit breaker. Both are lost on
restart and **not shared between replicas**. Running more than one `ai-service`
replica therefore degrades multi-turn quality in a way nothing detects. The Helm
chart's `aiService.replicaCount` defaults to `1`; nothing enforces it.
## Module boundaries inside `ai-service`
Ports-and-adapters, enforced by import discipline rather than by tooling:
- `rag/ports.py` declares `Retriever`, `SectionRetriever`, `ParentStore`,
`AnswerGenerator`, `Reranker`, plus three exception types
(`QueryEmbeddingUnavailable`, `AnswerGenerationUnavailable`,
`RerankUnavailable`) that adapters raise and the domain catches.
- `adapters/` is the only place `qdrant_client`, `boto3`, `psycopg` and
`prometheus_client` are imported — and always **lazily**, inside a method, so
the domain imports cleanly on a machine with none of them installed.
- `bootstrap.py` is the composition root. Nothing else constructs an adapter.
One boundary is looser than the protocol suggests: `RetrievalService` reaches
optional retriever capabilities with `getattr(self._retriever, "find_by_section",
None)` rather than through a declared protocol. `find_by_indication`,
`search_indication`, `search_lexical` and `find_by_drug` are all discovered this
way and none of them appear in `ports.py`. A retriever missing one silently
disables a whole route instead of failing a type check.
## Synchronous communication
Every hop is a blocking HTTP or SDK call. One answerable turn issues, in
sequence:
1. `POST /api/chat` (browser → web)
2. `POST /v1/rag/query` (web → ai-service)
3. Bedrock Converse — understanding
4. Qdrant `scroll`/`query_points` — retrieval (1N calls)
5. Bedrock Converse — generation (plus one retry on `evidence_sufficient=false`)
6. Bedrock Converse — entailment
7. optional Bedrock Converse ×2 — completeness repair + its re-verification
8. PostgreSQL insert — trace
Measured production latencies recorded in `ChatPanel.tsx` (n=8, 2026-08-11):
6.2 / 6.4 / 8.4 / 10.9 / 12.4 / 21.7 / 25.1 / 40.3 seconds.
## Asynchronous communication
**Not found.** No broker, no queue, no background worker, no SSE, no
WebSocket, no streaming response. `web`'s `/api/chat` awaits the full upstream
response before replying.
## Failure boundaries
| Boundary | Policy | Implemented in |
|---|---|---|
| Corpus/model manifest mismatch at startup | **Fail closed, crash the process** | `bootstrap.py::_verify_corpus_manifest``rag/manifest.py` |
| Query embedder unreachable | **Fail closed** — abstain, never a 500 | `rag/service.py`, `rag/ports.py` |
| Generator unreachable / malformed / budget exhausted | **Fail closed** — abstain with a specific reason code | `rag/answer.py::_generate` |
| Understanding call fails | **Fail closed** — abstain, tagged `system_error` | `rag/understanding.py::understand` |
| Grounding or entailment rejects | **Fail closed** — abstain | `rag/answer.py` |
| Reranker unreachable | **Fail open** — keep original order | `rag/service.py::_rerank` |
| Sufficiency check unreachable | **Fail open** — proceed to generate | `rag/answer.py::_check_sufficiency` |
| PostgreSQL trace write fails | **Fail open** — answer returned, `TRACE_WRITE_FAILED` counter | `routers/rag.py` |
| Conversation store read/write fails | **Fail open** — this turn has no memory | `rag/agent.py::_get_history` / `_remember` |
| Metrics package missing | **Degrade**`NullMetrics` | `bootstrap.py::_build_metrics` |
| OpenTelemetry packages missing / `OTEL_ENABLED=false` | **Degrade** — no-op tracer | `rag/telemetry.py` |
The asymmetry is deliberate and documented in-code: anything that could change
*what is stated* fails closed; anything that only affects quality or
observability fails open.
## Deployment units
| Unit | Image | Built by |
|---|---|---|
| `ai-service` | `apps/ai-service/Dockerfile``python:3.12-slim`, deps pinned inline (not from `pyproject.toml`) | `docker compose up --build` on the EC2 host |
| `web` | `apps/web/Dockerfile` — 3-stage node:20-slim, `pnpm --filter @duoc-thu/web build` | same |
| `postgres`, `qdrant`, `caddy`, `prometheus`, `tempo`, `grafana`, `otel-collector` | Upstream images | pulled |
There is **no container registry**. Images are built on the production host at
deploy time. The Helm chart assumes registry images
(`duocthu-ai-service:<tag>`) that nothing currently produces.
## Scaling implications
- `ai-service` is CPU-light and latency-bound on Bedrock. Horizontal scaling is
blocked by the in-process `_last_frame`/`_clarify_streak` state above.
- `web`'s rate limiter is per-process; a second replica doubles the effective
allowance. `middleware.ts` says so explicitly.
- Qdrant and PostgreSQL are single containers with named Docker volumes on one
EBS-backed host. No replication, no backup job in the repository.
- The `evidence_limit`/`max_context_tokens` policy (`rag/service.py`,
`EvidencePolicy`) bounds prompt size; nothing bounds concurrent Bedrock calls
beyond the per-request budget.