ai-service
FastAPI service holding the entire live RAG path: query understanding, drug
resolution, deterministic section routing, guarded retrieval, grounded
generation, claim/citation verification, printed-page citations, and
PostgreSQL conversation + retrieval traces. In production this is what
apps/web calls — there is no gateway in front of it.
Local infrastructure
docker compose -f ..\..\infra\docker\docker-compose.yml up -d postgres qdrant
python -m migrate
python -m uvicorn main:app --port 8079
Use a plain restart rather than
--reloadon Windows. The reloader has been observed serving the previous code after an edit on this project, so a change can appear to have no effect (or appear to work when it has not been loaded). Restarting the process avoids the ambiguity.
GET /health is always available. POST /v1/rag/query is the live agent
endpoint. Structured subject_scope/intent requests fail closed on
unknown/non-human input.
Bedrock is required for real answers
Since 2026-08-05 the live path calls AWS Bedrock and needs credentials with invoke permission:
| Role | Default | Setting |
|---|---|---|
| Query embedding | cohere.embed-v4:0 |
embedding_provider (default cohere-v4) |
| Answer generation | Bedrock Converse | generation_provider (bedrock-converse / bedrock-claude) |
| Rerank | Cohere rerank | rerank_enabled — off by default; only affects the similarity fallback, never the deterministic section route |
Credentials come from the environment (an IAM instance role in production, no long-lived keys). With no working generator configured, a rejected generation abstains — it must never degrade into dumping raw source text at a clinician.
EMBEDDING_PROVIDER=local-smoke is only for local plumbing checks. Its
hashing vectors are deterministic but not semantic and must not be used for
any retrieval-quality claim.
Corpus coupling
Startup verifies the duocthu_v1__manifest sidecar and refuses to run
against a collection whose corpus sha / model / dimensions do not match
(F-05). A machine with an empty Qdrant cannot serve answers until a snapshot
is restored or ingestion/ is re-run — the latter costs real Bedrock spend.
Request budget and timeouts
A turn makes several sequential Bedrock calls, bounded by
max_wall_clock_ms (40s) and max_llm_calls_per_turn (8) in config.py.
The budget is checked between calls and cannot cancel one already in
flight, which is separately bounded by read_timeout in
adapters/bedrock_converse.py. Real ceiling ≈ 40s + one in-flight call, so
any client calling this service must allow more than that — apps/web's
ChatPanel.tsx uses 65s for exactly this reason.
Tests
python -m pytest -q # needs Postgres + Qdrant up
python -m pytest -q --ignore=tests/test_api.py --ignore=tests/test_live_datastores.py
230 pass without live datastores. Coverage is uneven and worth checking
before relying on it: rag/agent.py's dosing state machine has been covered
only since 2026-08-11, and the repo has no frontend test setup, so a green
suite does not by itself confirm a user-visible change works.