9.8 KiB
9.8 KiB
25 — Troubleshooting
Every row is derived from a specific code path, comment, or observed failure in this repository. Nothing here is speculative.
Startup
| Symptom | Likely cause | How to verify | Fix |
|---|---|---|---|
ai-service exits immediately on start, ManifestMismatch in the log |
EMBEDDING_DIMENSIONS/model does not match duocthu_v1__manifest, or the sidecar collection is missing entirely |
docker logs docker-ai-service-1; then GET /collections/duocthu_v1__manifest/points/00000000-0000-5000-8000-000000000001 on Qdrant |
Point QDRANT_COLLECTION at the collection the manifest was written for, or restore/reload the corpus. Do not bypass the check |
ResponseHandlingException … connection refused at startup or during pytest collection |
EMBEDDING_PROVIDER=cohere-v4 with no reachable Qdrant. main.py builds the runtime at import time |
Try to reach QDRANT_URL |
Start Qdrant, or set EMBEDDING_PROVIDER=disabled |
ValueError: No production query embedder is configured |
EMBEDDING_PROVIDER is neither cohere-v4 nor disabled |
bootstrap.py::build_runtime |
Use one of the two supported values |
ValueError: Unknown ANSWER_PROVIDER |
Typo in ANSWER_PROVIDER |
bootstrap.py::_build_generator |
disabled | stub | bedrock-claude | bedrock-converse |
| Startup fails reading the entities file | ENTITIES_PATH default assumes a full monorepo checkout; the container flattens apps/ai-service into /app |
Check ENTITIES_PATH in .env.prod |
Set ENTITIES_PATH=./ingestion_data/drug_entities.json (the Dockerfile bakes it there) |
Request-time
| Symptom | Likely cause | How to verify | Fix |
|---|---|---|---|
503 RAG backend is not configured |
app.state.answer_service is None — i.e. EMBEDDING_PROVIDER=disabled |
GET /ready returns 200 in this mode, so check the env, not the probe |
Configure a real embedding provider |
| Every question returns an abstain | One dominant failure upstream | duocthu_abstention_total{reason} and duocthu_generation_rejected_total{reason} |
Follow the reason code in the table below |
provider_unavailable |
Bedrock unreachable, throttled, or IAM denied | duocthu_provider_failure_total{provider,operation,reason}; ai-service logs |
Check the instance role, model access, and region |
request_budget_exhausted |
40 s wall clock or 8 calls used. Most often the completeness-repair path (observed live at 40.3 s on an Isosorbid dinitrat dosage turn) | Tempo span durations per stage | Raise MAX_WALL_CLOCK_MS, or investigate why repair triggered |
unsupported_claim |
The entailment judge did not confirm a claim against its cited block | Trace row + Tempo rag.stage.entailment |
Usually genuine; if it recurs on correct answers, inspect the evidence labelling |
incomplete_answer |
The judge found a quote-validated omission and the repair still failed | answer.py logs answer completeness repair: at WARNING with the missing items |
Inspect the evidence; the repair doubles model calls, so it may also be a budget issue |
ungrounded_number |
A figure in the answer is not verbatim in the block it cites | Grounding is deterministic — reproduce with the same evidence | Working as designed; the answer was correctly discarded |
evidence_insufficient |
The model self-reported insufficiency twice | — | Often a genuinely unanswerable question for the retrieved section |
drug_not_in_formulary |
The name is not in the 684-drug catalog, or fuzzy matching did not put it in the candidate set | GET /v1/rag/suggest?q=<prefix> |
Correct behaviour for a real absence; the corpus is Part 2 monographs only |
out_of_scope |
looks_non_human matched, or the turn is about Part 1/Part 3 content |
rag/policy.py phrase list |
Correct behaviour |
| The bot re-asks the same clarifying question | Understanding did not merge a known field | Look for clarify_loop_exhausted after four turns |
Restate the whole question in one message or start a new session; the circuit breaker says so |
clarify_loop_exhausted |
Four consecutive clarifies | duocthu_clarify_asked_total{reason} |
As above |
| An abstain reads as "no data in the formulary" but the logs show an outage | A reason code with no entry in REFUSALS fell through to GENERIC_REFUSAL |
Compare the code against the map in apps/web/app/api/chat/route.ts |
Add the missing entry — the file's comment calls this out explicitly |
| Answers come back verbatim and unpolished | No generator configured — retrieval-only mode | duocthu_answer_extractive_total is incrementing |
Set a real ANSWER_PROVIDER |
| Section answer starts mid-sentence / wrong population first | part_index ordering lost |
adapters/qdrant.py::find_by_section re-sorts; check the payload has part_index |
Reload the corpus if payloads are missing the field |
Frontend
| Symptom | Likely cause | How to verify | Fix |
|---|---|---|---|
| "Hệ thống xử lý quá 65 giây nên đã dừng yêu cầu này" | Client abort at REQUEST_TIMEOUT_MS |
The backend may still have answered — check rag_retrieval_trace for the turn |
Retry; if frequent, look at Bedrock latency |
429 rate_limited |
middleware.ts: 12/min or 120/hour per IP on /api/chat |
Retry-After, X-RateLimit-* headers |
Wait, or adjust RULES — note the limiter is per process |
| "Dịch vụ AI Service đang khởi động hoặc gặp sự cố tạm thời" | Upstream returned non-OK — reason: upstream_error |
docker logs docker-ai-service-1 |
Fix the upstream |
| "Không thể kết nối đến AI Service (…)" | fetch threw — reason: upstream_unreachable |
Check AI_SERVICE_URL / API_GATEWAY_URL |
Correct the URL or start the service |
| Each starter-question click sends two requests | React 18 Strict Mode replay in dev | Only in pnpm dev |
initialQuerySentRef already guards it; do not remove |
/api/pdf 404 with a Vietnamese message |
The PDF is not at ../../ingestion/data/raw/… relative to process.cwd() |
ls inside the web container |
The path is resolved from apps/web, so the image must contain the repo layout |
Ingestion
| Symptom | Likely cause | How to verify | Fix |
|---|---|---|---|
chunk_all requires a verified printed_page_map |
--pdf not passed to chunk |
The exception text | Pass the source PDF; the folio map is built from it |
cannot cite …: printed folio missing for physical pages [...] |
page_map could not resolve a folio (two same-size candidates) |
Render the page and look at the header band | Investigate that page; the code deliberately refuses to guess |
cannot map … chunk source text uniquely to its section |
source_text occurs zero or multiple times in the section |
Gate chunk_source_text_not_unique |
A packer or normalisation change; do not relax the check |
DuplicateDrugIdError |
Two monographs slugify to the same drug_id |
The exception names it | Disambiguate in segment/ |
CorpusMismatch: refusing to load into '…' |
The collection was built from a different corpus/model/dimension | Compare corpus_sha256 and model_id |
Load into a new collection name |
collection '…' already holds N points but has no manifest |
The collection was written by something that did not record what it wrote | — | Recreate it via the loader |
| Loader exits 1 after upserting | collection_count != points_upserted |
The printed report | Investigate before querying — the corpus is not trustworthy |
NotImplementedError: 'visual-diff' is planned… |
Declared but unbuilt CLI subcommand | cli.py::_cmd_not_implemented |
Not a bug |
UnicodeEncodeError printing Vietnamese |
Windows console codepage | — | ingestion.cli reconfigures stdout; for other scripts set PYTHONIOENCODING=utf-8 |
Observability
| Symptom | Likely cause | How to verify | Fix |
|---|---|---|---|
| No traces in Grafana | The observability overlay was not included in docker compose up |
docker ps for otel-collector/tempo; check OTEL_ENABLED |
Include both -f files |
/metrics returns 404 |
No exporter on app.state — METRICS_ENABLED=false or prometheus_client missing |
main.py returns 404 rather than an empty 200 on purpose |
Install the extra / enable the flag |
/metrics returns 401 |
METRICS_TOKEN is set |
— | Send Authorization: Bearer <token> |
duocthu_loop_* and duocthu_followup_inherited_total are always 0 |
Registered but never incremented — leftovers of the retired ADR 0007 design | grep confirms no increment call |
Expected; not a data-loss symptom |
| Trace id present in the response but absent from Tempo | Batch export delay, or the collector is down | The deploy workflow retries for 60 s for this reason | Wait, then check the collector |
duocthu_trace_write_failed_total climbing |
PostgreSQL unreachable — answers still return (fail-open) | docker logs docker-postgres-1 |
Restore the database; no answers were lost |
Deployment
| Symptom | Likely cause | How to verify | Fix |
|---|---|---|---|
| Deploy fails at the gout smoke query | The corpus or the generator is broken on the new build | The workflow dumps the last 200 ai-service log lines | Investigate before retrying; the gate is doing its job |
| Deploy fails asserting a Grafana datasource | Provisioning files changed or Grafana did not finish starting | docker logs docker-grafana-1 |
Fix provisioning under infra/docker/grafana/ |
Deploy fails at test -n "$GRAFANA_ADMIN_PASSWORD" |
The GitHub secret is unset | Repository secrets | Set it |
A change to the postgres/qdrant service definition has no effect |
They are not in the workflow's up -d list |
docker inspect the container |
Restart them manually and deliberately |
| The host checkout moved but the app did not update | The build failed after git reset --hard |
docker compose ps |
Re-run the build; there is no automatic revert |