Files
duocthu/docs-legacy/25-troubleshooting.md
T

9.8 KiB

25 — Troubleshooting

Every row is derived from a specific code path, comment, or observed failure in this repository. Nothing here is speculative.

Startup

Symptom Likely cause How to verify Fix
ai-service exits immediately on start, ManifestMismatch in the log EMBEDDING_DIMENSIONS/model does not match duocthu_v1__manifest, or the sidecar collection is missing entirely docker logs docker-ai-service-1; then GET /collections/duocthu_v1__manifest/points/00000000-0000-5000-8000-000000000001 on Qdrant Point QDRANT_COLLECTION at the collection the manifest was written for, or restore/reload the corpus. Do not bypass the check
ResponseHandlingException … connection refused at startup or during pytest collection EMBEDDING_PROVIDER=cohere-v4 with no reachable Qdrant. main.py builds the runtime at import time Try to reach QDRANT_URL Start Qdrant, or set EMBEDDING_PROVIDER=disabled
ValueError: No production query embedder is configured EMBEDDING_PROVIDER is neither cohere-v4 nor disabled bootstrap.py::build_runtime Use one of the two supported values
ValueError: Unknown ANSWER_PROVIDER Typo in ANSWER_PROVIDER bootstrap.py::_build_generator disabled | stub | bedrock-claude | bedrock-converse
Startup fails reading the entities file ENTITIES_PATH default assumes a full monorepo checkout; the container flattens apps/ai-service into /app Check ENTITIES_PATH in .env.prod Set ENTITIES_PATH=./ingestion_data/drug_entities.json (the Dockerfile bakes it there)

Request-time

Symptom Likely cause How to verify Fix
503 RAG backend is not configured app.state.answer_service is None — i.e. EMBEDDING_PROVIDER=disabled GET /ready returns 200 in this mode, so check the env, not the probe Configure a real embedding provider
Every question returns an abstain One dominant failure upstream duocthu_abstention_total{reason} and duocthu_generation_rejected_total{reason} Follow the reason code in the table below
provider_unavailable Bedrock unreachable, throttled, or IAM denied duocthu_provider_failure_total{provider,operation,reason}; ai-service logs Check the instance role, model access, and region
request_budget_exhausted 40 s wall clock or 8 calls used. Most often the completeness-repair path (observed live at 40.3 s on an Isosorbid dinitrat dosage turn) Tempo span durations per stage Raise MAX_WALL_CLOCK_MS, or investigate why repair triggered
unsupported_claim The entailment judge did not confirm a claim against its cited block Trace row + Tempo rag.stage.entailment Usually genuine; if it recurs on correct answers, inspect the evidence labelling
incomplete_answer The judge found a quote-validated omission and the repair still failed answer.py logs answer completeness repair: at WARNING with the missing items Inspect the evidence; the repair doubles model calls, so it may also be a budget issue
ungrounded_number A figure in the answer is not verbatim in the block it cites Grounding is deterministic — reproduce with the same evidence Working as designed; the answer was correctly discarded
evidence_insufficient The model self-reported insufficiency twice Often a genuinely unanswerable question for the retrieved section
drug_not_in_formulary The name is not in the 684-drug catalog, or fuzzy matching did not put it in the candidate set GET /v1/rag/suggest?q=<prefix> Correct behaviour for a real absence; the corpus is Part 2 monographs only
out_of_scope looks_non_human matched, or the turn is about Part 1/Part 3 content rag/policy.py phrase list Correct behaviour
The bot re-asks the same clarifying question Understanding did not merge a known field Look for clarify_loop_exhausted after four turns Restate the whole question in one message or start a new session; the circuit breaker says so
clarify_loop_exhausted Four consecutive clarifies duocthu_clarify_asked_total{reason} As above
An abstain reads as "no data in the formulary" but the logs show an outage A reason code with no entry in REFUSALS fell through to GENERIC_REFUSAL Compare the code against the map in apps/web/app/api/chat/route.ts Add the missing entry — the file's comment calls this out explicitly
Answers come back verbatim and unpolished No generator configured — retrieval-only mode duocthu_answer_extractive_total is incrementing Set a real ANSWER_PROVIDER
Section answer starts mid-sentence / wrong population first part_index ordering lost adapters/qdrant.py::find_by_section re-sorts; check the payload has part_index Reload the corpus if payloads are missing the field

Frontend

Symptom Likely cause How to verify Fix
"Hệ thống xử lý quá 65 giây nên đã dừng yêu cầu này" Client abort at REQUEST_TIMEOUT_MS The backend may still have answered — check rag_retrieval_trace for the turn Retry; if frequent, look at Bedrock latency
429 rate_limited middleware.ts: 12/min or 120/hour per IP on /api/chat Retry-After, X-RateLimit-* headers Wait, or adjust RULES — note the limiter is per process
"Dịch vụ AI Service đang khởi động hoặc gặp sự cố tạm thời" Upstream returned non-OK — reason: upstream_error docker logs docker-ai-service-1 Fix the upstream
"Không thể kết nối đến AI Service (…)" fetch threw — reason: upstream_unreachable Check AI_SERVICE_URL / API_GATEWAY_URL Correct the URL or start the service
Each starter-question click sends two requests React 18 Strict Mode replay in dev Only in pnpm dev initialQuerySentRef already guards it; do not remove
/api/pdf 404 with a Vietnamese message The PDF is not at ../../ingestion/data/raw/… relative to process.cwd() ls inside the web container The path is resolved from apps/web, so the image must contain the repo layout

Ingestion

Symptom Likely cause How to verify Fix
chunk_all requires a verified printed_page_map --pdf not passed to chunk The exception text Pass the source PDF; the folio map is built from it
cannot cite …: printed folio missing for physical pages [...] page_map could not resolve a folio (two same-size candidates) Render the page and look at the header band Investigate that page; the code deliberately refuses to guess
cannot map … chunk source text uniquely to its section source_text occurs zero or multiple times in the section Gate chunk_source_text_not_unique A packer or normalisation change; do not relax the check
DuplicateDrugIdError Two monographs slugify to the same drug_id The exception names it Disambiguate in segment/
CorpusMismatch: refusing to load into '…' The collection was built from a different corpus/model/dimension Compare corpus_sha256 and model_id Load into a new collection name
collection '…' already holds N points but has no manifest The collection was written by something that did not record what it wrote Recreate it via the loader
Loader exits 1 after upserting collection_count != points_upserted The printed report Investigate before querying — the corpus is not trustworthy
NotImplementedError: 'visual-diff' is planned… Declared but unbuilt CLI subcommand cli.py::_cmd_not_implemented Not a bug
UnicodeEncodeError printing Vietnamese Windows console codepage ingestion.cli reconfigures stdout; for other scripts set PYTHONIOENCODING=utf-8

Observability

Symptom Likely cause How to verify Fix
No traces in Grafana The observability overlay was not included in docker compose up docker ps for otel-collector/tempo; check OTEL_ENABLED Include both -f files
/metrics returns 404 No exporter on app.stateMETRICS_ENABLED=false or prometheus_client missing main.py returns 404 rather than an empty 200 on purpose Install the extra / enable the flag
/metrics returns 401 METRICS_TOKEN is set Send Authorization: Bearer <token>
duocthu_loop_* and duocthu_followup_inherited_total are always 0 Registered but never incremented — leftovers of the retired ADR 0007 design grep confirms no increment call Expected; not a data-loss symptom
Trace id present in the response but absent from Tempo Batch export delay, or the collector is down The deploy workflow retries for 60 s for this reason Wait, then check the collector
duocthu_trace_write_failed_total climbing PostgreSQL unreachable — answers still return (fail-open) docker logs docker-postgres-1 Restore the database; no answers were lost

Deployment

Symptom Likely cause How to verify Fix
Deploy fails at the gout smoke query The corpus or the generator is broken on the new build The workflow dumps the last 200 ai-service log lines Investigate before retrying; the gate is doing its job
Deploy fails asserting a Grafana datasource Provisioning files changed or Grafana did not finish starting docker logs docker-grafana-1 Fix provisioning under infra/docker/grafana/
Deploy fails at test -n "$GRAFANA_ADMIN_PASSWORD" The GitHub secret is unset Repository secrets Set it
A change to the postgres/qdrant service definition has no effect They are not in the workflow's up -d list docker inspect the container Restart them manually and deliberately
The host checkout moved but the app did not update The build failed after git reset --hard docker compose ps Re-run the build; there is no automatic revert