7.5 KiB
07 — Indexing and storage
Qdrant collections
| Collection | Points | Vector | Purpose |
|---|---|---|---|
duocthu_v1 |
15,100 | 1,024-d, Cosine | The corpus |
duocthu_v1__manifest |
1 | 1-d [0.0], never searched |
Corpus binding record |
Why a sidecar collection
Qdrant has no collection-level metadata field, so the manifest must live in a
point. Putting it inside the data collection would make count() one larger
than the chunk count — and qdrant_point_count == chunk_count is an acceptance
gate. ingestion/load/manifest.py states the reasoning:
A gate that needs an "except the manifest" footnote is a gate that will eventually be read wrong.
Manifest point id is the fixed UUID 00000000-0000-5000-8000-000000000001,
defined identically in ingestion/load/manifest.py and
apps/ai-service/rag/manifest.py.
Manifest payload
| Field | Example | Compared at |
|---|---|---|
corpus_sha256 |
sha256 of the whole chunks.jsonl |
load time |
chunk_count |
15100 | load time |
model_id |
cohere.embed-v4:0 |
load time and startup |
dimensions |
1024 | load time and startup |
input_kind |
search_document |
load time |
provider, distance |
cohere-v4, Cosine |
load time |
Two independent checks use it:
- Load time —
assert_compatible()raisesCorpusMismatchon any conflict, before creating or writing anything, so a refused load leaves the store untouched. A data collection that already holds points but has no manifest is itself a refusal. - Startup —
bootstrap.py::_verify_corpus_manifestreads the sidecar and callsrag/manifest.py::check_manifest, comparingmodel_idanddimensionsagainst the configured query embedder. A mismatch — or a missing manifest — raisesManifestMismatch, which crashes the process at import time, so the service never serves a query against an unattested corpus.
The failure this prevents is silent: two embedding models can produce vectors of the same dimensionality, and Qdrant returns plausible nearest neighbours with no error.
Point ids
POINT_NAMESPACE = uuid.UUID("6f0d6d1e-4c2a-5f6b-9a3d-2f8e1c7b4a90")
point_id_for(chunk_id) = str(uuid.uuid5(POINT_NAMESPACE, chunk_id))
Derived, never random, so a re-load converges instead of doubling. The namespace is described in-code as "a constant of the project, not a tunable" — changing it re-ids the whole corpus and orphans every loaded point.
Consequence documented in adapters/qdrant.py: because ids are UUIDs, Qdrant's
natural scroll order (point-id order) is effectively random. find_by_section
therefore re-sorts by part_index before returning — PARACETAMOL's dosing
section came back 3, 4, 1, 2, 0, opening mid-sentence on paediatric doses. A
section served out of order is a clinical hazard, not a formatting one.
Payload
The whole chunk record passes through intact — build_point does
payload=dict(record) with no whitelist. ingestion/load/models.py explains
why: a whitelist would silently drop any field a later chunker adds.
Indexed payload fields
CollectionSpec.indexed_fields, created once at collection creation:
| Field | Schema | Used by |
|---|---|---|
chunk_id |
keyword | QdrantParentStore.get |
drug_id |
keyword | every retrieval route |
section_key |
keyword | find_by_section, find_by_indication, search_indication, search_lexical |
atc_codes |
keyword | no runtime query filters on it today |
chunk_kind |
keyword | find_by_drug, find_by_indication, search_indication |
has_quarantined_content |
bool | no runtime query filters on it today; it is read off the payload instead |
text is not in INDEXED_PAYLOAD_FIELDS, yet search_lexical issues
MatchText conditions against it. Qdrant requires an explicit full-text index
for MatchText; without one the condition does not match as intended. This is
recorded in 27-technical-debt.md — the lexical route may
be relying on the post-filter re-scoring in Python (matched = sum(1 for t in tokens if t in text_normalized.split())) rather than on the index.
Loading
ChunkLoader.load() (ingestion/load/upsert.py), in a fixed order:
assert_compatible()— corpus binding gate, before any write.- Create the collection + payload indexes if absent.
- Write the manifest.
- Validate each record (
validate_chunk_record) and each vector's length againstspec.vector_size— a wrong-sized vector is a whole-run defect, and failing on the first is cheaper than discovering it after 15,000 upserts. - Upsert in batches of 256 with
wait=True. - Report
collection_countvspoints_upserted;run.pyexits non-zero on mismatch.
assert_point_count(expected_chunks) exists as the stricter v1 gate but
run.py does not call it — it compares against points_upserted instead.
PostgreSQL schema
Four migrations, applied in sorted filename order by python -m migrate
(apps/ai-service/migrate.py). All are IF NOT EXISTS, so re-running is safe.
erDiagram
rag_retrieval_trace ||--o| rag_answer_feedback : "trace_id FK, ON DELETE CASCADE"
rag_conversation_turn }o..o{ rag_retrieval_trace : "conversation_id, no FK"
rag_retrieval_trace {
uuid trace_id PK
text query_text
text subject_scope
text query_intent
text decision
text reason
text resolved_drug_id
jsonb citations
text correlation_id
varchar32 otel_trace_id
timestamptz created_at
}
rag_conversation_turn {
bigserial id PK
text conversation_id
text line
timestamptz created_at
}
rag_answer_feedback {
uuid feedback_id PK
uuid trace_id FK "UNIQUE"
varchar128 conversation_id
varchar16 rating "helpful|not_helpful"
text comment "<=2000 chars"
timestamptz created_at
timestamptz updated_at
}
Indexes: rag_retrieval_trace (created_at DESC); partial indexes on
correlation_id and otel_trace_id where not null;
rag_conversation_turn (conversation_id, id);
rag_answer_feedback (created_at DESC).
Notes:
rag_conversation_turnis append-only. There is no retention or deletion path anywhere in the repository — every user turn accumulates forever. See 16-security.md.subject_scopeandquery_intenton the trace are the server-resolved values, not the caller's claim (routers/rag.pycomment).- Access is
psycopgwith a new connection per call and no pool, withconnect_timeout=5. The timeout matters: an unreachable-but-not-refusing host otherwise hangs on the OS TCP timeout, defeating the caller's fail-opentry/except.
Other storage
| Location | Contents | Lifecycle |
|---|---|---|
Docker volume postgres-data |
PostgreSQL data | Host-local, no backup job in repo |
Docker volume qdrant-data |
Qdrant storage | Host-local, no backup job in repo |
Docker volumes caddy-data, caddy-config |
ACME certs | Managed by Caddy |
Docker volumes prometheus-data, tempo-data, grafana-data |
Observability | Retention configured in Helm values only (7d / 24h); the Compose overlay sets no retention flags |
ingestion/data/processed/embeddings/*.jsonl |
Embedding cache keyed by (model_id, input_kind, sha256(text)) |
Local disk, reused across runs |
To move the corpus between machines, ingestion/README.md instructs snapshot +
restore of the Qdrant collection rather than re-embedding — it is free and
exact, whereas re-embedding costs real Bedrock spend.