Files
duocthu/docs-legacy/07-indexing-and-storage.md
T

7.5 KiB

07 — Indexing and storage

Qdrant collections

Collection Points Vector Purpose
duocthu_v1 15,100 1,024-d, Cosine The corpus
duocthu_v1__manifest 1 1-d [0.0], never searched Corpus binding record

Why a sidecar collection

Qdrant has no collection-level metadata field, so the manifest must live in a point. Putting it inside the data collection would make count() one larger than the chunk count — and qdrant_point_count == chunk_count is an acceptance gate. ingestion/load/manifest.py states the reasoning:

A gate that needs an "except the manifest" footnote is a gate that will eventually be read wrong.

Manifest point id is the fixed UUID 00000000-0000-5000-8000-000000000001, defined identically in ingestion/load/manifest.py and apps/ai-service/rag/manifest.py.

Manifest payload

Field Example Compared at
corpus_sha256 sha256 of the whole chunks.jsonl load time
chunk_count 15100 load time
model_id cohere.embed-v4:0 load time and startup
dimensions 1024 load time and startup
input_kind search_document load time
provider, distance cohere-v4, Cosine load time

Two independent checks use it:

  • Load timeassert_compatible() raises CorpusMismatch on any conflict, before creating or writing anything, so a refused load leaves the store untouched. A data collection that already holds points but has no manifest is itself a refusal.
  • Startupbootstrap.py::_verify_corpus_manifest reads the sidecar and calls rag/manifest.py::check_manifest, comparing model_id and dimensions against the configured query embedder. A mismatch — or a missing manifest — raises ManifestMismatch, which crashes the process at import time, so the service never serves a query against an unattested corpus.

The failure this prevents is silent: two embedding models can produce vectors of the same dimensionality, and Qdrant returns plausible nearest neighbours with no error.

Point ids

POINT_NAMESPACE = uuid.UUID("6f0d6d1e-4c2a-5f6b-9a3d-2f8e1c7b4a90")
point_id_for(chunk_id) = str(uuid.uuid5(POINT_NAMESPACE, chunk_id))

Derived, never random, so a re-load converges instead of doubling. The namespace is described in-code as "a constant of the project, not a tunable" — changing it re-ids the whole corpus and orphans every loaded point.

Consequence documented in adapters/qdrant.py: because ids are UUIDs, Qdrant's natural scroll order (point-id order) is effectively random. find_by_section therefore re-sorts by part_index before returning — PARACETAMOL's dosing section came back 3, 4, 1, 2, 0, opening mid-sentence on paediatric doses. A section served out of order is a clinical hazard, not a formatting one.

Payload

The whole chunk record passes through intact — build_point does payload=dict(record) with no whitelist. ingestion/load/models.py explains why: a whitelist would silently drop any field a later chunker adds.

Indexed payload fields

CollectionSpec.indexed_fields, created once at collection creation:

Field Schema Used by
chunk_id keyword QdrantParentStore.get
drug_id keyword every retrieval route
section_key keyword find_by_section, find_by_indication, search_indication, search_lexical
atc_codes keyword no runtime query filters on it today
chunk_kind keyword find_by_drug, find_by_indication, search_indication
has_quarantined_content bool no runtime query filters on it today; it is read off the payload instead

text is not in INDEXED_PAYLOAD_FIELDS, yet search_lexical issues MatchText conditions against it. Qdrant requires an explicit full-text index for MatchText; without one the condition does not match as intended. This is recorded in 27-technical-debt.md — the lexical route may be relying on the post-filter re-scoring in Python (matched = sum(1 for t in tokens if t in text_normalized.split())) rather than on the index.

Loading

ChunkLoader.load() (ingestion/load/upsert.py), in a fixed order:

  1. assert_compatible() — corpus binding gate, before any write.
  2. Create the collection + payload indexes if absent.
  3. Write the manifest.
  4. Validate each record (validate_chunk_record) and each vector's length against spec.vector_size — a wrong-sized vector is a whole-run defect, and failing on the first is cheaper than discovering it after 15,000 upserts.
  5. Upsert in batches of 256 with wait=True.
  6. Report collection_count vs points_upserted; run.py exits non-zero on mismatch.

assert_point_count(expected_chunks) exists as the stricter v1 gate but run.py does not call it — it compares against points_upserted instead.

PostgreSQL schema

Four migrations, applied in sorted filename order by python -m migrate (apps/ai-service/migrate.py). All are IF NOT EXISTS, so re-running is safe.

erDiagram
    rag_retrieval_trace ||--o| rag_answer_feedback : "trace_id FK, ON DELETE CASCADE"
    rag_conversation_turn }o..o{ rag_retrieval_trace : "conversation_id, no FK"

    rag_retrieval_trace {
        uuid trace_id PK
        text query_text
        text subject_scope
        text query_intent
        text decision
        text reason
        text resolved_drug_id
        jsonb citations
        text correlation_id
        varchar32 otel_trace_id
        timestamptz created_at
    }
    rag_conversation_turn {
        bigserial id PK
        text conversation_id
        text line
        timestamptz created_at
    }
    rag_answer_feedback {
        uuid feedback_id PK
        uuid trace_id FK "UNIQUE"
        varchar128 conversation_id
        varchar16 rating "helpful|not_helpful"
        text comment "<=2000 chars"
        timestamptz created_at
        timestamptz updated_at
    }

Indexes: rag_retrieval_trace (created_at DESC); partial indexes on correlation_id and otel_trace_id where not null; rag_conversation_turn (conversation_id, id); rag_answer_feedback (created_at DESC).

Notes:

  • rag_conversation_turn is append-only. There is no retention or deletion path anywhere in the repository — every user turn accumulates forever. See 16-security.md.
  • subject_scope and query_intent on the trace are the server-resolved values, not the caller's claim (routers/rag.py comment).
  • Access is psycopg with a new connection per call and no pool, with connect_timeout=5. The timeout matters: an unreachable-but-not-refusing host otherwise hangs on the OS TCP timeout, defeating the caller's fail-open try/except.

Other storage

Location Contents Lifecycle
Docker volume postgres-data PostgreSQL data Host-local, no backup job in repo
Docker volume qdrant-data Qdrant storage Host-local, no backup job in repo
Docker volumes caddy-data, caddy-config ACME certs Managed by Caddy
Docker volumes prometheus-data, tempo-data, grafana-data Observability Retention configured in Helm values only (7d / 24h); the Compose overlay sets no retention flags
ingestion/data/processed/embeddings/*.jsonl Embedding cache keyed by (model_id, input_kind, sha256(text)) Local disk, reused across runs

To move the corpus between machines, ingestion/README.md instructs snapshot + restore of the Qdrant collection rather than re-embedding — it is free and exact, whereas re-embedding costs real Bedrock spend.