164 lines
8.6 KiB
Markdown
164 lines
8.6 KiB
Markdown
# 18 — Testing
|
|
|
|
## What exists
|
|
|
|
| Suite | Location | Framework | Tests |
|
|
|---|---|---|---|
|
|
| ai-service | `apps/ai-service/tests/` | pytest | 278 passed, 6 skipped |
|
|
| ingestion | `ingestion/tests/` | pytest | 277 passed, 12 skipped |
|
|
| web | — | — | **None** |
|
|
| packages | — | — | **None** |
|
|
| E2E / browser | — | — | **None** |
|
|
|
|
Total automated coverage: **555 Python tests, 0 JavaScript tests.**
|
|
|
|
## Running them
|
|
|
|
```bash
|
|
# ingestion — no external services needed
|
|
cd ingestion
|
|
python -m pytest tests -q
|
|
|
|
# ai-service — see the caveat below
|
|
cd apps/ai-service
|
|
EMBEDDING_PROVIDER=disabled python -m pytest tests -q
|
|
```
|
|
|
|
### The ai-service collection caveat
|
|
|
|
Running `python -m pytest tests -q` with the repository's own
|
|
`apps/ai-service/.env` present **fails at collection**:
|
|
|
|
```
|
|
ERROR tests/test_api.py - qdrant_client.http.exceptions.ResponseHandlingException:
|
|
[WinError 10061] No connection could be made because the target machine actively refused it
|
|
Interrupted: 1 error during collection
|
|
```
|
|
|
|
Cause: `tests/test_api.py` imports `main`, and `main.py` calls
|
|
`build_runtime(get_settings())` at module scope. With
|
|
`EMBEDDING_PROVIDER=cohere-v4` (the code default, and what `.env` sets) that
|
|
constructs a `QdrantClient` and calls `get_collections()` for the manifest
|
|
check. No unit test needs that.
|
|
|
|
`EMBEDDING_PROVIDER=disabled` makes `build_runtime` return early and the suite
|
|
passes in 2.6 s. This is a real usability defect for a new contributor: it is
|
|
documented nowhere in the repository, and the failure looks like a broken test
|
|
suite rather than a missing environment variable. Recorded in
|
|
[27-technical-debt.md](27-technical-debt.md).
|
|
|
|
Both suites are also run with no dependency install step of their own —
|
|
`pyproject.toml` declares `test = ["pytest>=7.4,<9"]` as an optional extra, and
|
|
neither project has a lockfile.
|
|
|
|
## ai-service — coverage by module
|
|
|
|
| Test module | Tests | What it exercises |
|
|
|---|---|---|
|
|
| `test_agent.py` | 43 | The routing table, clarify gates, the circuit breaker, `_synthesize_query`, interaction and condition paths |
|
|
| `test_grounded_generation.py` | 35 | Generation, the insufficiency retry, grounding integration, entailment, the completeness repair, budget exhaustion |
|
|
| `test_retrieval_service.py` | 29 | Every retrieval route, `_decide`, hydration, patient safety facets |
|
|
| `test_understanding.py` | 26 | Frame parsing, candidate bounding, the deterministic cues, prior-frame merge, fail-closed paths |
|
|
| `test_section_routing.py` | 20 | Longest-phrase-wins, no-match-means-`None`, neighbour pooling |
|
|
| `test_api.py` | 15 | Route contracts, disclaimer presence, trace fail-open, feedback errors, `/metrics` token |
|
|
| `test_qdrant_adapter.py` | 12 | Payload mapping, scroll paging, `part_index` ordering, phrase anchoring |
|
|
| `test_grounding.py` | 12 | Per-citation binding, ungrounded numbers, invalid/absent citations |
|
|
| `test_clinical_condition_flow.py` | 12 | `PatientContext`, condition normalisation, candidate assessment |
|
|
| `test_citation_and_intro.py` | 11 | Citation indexing, intro mode, `list_mode` skipping sufficiency |
|
|
| `test_bedrock_converse.py` | 9 | JSON extraction, provider-error translation, rerank |
|
|
| `test_policy.py` | 6 | Subject-scope narrowing/widening rules |
|
|
| `test_manifest.py` | 6 | `check_manifest` mismatch and missing-manifest refusal |
|
|
| `test_live_datastores.py` | 6 | **Integration — skipped unless `RUN_INTEGRATION=1`** |
|
|
| `test_budget.py` | 6 | Wall-clock and call-count exhaustion |
|
|
| `test_prompt_untrusted_input.py` | 5 | Question fencing, marker stripping |
|
|
| `test_fusion.py` | 5 | RRF — **for code with no runtime caller** |
|
|
| `test_observability.py` | 4 | Span creation, stage timing, correlation ids |
|
|
| `test_embedding_outage.py` | 4 | `QueryEmbeddingUnavailable` → abstain |
|
|
| `test_bootstrap.py` | 4 | Runtime wiring decisions |
|
|
| `test_answer_guardrails.py` | 4 | Abstain-vs-extractive rules |
|
|
| `test_section_order.py` | 3 | Book-order presentation |
|
|
| `test_rerank_overview.py` | 3 | Rerank fail-open and top-k capping |
|
|
| `test_calculators.py` | 3 | BSA — **for code with no runtime caller** |
|
|
| `test_condition_evaluation.py` | 1 | Metric summarisation |
|
|
|
|
## ingestion — coverage by module
|
|
|
|
| Test module | Tests | What it exercises |
|
|
|---|---|---|
|
|
| `test_load_qdrant.py` | 52 | Point ids, payload passthrough, record validation, manifest conflicts, batching, count gate |
|
|
| `test_segment_assembler.py` | 23 | Event classification, section assembly, quarantine, preamble, duplicate ids |
|
|
| `test_segment_atc.py` | 22 | ATC code parsing |
|
|
| `test_embed_providers.py` | 22 | Cohere/Titan/local adapters, request shapes |
|
|
| `test_chunk.py` | 22 | Packing, overlap, label carry-forward, provenance, block descriptors |
|
|
| `test_segment_merge.py` | 13 | Multi-line heading merge |
|
|
| `test_load_qdrant_integration.py` | 12 | Loader against the in-memory store |
|
|
| `test_embed_cache.py` | 12 | Content-hash cache hits/misses |
|
|
| `test_segment_detector.py` | 11 | Title/heading detection, the `HMG-CoA` and `Mã ATC:` cases |
|
|
| `test_validation_residual_ink.py` | 10 | Residual-ink classification |
|
|
| `test_validation_metrics.py` | 10 | Back-index recall/precision |
|
|
| `test_segment_vocab.py` | 9 | Section vocabulary, part dividers |
|
|
| `test_normalize.py` | 9 | Glyph substitution, text flow |
|
|
| `test_extract_glyph_order.py` | 9 | Glyph/reading-order scanning |
|
|
| `test_segment_tables.py` | 8 | Table lift-out and quarantine marking |
|
|
| `test_extract_spans.py` | 7 | Span extraction |
|
|
| `test_extract_formulas.py` | 7 | Verified formula regions |
|
|
| `test_cli.py` | 7 | Subcommand wiring, including the two `NotImplementedError` stubs |
|
|
| `test_segment_units.py` | 6 | Unit handling |
|
|
| `test_validation_readiness.py` | 4 | Gate evaluation |
|
|
| `test_segment_io.py` | 4 | JSONL round-trip |
|
|
| `test_extract_page_map.py` | 4 | Printed-folio mapping, including the RIBOFLAVIN conflict |
|
|
| `test_embed_benchmark_local.py` | 4 | Local benchmark case loading |
|
|
| `test_entities_catalog.py` | 2 | Entity catalog build (skipped without the source artifact) |
|
|
|
|
## Test categories
|
|
|
|
| Category | Present? | Where |
|
|
|---|---|---|
|
|
| Unit | Yes | The bulk of both suites |
|
|
| Integration (in-memory doubles) | Yes | `test_load_qdrant_integration.py`, `rag/in_memory.py` |
|
|
| Integration (real datastores) | Yes but **gated off** | `tests/test_live_datastores.py`, `RUN_INTEGRATION=1` |
|
|
| Contract (API shape) | Partial | `test_api.py` via `TestClient` |
|
|
| Parser regression | Yes | The `test_segment_*` / `test_extract_*` family, each pinned to a named real-document case |
|
|
| Retrieval | Yes | `test_retrieval_service.py`, `test_section_routing.py` |
|
|
| RAG behaviour | Yes | `test_agent.py`, `test_grounded_generation.py` — all with stub LLMs |
|
|
| Frontend | **No** | — |
|
|
| E2E / browser | **No** | — |
|
|
| Deployment | Partial | The smoke assertions inside `deploy.yml` ([22](22-ci-cd.md)) |
|
|
| Load / performance | **No** | — |
|
|
| Security | **No** | — |
|
|
|
|
## Test design notes worth knowing
|
|
|
|
- **No test calls a real LLM or a real AWS endpoint.** Generators are stubbed
|
|
with objects implementing the `AnswerGenerator` protocol, and stubs are told
|
|
apart by which schema they receive — `tests/test_grounded_generation.py`
|
|
explains the technique.
|
|
- `ruff` config carries a per-file ignore for `tests/*` (`ARG001`, `ARG002`)
|
|
with a written justification: test doubles implement the domain protocols, so
|
|
conformance requires full signatures even where an argument is unused.
|
|
- Skips are honest: `pytest.importorskip` for `botocore` and the OpenTelemetry
|
|
SDK, and a module-level `skipif` for the live-datastore suite. Nothing is
|
|
`xfail`-marked.
|
|
|
|
## What is not tested
|
|
|
|
- **The entire frontend** — including the 65 s timeout derivation, abort
|
|
handling, the Strict-Mode duplicate-request guard, the `REFUSALS` mapping and
|
|
the citation grouping. All of those encode real production bugs that were
|
|
fixed by hand and could silently regress.
|
|
- **`middleware.ts` rate limiting** — the sweep logic, the "do not record a
|
|
rejected request" rule, and the `X-Forwarded-For` parsing.
|
|
- **Real Qdrant/PostgreSQL behaviour** in the default run (integration is gated).
|
|
- **The Helm chart** — never rendered or linted in CI.
|
|
- **Migrations** — no test applies them or checks their result.
|
|
- **Prompt content** — no snapshot test pins `SYSTEM_PROMPT`; a rule can be
|
|
edited away without any test failing.
|
|
- **End-to-end answer quality** — that is the eval sets' job, and none of them
|
|
runs automatically ([19](19-rag-evaluation.md)).
|
|
|
|
## CI
|
|
|
|
**No test runs in CI.** `.github/workflows/deploy.yml` deploys on push to
|
|
`master` without linting, type-checking, or executing either suite. See
|
|
[22-ci-cd.md](22-ci-cd.md).
|