# Review: `apps/ai-service/rag` and the hard-10 "10/10" Reviewer: Claude, 2026-08-03. Every number below was produced by running the code, not by reading it. Reproduction commands are given per finding. **Headline: the 10/10 reproduces, and it does not mean what it appears to mean.** Four of the ten passes are bought by a term list drawn from the ten scored queries, one passes for a reason unrelated to what it tests, and the whole eval can only run against an artifact built from the same ten pages. Baseline, reproduced: ``` cd apps/ai-service python -m rag.run_eval \ --cases evals/manual_adversarial_hard10.jsonl \ --documents ../../ingestion/scratch/rag-table-pilot/out/hard10/retrieval_documents.jsonl \ --parents ../../ingestion/scratch/rag-table-pilot/out/hard10/logical_tables.jsonl -> release_gate: {"cases": 10, "passed": 10, "pass_rate": 1.0} ``` `pytest -q` in `apps/ai-service` → **7 passed**. `ruff check --select F,E9,B,ARG .` → **2 errors** (both ARG001, one in `tests/test_retrieval_service.py::FixedRetriever.search`). --- ## 1. The abstain case passes by coincidence, and the same path refuses a valid question `unsupported-veterinary` ("Liều famciclovir điều trị cho mèo là bao nhiêu?", `expected_id: null`) is the case that is supposed to show the system refusing an unsupported question. It abstains — with `reason: "ambiguous_top_evidence"`, not a scope check. Measured: its top two hits tie at **0.675969 and 0.675969, a gap of exactly 0.000000**. `_is_ambiguous` fires on the tie. Both scores are far above either threshold (0.08 in `run_eval`, 0.12 by default), so the refusal has nothing to do with the question being unanswerable. Two checks that settle it: - With `ambiguity_margin=0.0` the identical query returns `decision=verify_pdf` and three pieces of evidence — the case **fails**. The pass rests entirely on one tie-breaking constant. - Replacing `cho mèo` (for cats) with `cho người lớn` (for adults) — a perfectly answerable clinical question — produces the **same** `abstain / ambiguous_top_evidence`. The word "mèo" changes nothing. So there is no out-of-scope detection in this service, and the eval reports that there is. For a drug reference aimed at clinicians this is the worst shape of defect available: a refusal mechanism that looks validated, fires on ties rather than on scope, and will refuse real dosing questions at the same rate. ## 2. `_structured_boost` is tuned on the queries it is scored against `in_memory.QUANTITATIVE_TERMS` has 13 entries. **12 of the 13 appear literally in the 10 scored queries**; only `thành` does not: ``` in queries : bao, clcr, kg, liều, lít, mg, ml, nồng, phút, thể, tích, tốc not in them: thành ``` Load-bearing, measured by monkey-patching and re-running the same 10 cases: | Configuration | Score | |---|---| | as shipped | 10/10 | | `QUANTITATIVE_TERMS` emptied | **8/10** | | `_structured_boost` disabled entirely | **6/10** | Failures when the boost is removed: `formula-no-printed-bar`, `renal-herpes-typo`, `spatial-dose-formula`, `ors-who-composition`. Four of the ten passes come from a hand-written list whose contents overlap the test queries almost exactly. That is fitting the test set; the resulting number predicts nothing about a query written by someone else. Stated precisely, because it matters: these files are untracked, so there is no commit history to prove the term list was written *after* the queries. The 12/13 overlap is strong evidence of it, not proof of the order. ## 3. "10/10" is Recall@3, not Recall@1 `EvaluationOutcome.passed` is `expected_id in retrieved_ids`, and `EvidencePolicy.evidence_limit` is 3. Scored at Recall@1 the same run gives **9/10**. The one that moves is `formula-no-printed-bar` — the ADENOSIN formula printed without a fraction bar, i.e. exactly the case the outlier catalog flags as hardest. It lands at **rank 3 of 3**, behind `adenosin__lieu_luong_va_cach_dung__1` and `adenosin__than_trong__0`. An answer layer handed those three in that order sees two prose sections before the formula it actually needs. ## 4. The eval cannot be run on anything but the ten pages it was built from `artifacts.load_documents` requires `section_key`, plus `source_refs` and `requires_visual_check`. Those fields exist **only** in `out/hard10/retrieval_documents.jsonl`: ``` out/all : KeyError 'section_key' out/100 : KeyError 'section_key' out/hard10: loads ``` The corpus-wide artifact — the 151-block, all-monograph one — cannot be loaded by this code at all. The retriever's entire universe is 164 documents across **6 drug_ids**, and those 6 are exactly the 6 under test (`set(artifact) - set(cases)` is empty). Per-query candidate pools are 19-31 documents, because `search` filters on `drug_id` first. `CLAUDE.md` is explicit that a selected-page scope must not be reported as a whole-document one. Widening this eval requires fixing either the loader or the artifact writer; until then no number from it generalises. ## 5. Two cases do not test what their names say `run_eval.run()` calls `service.retrieve(case.query, case.drug_id)` — the correct `drug_id` is handed in from the fixture. - `renal-herpes-typo` deliberately misspells "famciclovia", but the case carries `drug_id: "famciclovir"`. Entity resolution is bypassed, so the typo never reaches the thing that would have to survive it; it only perturbs lexical scoring *inside* the already-correct drug. - `unsupported-veterinary` likewise gets the right drug handed to it. Both are still useful as within-drug ranking cases. Neither is evidence about name resolution, which is where `docs/v1-delivery-plan.md` §B4 puts the 19 measured substring traps. ## 6. `manual_adversarial` sits in the same release gate as `expert` `RELEASE_GATE_ORIGINS = {EXPERT, MANUAL_ADVERSARIAL}`, and all ten cases are `manual_adversarial` — written by the same agent that wrote the retriever. `docs/v1-delivery-plan.md` §8 says this in as many words: self-written, self-graded questions measure the author's imagination, not clinical reality. In fairness these are *routing* cases (did it fetch the right block id), not content-accuracy cases, and routing is legitimately self-checkable. The problem is the label: bucketing them with `expert` and calling the result a release gate reads as clinical validation to anyone who did not write it. ## 7. The eval does not exercise the policy that ships `run_eval.run()` hardcodes `EvidencePolicy(minimum_score=0.08, ambiguity_margin=0.01)`; the class defaults are `0.12` and `0.015`. I expected this to inflate the score. **It does not** — re-running with the default policy also gives 10/10. Reporting that because it was checked. It remains a smell that the benchmark and the shipped default are different constants, especially given finding 1, where the whole result turns on `ambiguity_margin`. ## 8. `_char_ngrams` is not character n-grams of the text It builds `" ".join(sorted(_terms(text)))` — the unique words, alphabetised — then takes 3-grams of that. Word adjacency is destroyed and the resulting n-grams straddle alphabetically-neighbouring word boundaries. It still measures some overlap, and I did **not** trace a specific eval failure to it, so this is a naming/design objection rather than a demonstrated bug. But it should not be described as character n-gram matching in any writeup. --- ## What is genuinely good Not everything here is a complaint, and these should survive any rework: - The ports/adapters split is clean. `rag/ports.py` is `Protocol`-only and the domain imports no SDK — exactly the dependency inversion `CLAUDE.md` asks for, and it is why finding 1 could be tested at all. - `parent_hydration_failed` refuses to answer from a table-row fragment whose parent table is missing. That is the ADR 0006 contract enforced in code, and it is the right instinct. - `missing_provenance` abstains when a document has no `source_refs`. Also right, also load-bearing for citations. - `requires_visual_check` propagates from row *or* parent into `VERIFY_PDF`, which honours the quarantine rule rather than paraphrasing a table. ## Suggested order of work 1. Separate scope refusal from tie detection. A tie is not a reason to refuse; an out-of-drug or out-of-corpus question is. Right now only the first exists, and finding 1 shows it is standing in for the second. 2. Make `artifacts.py` read the corpus-wide artifact, then re-run. Any number from a 6-drug universe is provisional. 3. Report Recall@1 and Recall@3 separately, always both. 4. Move `QUANTITATIVE_TERMS` out of the scorer or derive it from the corpus rather than by hand — and re-measure. A number produced with a query-derived boost list should carry that caveat wherever it is quoted. 5. Rename the bucket, or split `manual_adversarial` out of the release gate until a pharmacist has written cases.