8.8 KiB
Review: apps/ai-service/rag and the hard-10 "10/10"
Reviewer: Claude, 2026-08-03. Every number below was produced by running the code, not by reading it. Reproduction commands are given per finding.
Headline: the 10/10 reproduces, and it does not mean what it appears to mean. Four of the ten passes are bought by a term list drawn from the ten scored queries, one passes for a reason unrelated to what it tests, and the whole eval can only run against an artifact built from the same ten pages.
Baseline, reproduced:
cd apps/ai-service
python -m rag.run_eval \
--cases evals/manual_adversarial_hard10.jsonl \
--documents ../../ingestion/scratch/rag-table-pilot/out/hard10/retrieval_documents.jsonl \
--parents ../../ingestion/scratch/rag-table-pilot/out/hard10/logical_tables.jsonl
-> release_gate: {"cases": 10, "passed": 10, "pass_rate": 1.0}
pytest -q in apps/ai-service → 7 passed.
ruff check --select F,E9,B,ARG . → 2 errors (both ARG001, one in
tests/test_retrieval_service.py::FixedRetriever.search).
1. The abstain case passes by coincidence, and the same path refuses a valid question
unsupported-veterinary ("Liều famciclovir điều trị cho mèo là bao nhiêu?",
expected_id: null) is the case that is supposed to show the system refusing
an unsupported question. It abstains — with
reason: "ambiguous_top_evidence", not a scope check.
Measured: its top two hits tie at 0.675969 and 0.675969, a gap of exactly
0.000000. _is_ambiguous fires on the tie. Both scores are far above either
threshold (0.08 in run_eval, 0.12 by default), so the refusal has nothing to
do with the question being unanswerable.
Two checks that settle it:
- With
ambiguity_margin=0.0the identical query returnsdecision=verify_pdfand three pieces of evidence — the case fails. The pass rests entirely on one tie-breaking constant. - Replacing
cho mèo(for cats) withcho người lớn(for adults) — a perfectly answerable clinical question — produces the sameabstain / ambiguous_top_evidence. The word "mèo" changes nothing.
So there is no out-of-scope detection in this service, and the eval reports that there is. For a drug reference aimed at clinicians this is the worst shape of defect available: a refusal mechanism that looks validated, fires on ties rather than on scope, and will refuse real dosing questions at the same rate.
2. _structured_boost is tuned on the queries it is scored against
in_memory.QUANTITATIVE_TERMS has 13 entries. 12 of the 13 appear
literally in the 10 scored queries; only thành does not:
in queries : bao, clcr, kg, liều, lít, mg, ml, nồng, phút, thể, tích, tốc
not in them: thành
Load-bearing, measured by monkey-patching and re-running the same 10 cases:
| Configuration | Score |
|---|---|
| as shipped | 10/10 |
QUANTITATIVE_TERMS emptied |
8/10 |
_structured_boost disabled entirely |
6/10 |
Failures when the boost is removed: formula-no-printed-bar,
renal-herpes-typo, spatial-dose-formula, ors-who-composition.
Four of the ten passes come from a hand-written list whose contents overlap the test queries almost exactly. That is fitting the test set; the resulting number predicts nothing about a query written by someone else.
Stated precisely, because it matters: these files are untracked, so there is no commit history to prove the term list was written after the queries. The 12/13 overlap is strong evidence of it, not proof of the order.
3. "10/10" is Recall@3, not Recall@1
EvaluationOutcome.passed is expected_id in retrieved_ids, and
EvidencePolicy.evidence_limit is 3. Scored at Recall@1 the same run gives
9/10.
The one that moves is formula-no-printed-bar — the ADENOSIN formula printed
without a fraction bar, i.e. exactly the case the outlier catalog flags as
hardest. It lands at rank 3 of 3, behind
adenosin__lieu_luong_va_cach_dung__1 and adenosin__than_trong__0. An
answer layer handed those three in that order sees two prose sections before
the formula it actually needs.
4. The eval cannot be run on anything but the ten pages it was built from
artifacts.load_documents requires section_key, plus source_refs and
requires_visual_check. Those fields exist only in
out/hard10/retrieval_documents.jsonl:
out/all : KeyError 'section_key'
out/100 : KeyError 'section_key'
out/hard10: loads
The corpus-wide artifact — the 151-block, all-monograph one — cannot be loaded
by this code at all. The retriever's entire universe is 164 documents across
6 drug_ids, and those 6 are exactly the 6 under test (set(artifact) - set(cases) is empty). Per-query candidate pools are 19-31 documents, because
search filters on drug_id first.
CLAUDE.md is explicit that a selected-page scope must not be reported as a
whole-document one. Widening this eval requires fixing either the loader or
the artifact writer; until then no number from it generalises.
5. Two cases do not test what their names say
run_eval.run() calls service.retrieve(case.query, case.drug_id) — the
correct drug_id is handed in from the fixture.
renal-herpes-typodeliberately misspells "famciclovia", but the case carriesdrug_id: "famciclovir". Entity resolution is bypassed, so the typo never reaches the thing that would have to survive it; it only perturbs lexical scoring inside the already-correct drug.unsupported-veterinarylikewise gets the right drug handed to it.
Both are still useful as within-drug ranking cases. Neither is evidence about
name resolution, which is where docs/v1-delivery-plan.md §B4 puts the 19
measured substring traps.
6. manual_adversarial sits in the same release gate as expert
RELEASE_GATE_ORIGINS = {EXPERT, MANUAL_ADVERSARIAL}, and all ten cases are
manual_adversarial — written by the same agent that wrote the retriever.
docs/v1-delivery-plan.md §8 says this in as many words: self-written,
self-graded questions measure the author's imagination, not clinical reality.
In fairness these are routing cases (did it fetch the right block id), not
content-accuracy cases, and routing is legitimately self-checkable. The
problem is the label: bucketing them with expert and calling the result a
release gate reads as clinical validation to anyone who did not write it.
7. The eval does not exercise the policy that ships
run_eval.run() hardcodes EvidencePolicy(minimum_score=0.08, ambiguity_margin=0.01); the class defaults are 0.12 and 0.015.
I expected this to inflate the score. It does not — re-running with the
default policy also gives 10/10. Reporting that because it was checked. It
remains a smell that the benchmark and the shipped default are different
constants, especially given finding 1, where the whole result turns on
ambiguity_margin.
8. _char_ngrams is not character n-grams of the text
It builds " ".join(sorted(_terms(text))) — the unique words, alphabetised —
then takes 3-grams of that. Word adjacency is destroyed and the resulting
n-grams straddle alphabetically-neighbouring word boundaries. It still
measures some overlap, and I did not trace a specific eval failure to it,
so this is a naming/design objection rather than a demonstrated bug. But it
should not be described as character n-gram matching in any writeup.
What is genuinely good
Not everything here is a complaint, and these should survive any rework:
- The ports/adapters split is clean.
rag/ports.pyisProtocol-only and the domain imports no SDK — exactly the dependency inversionCLAUDE.mdasks for, and it is why finding 1 could be tested at all. parent_hydration_failedrefuses to answer from a table-row fragment whose parent table is missing. That is the ADR 0006 contract enforced in code, and it is the right instinct.missing_provenanceabstains when a document has nosource_refs. Also right, also load-bearing for citations.requires_visual_checkpropagates from row or parent intoVERIFY_PDF, which honours the quarantine rule rather than paraphrasing a table.
Suggested order of work
- Separate scope refusal from tie detection. A tie is not a reason to refuse; an out-of-drug or out-of-corpus question is. Right now only the first exists, and finding 1 shows it is standing in for the second.
- Make
artifacts.pyread the corpus-wide artifact, then re-run. Any number from a 6-drug universe is provisional. - Report Recall@1 and Recall@3 separately, always both.
- Move
QUANTITATIVE_TERMSout of the scorer or derive it from the corpus rather than by hand — and re-measure. A number produced with a query-derived boost list should carry that caveat wherever it is quoted. - Rename the bucket, or split
manual_adversarialout of the release gate until a pharmacist has written cases.