Fix migration workflow: upload as artifact instead of scp to practice EC2
This commit is contained in:
@@ -0,0 +1,140 @@
|
||||
# 19 — RAG evaluation
|
||||
|
||||
## What exists
|
||||
|
||||
| Asset | Location | Size | Runner |
|
||||
|---|---|---|---|
|
||||
| Retrieval eval harness | `rag/run_eval.py` | — | Manual CLI; uses in-memory stores, **not** Qdrant |
|
||||
| Retrieval eval types | `rag/evaluation.py` | — | — |
|
||||
| Condition→drug metrics | `rag/condition_evaluation.py` | — | **No runner** — only `tests/test_condition_evaluation.py` |
|
||||
| Adversarial hard set | `evals/manual_adversarial_hard10.jsonl` | 10 cases | No automated runner |
|
||||
| Condition→drug set | `evals/condition_to_drug_v1.jsonl` | 20 cases | No automated runner |
|
||||
| Production battery | `evals/production_manual_60.jsonl` | 60 cases | `scripts/run_manual_battery.py` (live HTTP) |
|
||||
| Golden datasets | `Golden Dataset/*.csv` | 5 files, 209 rows | **No runner anywhere** |
|
||||
| Deploy smoke assertion | `.github/workflows/deploy.yml` | 1 case | Runs on every deploy |
|
||||
|
||||
## The datasets
|
||||
|
||||
### `Golden Dataset/` — hand-labelled, Vietnamese, unwired
|
||||
|
||||
| File | Rows | Columns |
|
||||
|---|---|---|
|
||||
| `golden_intent_v1.csv` | 73 | question, correct intent, labelling rationale, group, difficulty |
|
||||
| `golden_entity_v1.csv` | 50 | question, correct drug, correct attribute, disease, symptom |
|
||||
| `golden_e2e_v1.csv` | 36 | scenario, question, expected intent/drug/attribute, **required content**, expected citation, pass criteria, actual-result column |
|
||||
| `golden_summary_v1.csv` | 32 | drug, attribute, source page, **verbatim source text**, meanings that must be preserved, numbers that must be copied exactly, max length, faithfulness / coverage / readability scores 0–2 |
|
||||
| `golden_multiturn_v1.csv` | 19 | conversation id, turn, question, expected behaviour, expected drug/section/population, what should be inherited |
|
||||
|
||||
These are genuinely useful — `golden_summary_v1.csv` carries the exact source
|
||||
paragraph and the exact numbers that must survive, which is precisely the
|
||||
property `grounding.verify` enforces. But **no code in the repository reads
|
||||
them**. The scoring columns are blank, i.e. filled in by hand.
|
||||
|
||||
### `evals/production_manual_60.jsonl`
|
||||
|
||||
The most structured set. Each case declares observable invariants:
|
||||
|
||||
```json
|
||||
{"id":"G01","category":"general_condition","query":"Tăng huyết áp dùng thuốc gì?",
|
||||
"decision":"answerable","condition_mode":"general",
|
||||
"expected_any_drug_ids":["methyldopa","quinapril","labetalol_hydroclorid"],
|
||||
"must_have_citations":true,"max_drugs":8}
|
||||
```
|
||||
|
||||
`scripts/run_manual_battery.py` is deliberately **a transparent HTTP recorder,
|
||||
not an LLM judge** — it posts each case to a running service and writes every
|
||||
full response to JSONL for human review against the rendered PDF pages. The
|
||||
docstring states the rationale: each case has observable invariants (decision,
|
||||
relation/section, candidate bound, citations, drug provenance), so an exact
|
||||
comparison is auditable in a way a judge model is not.
|
||||
|
||||
### `evals/condition_to_drug_v1.jsonl`
|
||||
|
||||
20 cases with `expected_intent`, `expected_condition`, `expected_relation`,
|
||||
`expected_clarification` — designed for `condition_evaluation.py`'s metrics.
|
||||
|
||||
### `evals/manual_adversarial_hard10.jsonl`
|
||||
|
||||
10 hard cases, each targeting a known parsing hazard — cross-page contrast
|
||||
dosing, a formula with no printed fraction bar — with `expected_drug_id` and an
|
||||
`expected_id` pointing at a specific block.
|
||||
|
||||
## Metrics the code can compute
|
||||
|
||||
`rag/condition_evaluation.py::summarize_condition_outcomes` is fully implemented
|
||||
and deterministic — no judge model:
|
||||
|
||||
| Metric | Definition |
|
||||
|---|---|
|
||||
| `intent_accuracy` | exact match on turn type |
|
||||
| `condition_normalization_accuracy` | exact match on the normalised condition |
|
||||
| `ambiguity_clarification_accuracy` | did it clarify exactly when it should |
|
||||
| `indication_recall_at_8` | any expected drug in the top-8 retrieved |
|
||||
| `drug_precision_at_8` | expected ∩ retrieved / retrieved |
|
||||
| `section_correctness` | every retrieved section is `chi_dinh` |
|
||||
| `relation_correctness` | indication vs adverse-effect vs contraindication |
|
||||
| `unsupported_drug_rate` | generated drugs not present in retrieval |
|
||||
| `citation_correctness` | mean over per-citation validity flags |
|
||||
| `groundedness` | mean over per-claim grounded flags |
|
||||
| `patient_context_extraction_accuracy` | field-by-field match on `PatientContext` |
|
||||
| `safety_evidence_retrieval_accuracy` | expected safety facets actually retrieved |
|
||||
|
||||
`rag/evaluation.py::summarize` covers retrieval-only outcomes (drug resolution
|
||||
status and retrieved-id match).
|
||||
|
||||
**Neither summariser has a production runner.** `run_eval.py` uses
|
||||
`InMemoryLexicalRetriever` over JSONL artifacts, so it measures the resolver and
|
||||
the section router — not the deployed Qdrant retrieval.
|
||||
|
||||
## What is *not* measured anywhere
|
||||
|
||||
| Standard RAG metric | State |
|
||||
|---|---|
|
||||
| Retrieval recall@k / precision@k against the live corpus | **Not found** — the code exists for condition→drug only, with no runner |
|
||||
| MRR / NDCG | **Not found** |
|
||||
| Hit-rate on the section route | Measured once by hand (0.544 overall, 0.05 on `chong_chi_dinh` for the *similarity* route, 2026-08-04) — that number is recorded in code comments and ADRs, not reproducible by any committed script |
|
||||
| Faithfulness / answer correctness scoring | Manual only (`golden_summary_v1.csv` columns) |
|
||||
| LLM-as-judge | **Deliberately absent** — `condition_evaluation.py` says so explicitly |
|
||||
| Latency distribution | Measured by hand once (n=8), recorded in `ChatPanel.tsx` |
|
||||
| Regression gate in CI | **Not found** — nothing blocks a merge on eval results |
|
||||
|
||||
## The one automated quality gate
|
||||
|
||||
`.github/workflows/deploy.yml` runs, on every deploy, a single condition→drug
|
||||
case:
|
||||
|
||||
```
|
||||
query: "Đợt gout cấp có thuốc nào được Dược thư ghi chỉ định?"
|
||||
assert: response contains "decision":"answerable"
|
||||
assert: response contains "section_key":"chi_dinh"
|
||||
```
|
||||
|
||||
Plus a second query used to assert trace propagation. If either fails, the
|
||||
deploy fails and the last 200 lines of `ai-service` logs are dumped. This is a
|
||||
smoke test on one behaviour, not an evaluation — but it is the only quality
|
||||
assertion that runs without a human.
|
||||
|
||||
## Honest assessment
|
||||
|
||||
The repository has **good evaluation *material*** and **no evaluation *system***.
|
||||
209 hand-labelled golden rows, 90 JSONL cases and two implemented deterministic
|
||||
metric summarisers exist; the wiring between them — a runner that executes a set
|
||||
against the live service, computes the metrics and compares against a baseline —
|
||||
does not.
|
||||
|
||||
Consequently, no claim of the form "retrieval quality is X" or "the system is
|
||||
production-ready because it passes evaluation" can be supported from this
|
||||
repository today. What *can* be supported is that the safety mechanisms are
|
||||
unit-tested (555 tests) and that one end-to-end behaviour is asserted on every
|
||||
deploy.
|
||||
|
||||
## Suggested minimum wiring (from what already exists)
|
||||
|
||||
1. A runner that feeds `evals/condition_to_drug_v1.jsonl` through the live
|
||||
service into `summarize_condition_outcomes` and prints the metric table.
|
||||
2. A CSV reader for `Golden Dataset/golden_e2e_v1.csv` that fills its
|
||||
`ket_qua_thuc_te` column automatically.
|
||||
3. A stored baseline plus a threshold comparison so a regression fails a check
|
||||
rather than being noticed in production.
|
||||
|
||||
All three are new code; none requires new design.
|
||||
Reference in New Issue
Block a user