142 lines
5.9 KiB
Markdown
142 lines
5.9 KiB
Markdown
# Ragas scoring — setup and pins
|
|
|
|
`scripts/run_all_evals.py` checks invariants (decision, citations, drug
|
|
provenance). `scripts/score_evals_ragas.py` is the other half: it scores the
|
|
*quality* of answers already recorded by that run, so a change to retrieval or
|
|
prompting cannot quietly degrade answers while every invariant still passes.
|
|
|
|
## Why a separate virtualenv
|
|
|
|
Ragas pulls the LangChain stack, which conflicts with this service's pinned
|
|
dependencies. Installing it into the service environment on 2026-08-18 broke
|
|
`botocore` and `langchain-core` system-wide. Keep it isolated.
|
|
|
|
## Install
|
|
|
|
```bash
|
|
python -m venv /path/to/ragas_env
|
|
/path/to/ragas_env/bin/python -m pip install --upgrade pip
|
|
/path/to/ragas_env/bin/python -m pip install ragas langchain-aws boto3
|
|
```
|
|
|
|
That alone does **not** work. `pip` resolves `langchain-openai` /
|
|
`langchain-community` to versions that need a newer `langchain-core` than
|
|
ragas accepts, and the import dies in `ragas/embeddings/base.py` on
|
|
`from langchain_openai.embeddings import OpenAIEmbeddings`. Pin all three:
|
|
|
|
```bash
|
|
/path/to/ragas_env/bin/python -m pip install \
|
|
"langchain-core<0.4,>=0.3.85" \
|
|
"langchain-community<0.4,>=0.3" \
|
|
"langchain-openai<0.4,>=0.3" \
|
|
"langchain-aws<1.0"
|
|
```
|
|
|
|
Verify before running anything real:
|
|
|
|
```bash
|
|
/path/to/ragas_env/bin/python -c \
|
|
"import ragas, langchain_aws; from ragas.llms import llm_factory; print(ragas.__version__)"
|
|
```
|
|
|
|
Known-good: ragas **0.4.3**, boto3 1.43.x.
|
|
|
|
## Models
|
|
|
|
Judged by Bedrock in `us-east-1`, using the same credentials as the rest of
|
|
this repo (instance role or `AWS_*` env). Pay per call: roughly three LLM
|
|
calls per case per metric, so a 90-case run is a few hundred calls.
|
|
|
|
| role | model | why |
|
|
|---|---|---|
|
|
| judge | `qwen.qwen3-next-80b-a3b` | same model production answers with |
|
|
| embeddings | `cohere.embed-multilingual-v3` | **not** cohere-v4 |
|
|
|
|
cohere-v4 is what production embeds the corpus with, but `langchain-aws`
|
|
cannot parse v4's response envelope — `BedrockEmbeddings.embed_query` raises a
|
|
bare `KeyError(0)`. It does not matter here: this embedder only measures how
|
|
close an answer sits to its question and never touches the index, so it has no
|
|
need to match the retrieval model. Multilingual does matter — the corpus and
|
|
the questions are Vietnamese.
|
|
|
|
## Run
|
|
|
|
```bash
|
|
# 1. record responses (hits the deployment)
|
|
python3 scripts/run_all_evals.py \
|
|
--base-url https://realvuxbaro.me --output-dir /tmp/evals
|
|
|
|
# 2. score them (offline, re-runnable, no production traffic)
|
|
/path/to/ragas_env/bin/python scripts/score_evals_ragas.py \
|
|
--input /tmp/evals/production60.jsonl \
|
|
--output /tmp/evals/production60.ragas.jsonl
|
|
```
|
|
|
|
## Reading the numbers
|
|
|
|
Only `answerable` turns are scored. An abstain or a clarify has no claims to be
|
|
faithful to, and averaging them in would move the mean for no reason.
|
|
|
|
- **faithfulness** — every claim traceable to the cited evidence. This is the
|
|
hallucination check and the one that must stay at 1.0. Anything below means
|
|
the answer asserted something its own citations do not support.
|
|
- **context_precision** — how much of what was retrieved was actually useful.
|
|
Low values are a *retrieval* signal, not a safety one: the answer can be
|
|
perfectly faithful while most of the retrieved chunks were noise.
|
|
- **answer_relevancy** — whether the answer addresses the question. Catches a
|
|
well-grounded answer to a different question. Expect below 1.0 on clinical
|
|
answers that legitimately add safety context the question did not ask for.
|
|
|
|
## The trap that already caught us once
|
|
|
|
Contexts must carry the drug name. A citation's `evidence_text` is raw section
|
|
prose that often never names its own drug ("Tăng huyết áp (dùng đơn trị
|
|
liệu...)"). The service knows the drug from a separate field; a judge handed
|
|
the bare text does not.
|
|
|
|
Scored that way on 2026-08-18, a multi-drug answer came out at **faithfulness
|
|
0.251** — every claim marked unsupported because no context could be
|
|
attributed to any drug. The identical run scored **1.000** once `[drug_name]`
|
|
was prefixed. That was a defect in the measurement, not in the service, and it
|
|
would have been reported as a model regression. `_contexts_for` in the scoring
|
|
script exists solely to prevent it; do not simplify it away.
|
|
|
|
## The judge confuses lookalike drug names — verify before believing a low score
|
|
|
|
Run of 2026-08-19, case **G14** ("Viêm phổi mắc phải ở cộng đồng dùng thuốc
|
|
gì?"), scored **faithfulness 0.43**. It is not a hallucination. The answer
|
|
reproduces the GEMIFLOXACIN indication line from printed page 720 almost
|
|
verbatim, pathogen for pathogen.
|
|
|
|
Dumping the per-claim verdicts shows the judge contradicting itself:
|
|
|
|
> "Mycoplasma pneumoniae chỉ được liệt kê trong chỉ định của GEMIFLOXACIN cho
|
|
> viêm phổi cộng đồng, nhưng không phải với GEMIFLOXACIN — mà là với
|
|
> GEMIFLOXACIN trong phần đầu context"
|
|
|
|
> "...chỉ được liệt kê trong chỉ định của GEMIFLOXACIN, không phải
|
|
> GEMIFLOXACIN"
|
|
|
|
The same name sits on both sides of the contradiction. The judge is mixing up
|
|
**GEMIFLOXACIN** and **GATIFLOXACIN**, which differ by two letters, and marks
|
|
correct claims unsupported on that basis.
|
|
|
|
This is a formulary full of near-identical stems — -floxacin, -azolam,
|
|
-tidine, -pril, -sartan — so expect it wherever one answer cites two drugs
|
|
from the same class.
|
|
|
|
**Treat faithfulness below 1.0 as a question, not a verdict.** Before reporting
|
|
one as a regression, dump the claim-level verdicts and read them against the
|
|
printed page:
|
|
|
|
```python
|
|
stmts = (await metric._create_statements(sample.to_dict(), None)).statements
|
|
verdicts = await metric._create_verdicts(sample.to_dict(), stmts, None)
|
|
for v in verdicts.statements:
|
|
print(v.verdict, v.statement, v.reason)
|
|
```
|
|
|
|
Both low scores this project has investigated turned out to be measurement
|
|
faults, not service faults: the missing drug names on 2026-08-18, and this on
|
|
2026-08-19. That record is the reason for the rule above.
|