5.9 KiB
Ragas scoring — setup and pins
scripts/run_all_evals.py checks invariants (decision, citations, drug
provenance). scripts/score_evals_ragas.py is the other half: it scores the
quality of answers already recorded by that run, so a change to retrieval or
prompting cannot quietly degrade answers while every invariant still passes.
Why a separate virtualenv
Ragas pulls the LangChain stack, which conflicts with this service's pinned
dependencies. Installing it into the service environment on 2026-08-18 broke
botocore and langchain-core system-wide. Keep it isolated.
Install
python -m venv /path/to/ragas_env
/path/to/ragas_env/bin/python -m pip install --upgrade pip
/path/to/ragas_env/bin/python -m pip install ragas langchain-aws boto3
That alone does not work. pip resolves langchain-openai /
langchain-community to versions that need a newer langchain-core than
ragas accepts, and the import dies in ragas/embeddings/base.py on
from langchain_openai.embeddings import OpenAIEmbeddings. Pin all three:
/path/to/ragas_env/bin/python -m pip install \
"langchain-core<0.4,>=0.3.85" \
"langchain-community<0.4,>=0.3" \
"langchain-openai<0.4,>=0.3" \
"langchain-aws<1.0"
Verify before running anything real:
/path/to/ragas_env/bin/python -c \
"import ragas, langchain_aws; from ragas.llms import llm_factory; print(ragas.__version__)"
Known-good: ragas 0.4.3, boto3 1.43.x.
Models
Judged by Bedrock in us-east-1, using the same credentials as the rest of
this repo (instance role or AWS_* env). Pay per call: roughly three LLM
calls per case per metric, so a 90-case run is a few hundred calls.
| role | model | why |
|---|---|---|
| judge | qwen.qwen3-next-80b-a3b |
same model production answers with |
| embeddings | cohere.embed-multilingual-v3 |
not cohere-v4 |
cohere-v4 is what production embeds the corpus with, but langchain-aws
cannot parse v4's response envelope — BedrockEmbeddings.embed_query raises a
bare KeyError(0). It does not matter here: this embedder only measures how
close an answer sits to its question and never touches the index, so it has no
need to match the retrieval model. Multilingual does matter — the corpus and
the questions are Vietnamese.
Run
# 1. record responses (hits the deployment)
python3 scripts/run_all_evals.py \
--base-url https://realvuxbaro.me --output-dir /tmp/evals
# 2. score them (offline, re-runnable, no production traffic)
/path/to/ragas_env/bin/python scripts/score_evals_ragas.py \
--input /tmp/evals/production60.jsonl \
--output /tmp/evals/production60.ragas.jsonl
Reading the numbers
Only answerable turns are scored. An abstain or a clarify has no claims to be
faithful to, and averaging them in would move the mean for no reason.
- faithfulness — every claim traceable to the cited evidence. This is the hallucination check and the one that must stay at 1.0. Anything below means the answer asserted something its own citations do not support.
- context_precision — how much of what was retrieved was actually useful. Low values are a retrieval signal, not a safety one: the answer can be perfectly faithful while most of the retrieved chunks were noise.
- answer_relevancy — whether the answer addresses the question. Catches a well-grounded answer to a different question. Expect below 1.0 on clinical answers that legitimately add safety context the question did not ask for.
The trap that already caught us once
Contexts must carry the drug name. A citation's evidence_text is raw section
prose that often never names its own drug ("Tăng huyết áp (dùng đơn trị
liệu...)"). The service knows the drug from a separate field; a judge handed
the bare text does not.
Scored that way on 2026-08-18, a multi-drug answer came out at faithfulness
0.251 — every claim marked unsupported because no context could be
attributed to any drug. The identical run scored 1.000 once [drug_name]
was prefixed. That was a defect in the measurement, not in the service, and it
would have been reported as a model regression. _contexts_for in the scoring
script exists solely to prevent it; do not simplify it away.
The judge confuses lookalike drug names — verify before believing a low score
Run of 2026-08-19, case G14 ("Viêm phổi mắc phải ở cộng đồng dùng thuốc gì?"), scored faithfulness 0.43. It is not a hallucination. The answer reproduces the GEMIFLOXACIN indication line from printed page 720 almost verbatim, pathogen for pathogen.
Dumping the per-claim verdicts shows the judge contradicting itself:
"Mycoplasma pneumoniae chỉ được liệt kê trong chỉ định của GEMIFLOXACIN cho viêm phổi cộng đồng, nhưng không phải với GEMIFLOXACIN — mà là với GEMIFLOXACIN trong phần đầu context"
"...chỉ được liệt kê trong chỉ định của GEMIFLOXACIN, không phải GEMIFLOXACIN"
The same name sits on both sides of the contradiction. The judge is mixing up GEMIFLOXACIN and GATIFLOXACIN, which differ by two letters, and marks correct claims unsupported on that basis.
This is a formulary full of near-identical stems — -floxacin, -azolam, -tidine, -pril, -sartan — so expect it wherever one answer cites two drugs from the same class.
Treat faithfulness below 1.0 as a question, not a verdict. Before reporting one as a regression, dump the claim-level verdicts and read them against the printed page:
stmts = (await metric._create_statements(sample.to_dict(), None)).statements
verdicts = await metric._create_verdicts(sample.to_dict(), stmts, None)
for v in verdicts.statements:
print(v.verdict, v.statement, v.reason)
Both low scores this project has investigated turned out to be measurement faults, not service faults: the missing drug names on 2026-08-18, and this on 2026-08-19. That record is the reason for the rule above.