Files
duocthu/apps/ai-service/evals/README-ragas.md
T

5.9 KiB

Ragas scoring — setup and pins

scripts/run_all_evals.py checks invariants (decision, citations, drug provenance). scripts/score_evals_ragas.py is the other half: it scores the quality of answers already recorded by that run, so a change to retrieval or prompting cannot quietly degrade answers while every invariant still passes.

Why a separate virtualenv

Ragas pulls the LangChain stack, which conflicts with this service's pinned dependencies. Installing it into the service environment on 2026-08-18 broke botocore and langchain-core system-wide. Keep it isolated.

Install

python -m venv /path/to/ragas_env
/path/to/ragas_env/bin/python -m pip install --upgrade pip
/path/to/ragas_env/bin/python -m pip install ragas langchain-aws boto3

That alone does not work. pip resolves langchain-openai / langchain-community to versions that need a newer langchain-core than ragas accepts, and the import dies in ragas/embeddings/base.py on from langchain_openai.embeddings import OpenAIEmbeddings. Pin all three:

/path/to/ragas_env/bin/python -m pip install \
  "langchain-core<0.4,>=0.3.85" \
  "langchain-community<0.4,>=0.3" \
  "langchain-openai<0.4,>=0.3" \
  "langchain-aws<1.0"

Verify before running anything real:

/path/to/ragas_env/bin/python -c \
  "import ragas, langchain_aws; from ragas.llms import llm_factory; print(ragas.__version__)"

Known-good: ragas 0.4.3, boto3 1.43.x.

Models

Judged by Bedrock in us-east-1, using the same credentials as the rest of this repo (instance role or AWS_* env). Pay per call: roughly three LLM calls per case per metric, so a 90-case run is a few hundred calls.

role model why
judge qwen.qwen3-next-80b-a3b same model production answers with
embeddings cohere.embed-multilingual-v3 not cohere-v4

cohere-v4 is what production embeds the corpus with, but langchain-aws cannot parse v4's response envelope — BedrockEmbeddings.embed_query raises a bare KeyError(0). It does not matter here: this embedder only measures how close an answer sits to its question and never touches the index, so it has no need to match the retrieval model. Multilingual does matter — the corpus and the questions are Vietnamese.

Run

# 1. record responses (hits the deployment)
python3 scripts/run_all_evals.py \
    --base-url https://realvuxbaro.me --output-dir /tmp/evals

# 2. score them (offline, re-runnable, no production traffic)
/path/to/ragas_env/bin/python scripts/score_evals_ragas.py \
    --input /tmp/evals/production60.jsonl \
    --output /tmp/evals/production60.ragas.jsonl

Reading the numbers

Only answerable turns are scored. An abstain or a clarify has no claims to be faithful to, and averaging them in would move the mean for no reason.

  • faithfulness — every claim traceable to the cited evidence. This is the hallucination check and the one that must stay at 1.0. Anything below means the answer asserted something its own citations do not support.
  • context_precision — how much of what was retrieved was actually useful. Low values are a retrieval signal, not a safety one: the answer can be perfectly faithful while most of the retrieved chunks were noise.
  • answer_relevancy — whether the answer addresses the question. Catches a well-grounded answer to a different question. Expect below 1.0 on clinical answers that legitimately add safety context the question did not ask for.

The trap that already caught us once

Contexts must carry the drug name. A citation's evidence_text is raw section prose that often never names its own drug ("Tăng huyết áp (dùng đơn trị liệu...)"). The service knows the drug from a separate field; a judge handed the bare text does not.

Scored that way on 2026-08-18, a multi-drug answer came out at faithfulness 0.251 — every claim marked unsupported because no context could be attributed to any drug. The identical run scored 1.000 once [drug_name] was prefixed. That was a defect in the measurement, not in the service, and it would have been reported as a model regression. _contexts_for in the scoring script exists solely to prevent it; do not simplify it away.

The judge confuses lookalike drug names — verify before believing a low score

Run of 2026-08-19, case G14 ("Viêm phổi mắc phải ở cộng đồng dùng thuốc gì?"), scored faithfulness 0.43. It is not a hallucination. The answer reproduces the GEMIFLOXACIN indication line from printed page 720 almost verbatim, pathogen for pathogen.

Dumping the per-claim verdicts shows the judge contradicting itself:

"Mycoplasma pneumoniae chỉ được liệt kê trong chỉ định của GEMIFLOXACIN cho viêm phổi cộng đồng, nhưng không phải với GEMIFLOXACIN — mà là với GEMIFLOXACIN trong phần đầu context"

"...chỉ được liệt kê trong chỉ định của GEMIFLOXACIN, không phải GEMIFLOXACIN"

The same name sits on both sides of the contradiction. The judge is mixing up GEMIFLOXACIN and GATIFLOXACIN, which differ by two letters, and marks correct claims unsupported on that basis.

This is a formulary full of near-identical stems — -floxacin, -azolam, -tidine, -pril, -sartan — so expect it wherever one answer cites two drugs from the same class.

Treat faithfulness below 1.0 as a question, not a verdict. Before reporting one as a regression, dump the claim-level verdicts and read them against the printed page:

stmts = (await metric._create_statements(sample.to_dict(), None)).statements
verdicts = await metric._create_verdicts(sample.to_dict(), stmts, None)
for v in verdicts.statements:
    print(v.verdict, v.statement, v.reason)

Both low scores this project has investigated turned out to be measurement faults, not service faults: the missing drug names on 2026-08-18, and this on 2026-08-19. That record is the reason for the rule above.