# Ragas scoring — setup and pins `scripts/run_all_evals.py` checks invariants (decision, citations, drug provenance). `scripts/score_evals_ragas.py` is the other half: it scores the *quality* of answers already recorded by that run, so a change to retrieval or prompting cannot quietly degrade answers while every invariant still passes. ## Why a separate virtualenv Ragas pulls the LangChain stack, which conflicts with this service's pinned dependencies. Installing it into the service environment on 2026-08-18 broke `botocore` and `langchain-core` system-wide. Keep it isolated. ## Install ```bash python -m venv /path/to/ragas_env /path/to/ragas_env/bin/python -m pip install --upgrade pip /path/to/ragas_env/bin/python -m pip install ragas langchain-aws boto3 ``` That alone does **not** work. `pip` resolves `langchain-openai` / `langchain-community` to versions that need a newer `langchain-core` than ragas accepts, and the import dies in `ragas/embeddings/base.py` on `from langchain_openai.embeddings import OpenAIEmbeddings`. Pin all three: ```bash /path/to/ragas_env/bin/python -m pip install \ "langchain-core<0.4,>=0.3.85" \ "langchain-community<0.4,>=0.3" \ "langchain-openai<0.4,>=0.3" \ "langchain-aws<1.0" ``` Verify before running anything real: ```bash /path/to/ragas_env/bin/python -c \ "import ragas, langchain_aws; from ragas.llms import llm_factory; print(ragas.__version__)" ``` Known-good: ragas **0.4.3**, boto3 1.43.x. ## Models Judged by Bedrock in `us-east-1`, using the same credentials as the rest of this repo (instance role or `AWS_*` env). Pay per call: roughly three LLM calls per case per metric, so a 90-case run is a few hundred calls. | role | model | why | |---|---|---| | judge | `qwen.qwen3-next-80b-a3b` | same model production answers with | | embeddings | `cohere.embed-multilingual-v3` | **not** cohere-v4 | cohere-v4 is what production embeds the corpus with, but `langchain-aws` cannot parse v4's response envelope — `BedrockEmbeddings.embed_query` raises a bare `KeyError(0)`. It does not matter here: this embedder only measures how close an answer sits to its question and never touches the index, so it has no need to match the retrieval model. Multilingual does matter — the corpus and the questions are Vietnamese. ## Run ```bash # 1. record responses (hits the deployment) python3 scripts/run_all_evals.py \ --base-url https://realvuxbaro.me --output-dir /tmp/evals # 2. score them (offline, re-runnable, no production traffic) /path/to/ragas_env/bin/python scripts/score_evals_ragas.py \ --input /tmp/evals/production60.jsonl \ --output /tmp/evals/production60.ragas.jsonl ``` ## Reading the numbers Only `answerable` turns are scored. An abstain or a clarify has no claims to be faithful to, and averaging them in would move the mean for no reason. - **faithfulness** — every claim traceable to the cited evidence. This is the hallucination check and the one that must stay at 1.0. Anything below means the answer asserted something its own citations do not support. - **context_precision** — how much of what was retrieved was actually useful. Low values are a *retrieval* signal, not a safety one: the answer can be perfectly faithful while most of the retrieved chunks were noise. - **answer_relevancy** — whether the answer addresses the question. Catches a well-grounded answer to a different question. Expect below 1.0 on clinical answers that legitimately add safety context the question did not ask for. ## The trap that already caught us once Contexts must carry the drug name. A citation's `evidence_text` is raw section prose that often never names its own drug ("Tăng huyết áp (dùng đơn trị liệu...)"). The service knows the drug from a separate field; a judge handed the bare text does not. Scored that way on 2026-08-18, a multi-drug answer came out at **faithfulness 0.251** — every claim marked unsupported because no context could be attributed to any drug. The identical run scored **1.000** once `[drug_name]` was prefixed. That was a defect in the measurement, not in the service, and it would have been reported as a model regression. `_contexts_for` in the scoring script exists solely to prevent it; do not simplify it away. ## The judge confuses lookalike drug names — verify before believing a low score Run of 2026-08-19, case **G14** ("Viêm phổi mắc phải ở cộng đồng dùng thuốc gì?"), scored **faithfulness 0.43**. It is not a hallucination. The answer reproduces the GEMIFLOXACIN indication line from printed page 720 almost verbatim, pathogen for pathogen. Dumping the per-claim verdicts shows the judge contradicting itself: > "Mycoplasma pneumoniae chỉ được liệt kê trong chỉ định của GEMIFLOXACIN cho > viêm phổi cộng đồng, nhưng không phải với GEMIFLOXACIN — mà là với > GEMIFLOXACIN trong phần đầu context" > "...chỉ được liệt kê trong chỉ định của GEMIFLOXACIN, không phải > GEMIFLOXACIN" The same name sits on both sides of the contradiction. The judge is mixing up **GEMIFLOXACIN** and **GATIFLOXACIN**, which differ by two letters, and marks correct claims unsupported on that basis. This is a formulary full of near-identical stems — -floxacin, -azolam, -tidine, -pril, -sartan — so expect it wherever one answer cites two drugs from the same class. **Treat faithfulness below 1.0 as a question, not a verdict.** Before reporting one as a regression, dump the claim-level verdicts and read them against the printed page: ```python stmts = (await metric._create_statements(sample.to_dict(), None)).statements verdicts = await metric._create_verdicts(sample.to_dict(), stmts, None) for v in verdicts.statements: print(v.verdict, v.statement, v.reason) ``` Both low scores this project has investigated turned out to be measurement faults, not service faults: the missing drug names on 2026-08-18, and this on 2026-08-19. That record is the reason for the rule above.