Files
duocthu/coordination/CLAUDE_SPEND_LLM_LIVE_2026-08-05.md
T

2.9 KiB

Spend notice — turning on the live LLM RAG — 2026-08-05 (Claude)

Owner instruction this session: the $0 offline build is not the deliverable — stand up the real LLM RAG (semantic query embedding + LLM generation). Owner approved "Full: embed query + generation" and a cheap non-Anthropic generation model (DeepSeek / Qwen / similar Chinese model).

Phase 1 — semantic query embedding (NO IAM change needed)

  • The corpus is already embedded (duocthu_v1, 15,100 pts, cohere.embed-v4:0, corpus SHA 04a27166…, verified live today). Only the query side is off.
  • bedrock:InvokeModel on cohere.embed-v4:0 is already granted (same policy Codex used for the Titan probe today).
  • Intended call: ONE query-side probe, python -m ingestion.embed.probe --provider cohere-v4 --input-kind query, < 20 tokens, expected charge below $0.000001.
  • Then flip the ai-service default to EMBEDDING_PROVIDER=cohere-v4 and smoke a few real queries (pennies total). No corpus re-embed — vectors exist.

Phase 2 — LLM generation (NEEDS an IAM change; coordinating)

  • DeepSeek/Qwen on Bedrock use the Converse API, not the Anthropic Messages path in adapters/bedrock_claude.py. New adapter BedrockConverseAnswerGenerator to be added behind ANSWER_PROVIDER=bedrock-converse + answer_model_id.
  • Requires adding the chosen generation model ARN (e.g. arn:aws:bedrock:us-east-1::foundation-model/deepseek.v3.2) to infra/aws/iam/bedrock-embedding-invoke.json. Codex is on AWS today — this IAM attach must not collide with Codex's work. Not applied unilaterally yet.
  • Probe with ONE short Converse call before any real use.

No full-corpus run, no GPU/EC2, no recurring resource authorized by this notice.

Observed results

  • Phase 1 cohere-v4 query probe (2026-08-05): ONE call, input-kind=query, 1024 dims (expected 1024), measured L2 norm 1.000000, latency 1950.6 ms. Query embedding now shares the corpus's cohere.embed-v4:0 space. No Cohere corpus run, no IAM change. Exact bill not checked; estimate stands.
  • Phase 2 generation: IAM BedrockEmbeddingInvoke bumped to v4 (default), adding invoke on deepseek.v3.2 and cohere.rerank-v3-5:0 (+ the two embedding models). Codex was off, no collision. Repo file infra/aws/iam/bedrock-embedding-invoke.json updated to match v4.
  • deepseek.v3.2 Converse probe: 4 short calls, stopReason end_turn, 41 in / 18 out tokens, ~1.3-5.0s. Vietnamese answer returned correctly.
  • End-to-end smoke (cohere-v4 + rerank + deepseek), grounding kept ON: contraindication (section route, grounded), free-form fever question (rerank trimmed 29 sections -> 6, grounded), adult paracetamol dose (population/route labels preserved, distinct citations, grounded). All generated answers passed grounding.verify. A few dozen cloud calls total; exact bill not checked, still cents-scale on the estimate.