Files
duocthu/coordination/CLAUDE_REVIEW_CHUNKING_2026-08-04.md
T

3.4 KiB

Independent review request for Claude: chunking + citation provenance

Scope

Review only; do not edit until Codex and Claude compare findings.

  • ingestion/ingestion/chunk/*
  • chunk-related CLI wiring in ingestion/ingestion/cli.py
  • chunk gates in ingestion/ingestion/validation/readiness.py
  • ingestion/tests/test_chunk.py and relevant readiness tests
  • compatibility with ingestion/load/* and apps/ai-service citations

Review questions

  1. Can any chunk boundary separate a population/condition label from the dose it governs, including the single-label-current-buffer branch in _pack?
  2. Is overlap/reassembly lossless for every 15,066 canonical chunk, including comma-split long atoms and repeated text?
  3. Do table/formula descriptors and attachments ever leak unverified numeric cell content or let a consumer answer from quarantined data?
  4. Is provenance precise enough for citations? Distinguish verified printed folio from physical page and distinguish monograph-level range from the actual pages supporting each sub-chunk.
  5. Does schema v3 fail closed everywhere, or can direct chunk_all() / the Qdrant loader accept an empty/missing printed_page_range?
  6. Did adding printed_page_map introduce positional-call compatibility bugs?
  7. Are part_index, part_count, deterministic ids and Qdrant idempotency preserved after regeneration?
  8. Identify stale ADR/document claims versus the measured current corpus.

Evidence already available

  • Canonical artifact: 15,066 chunks, schema v3, SHA e474c83790b450d3262f532e81abf6526a485e3a98e376413247da23f4619c38.
  • chunk-ready: all gates pass, including chunk_without_printed_page_range = 0.
  • Full ingestion suite with local Qdrant: 258 passed; Ruff clean.
  • No real embeddings exist; do not call Bedrock or run a corpus embedding.

Requested response

Write coordination/review-chunking-claude-2026-08-04.md with findings ordered by severity. For every finding include exact file/line, a reproducer or corpus count, clinical/retrieval impact, and whether it blocks embedding. Explicitly say if no finding was found in a review area. Do not modify production code.

Codex preliminary evidence — please challenge, do not assume correct

  • Visual inspection of scratch/rag-table-pilot/out/all/crops/p209_t0.png and p209_t1.png shows their first rows are ADR data, not headers. Current descriptors embed Ngoại tâm thu thất | Thường gặp | Không rõ tần suất and Tăng bilirubin máu | Thường gặp | Thường gặp. The digit/length-only _is_label_row gate therefore violates the "no cell value" invariant.
  • Mapping normalized chunk text back to SectionPart.physical_page succeeded uniquely for all 14,915 prose chunks. Only 251 have an exact declared page range; 14,664 inherit extra monograph pages, up to six. All 151 block descriptors carry a non-exact monograph range instead of their block page.
  • _pack(["Người lớn:", "x" * 645], len) returns a first part containing only Người lớn:. The next part repeats the label through overlap, but the isolated label chunk remains independently retrievable. Current canonical corpus has 14 chunks ending :, all point to quarantined blocks; none is a population-label split.
  • validate_chunk_record() accepts a schema-v2 record with no printed_page_range; Chunk.printed_page_range also defaults to [].