Skip to content

Clinician-Governed Learning

This module provides the validation contract for clinical mapping review. It does not provide a web UI, a database, patient-data persistence, Firestore, Cloud Storage, or HTTP endpoints. The main fax-automation project owns those responsibilities so this package remains portable and stateless.

mapper result
  -> main project stores a review case
  -> staff drafts corrections or a clinician approves a decision
  -> module validates the decision against the pinned LOINC release and safety rules
  -> main project compiles, replays, and publishes an immutable registry snapshot
  -> the next mapping job loads that exact snapshot

People and Responsibilities

Role Main responsibility
Staff reviewer Creates a draft, corrects obvious OCR facts, and attaches the report/page. Cannot activate a reusable mapping.
Authorized clinician Selects a LOINC code, chooses the scope, records rationale, and approves a reusable decision.
Main project Authenticates users, stores review history, serves the clinician UI, uploads CSV files, and starts mapping jobs.
Registry publisher Validates approvals, runs replay, writes immutable snapshots, and atomically changes the active pointer.
noise_to_loinc Searches active LOINC terms, validates decisions, compiles registry content, and replays mappings.

The UI must show a clinician the original OCR row, report/page link, mapper status and reason, ranked candidates, official LOINC name, six axes, example units, and active status. The clinician should never need to edit JSON.

What Happens to Each Outcome

Mapper result or review outcome Main-project action
mapped Use the existing accepted-only Elation path. Store provenance for audit.
abstain Show safe candidates and request a clinician choice or missing context.
invalid_candidate Show the hard safety reason. Let a reviewer correct OCR facts such as unit, analytical specimen, method, time, or scale, then remap.
no_candidate Let the clinician search active LOINC terms, select one, and provide supporting facts.
force_mapped Treat as a separately governed operational exception. Audit it outside ordinary mapping-quality metrics; it is not filed to Elation by default.
not_standardized review decision Keep the document for manual handling. Do not emit test.loinc.

An approved observation should map in the next newly launched job when the same clinical facts recur. A new row with a different or conflicting unit, specimen, method, time, or scale can still correctly abstain. The system must not promise that every future unknown clinical meaning will map automatically.

Doctor-Friendly Review Form

The main project should present this sequence, in plain language:

  1. Confirm the original report row and page.
  2. Correct an OCR fact only when the source document supports it: name, unit, analytical specimen, collection specimen, method, time, or scale.
  3. Search active LOINC terms by code, official name, or synonym.
  4. Inspect the chosen code's component, property, time, system, scale, method, and example units.
  5. Answer one scope question:
  6. Works at all laboratories: approve_universal and universal_template.
  7. Only this laboratory or test method: approve_source_specific and source_evidence.
  8. Enter the clinical rationale and submit with the authenticated reviewer identity.

fast_path_enabled, raw registry JSON, and routine force-exact controls must not appear in this UI. A normal clinician approval never disables UCUM, unit/property, specimen, method, time, scale, or active-code checks. A separately authorized backend-only operational workflow may create approve_force_exact; it is intentionally not a clinical confidence override.

Public Package Contract

The main project creates and persists a ReviewCase after an uncertain result:

from loinc_mapper import (
    ReviewDecision,
    PipelineSettings,
    PublicationPolicy,
    build_review_case,
    build_review_case_from_deidentified_facts,
    compile_registry_snapshot,
    correct_and_recheck,
    finalize_registry_snapshot,
    replay_registry_snapshot,
    validate_review_decision,
)

review_case = build_review_case(
    observation,
    mapping_result,
    report_reference={
        "report_id": "fax-123",
        "row_id": "row-7",
        "page": 2,
        "source_document_reference": "private://reports/fax-123/page-2",
    },
)

decision = ReviewDecision.from_dict({
    "case_id": review_case.case_id,
    "case_revision": review_case.case_revision,
    "decision": "approve_universal",
    "selected_loinc": "2345-7",
    "mapping_scope": "universal_template",
    "reviewer_id": "clinician-account-id",
    "reviewer_name": "Dr Example",
    "rationale": "Reviewed report label, mg/dL unit, and serum/plasma result style.",
})

validation = validate_review_decision(review_case, decision, catalog, active_registry)
if not validation.eligible_for_publish:
    raise ValueError(validation.errors)

policy = PublicationPolicy.CLINICIAN_FAST_REPLAY
candidate = compile_registry_snapshot(
    active_registry,
    [validation],
    catalog,
    publication_policy=policy,
)
if candidate.requires_publication:
    # mapper_factory uses PipelineSettings.clinician_review_replay().
    replay = replay_registry_snapshot(
        candidate,
        [validation],
        mapper_factory,
        publication_policy=policy,
    )
    active_snapshot = finalize_registry_snapshot(
        candidate,
        replay,
        publication_policy=policy,
    )

For a staff correction, persist the original case and decision, then use the storage-free helper to obtain the corrected facts and fresh mapper result:

recheck = correct_and_recheck(review_case, correction_decision, mapper)
# Persist recheck.corrected_observation and recheck.mapping_result as a new audit revision.

If the main project starts from already de-identified fact rows, use build_review_case_from_deidentified_facts. It accepts only the documented mapping fields and rejects patient/free-form report fields; the report reference must remain an opaque application identifier.

For normal clinician publication, mapper_factory must use PipelineSettings.clinician_review_replay(). It uses the pinned LOINC catalog and candidate registry but does not load or invoke UMLS, ScispaCy, FAISS, SapBERT, or the learned ranker. It still runs exact-alias, active-code, context, UCUM, and six-axis validation. finalize_registry_snapshot is pure: it does not write files or move an active pointer.

The candidate payload has registry_status: candidate and cannot be loaded by MappingRegistry.from_json. The finalized payload has registry_status: active and a manifest containing parent version, LOINC release, reviewer decision IDs, case IDs, checksum, replay result hash, and timestamps.

clinician_fast_replay maps all changed approvals through one map_many() batch against their original reviewed facts. A passed replay immediately enables the exact universal-template fast path. It is not a force-map: inactive codes, contained/fuzzy aliases, invalid units, conflicting specimens, and failed six-axis checks still cannot publish. Source-specific entries remain full-pipeline evidence, and clinicians never set fast-path behavior.

The mapper also ships the editable config/axis_context_aliases.json with the runtime bundle. Its method_families section contains reusable report-to-LOINC method surfaces, for example ImmunoBlot and IB, or IFA and IF. A method family is evidence normalization only; it never selects a code. Configuration changes require a version bump, regression tests, and worker restart. The package keeps conservative fallback families if the file cannot be read, so a bad deployment cannot silently remove safety behavior.

For collection-only reports, preserve the distinction between the draw and the material analyzed. collection_specimen_code=blood_venous does not mean the analytical specimen is Bld. A replay-certified universal approval may use a reviewed serum_or_plasma template with a venous collection-only future row, but an explicit urine, CSF, arterial, capillary, or incompatible analytical specimen remains a hard rejection. An approval response marked eligible_for_publish still needs the main project's publisher to finalize an active immutable snapshot before workers can use it.

strict_replay remains available for release audits. It performs the original and unseen-laboratory full replays plus an equivalence fast replay, but it is not the normal clinician publication path.

Restricted force-exact workflow

approve_force_exact with mapping_scope=force_exact is reserved for an explicitly authorized operational exception. It requires an active LOINC term eligible for a result row, an exact lossless mapping name, and a matching LOINC-derived result style, then publishes a separate force_exact_overrides registry section. A fax application may supply a traced repaired mapping_name, but must retain the original OCR label and source-cell evidence. The override deliberately bypasses unit and six-axis validation at runtime, returns status=force_mapped, and is excluded from ordinary mapping metrics and model-training data. The module blocks Elation conversion unless the trusted worker deployment, not a browser, sets IS_FORCE_MAPPED_ELATION_ON=true.

Use this only when policy specifically authorizes the exception. It is never a remedy for a unit, specimen, method, time, or property conflict that can be corrected from the source document.

Approved Alias Families

When an approval is compiled, the module generates a bounded verified_aliases list from the approved name. It contains only normalized, compact, and reviewed OCR-confusion surfaces; it does not invent semantic synonyms. Every surface remains attached to the original clinician decision and is subject to the same active-code, UCUM, specimen, method, time, scale, and six-axis checks.

An exact match to a replay-certified verified_aliases surface may receive the same clinician authority as the canonical alias. A contained display match or arbitrary fuzzy match never receives that authority and must still pass ordinary confidence and margin thresholds. This improves harmless OCR recall without turning fuzzy similarity into a clinical override.

The registry also folds the closed variations a report applies to an approved surface (normalization.surface_key): singular and plural, %/percent (pct only beside other words; alone it is procalcitonin), and word order, except after a slash or a relation word (ratio, index, per, vs), where order is meaning. A printed Lymphocyte %, Neutrophil % or Granulocytes, Immature, % therefore finds the approved Lymphocytes %, Neutrophils, percent and Immature Granulocytes % as a verified variant (evidence alias_fold: true), including on the registry fast path. Nothing else folds: an acronym (RDW Standard Deviation for RDW-SD) is a new surface a clinician approves, not a rule. tools/registry_lint.py lists, under its own classification divided (2026-10-07), a universal approval of a surface beside another code for the same surface in one laboratory's scope (Vitamin D, 25-Hydroxy -> 62292-8 universally and -> 1989-3 for one laboratory): the scopes keep the pair apart at runtime, but the practice has answered one printed surface with two codes, a clinician's decision to make.

Word breaks are not meaning either: LDL-P, LDLP and LDL P are one surface (normalization.joined_text, the normalized text without its spaces; it keeps %, unlike the compact form). A row that differs from an approved alias only in word breaks is that alias (match_kind joined), unless two approvals with different targets share the row's compact form: PSAFREE may be PSA, FREE or PSA, % FREE with the % lost, so there it is a matched surface, not the row's own. Publication treats a joined surface as the same alias.

Two facts keep this safe. Every registry candidate records exact_for_row: whether the row's own mapping name is that alias itself under the fold. A compact-form hit (Lymphocytes reaching Lymphocytes % because the compact form erases %) or an OCR hypothesis is still a matched surface, but not an approval of that row's surface. And clinician authority clears the thresholds only for one approved meaning: when approved candidates with different targets survive on one row, authority stays only with the row's own surface when it ranks first (an exact source test id or explicit discriminator entry keeps its documented precedence); otherwise the row reviews. So PSA, FREE → 10886-0 and PSA, % FREE → 12841-3 may both be published (the publish conflict rule compares folded surfaces, not compact forms), each row keeps its own approval, and an OCR row that is the exact surface of neither is reviewed rather than filed by code order.

Replay diagnostics are bounded to the top 10 candidates by default. They include status, code, confidence, margin, reasons, display names such as 718-7 - Hemoglobin [Mass/volume] in Blood, and registry evidence. They omit patient and report payloads.

Decision Values

decision Effect
approve_universal Creates a reusable source-neutral template after deterministic safety validation and one batched exact replay of the reviewed row.
approve_source_specific Creates evidence for the named laboratory/test context only. It never blocks universal retrieval for other laboratories.
approve_force_exact Restricted backend-only operational exception. Creates a separate exact-lossless override after active target/result-row validation; it is not a normal clinician approval.
correct_and_recheck Stores corrected OCR facts as an audited revision and returns them for remapping. It creates no registry entry.
reject_candidate Records why a candidate is clinically incorrect. It creates no registry entry.
not_standardized Records that no suitable structured LOINC should be emitted.
not_lab_result Records that the row does not belong in this laboratory mapper.
defer Saves no reusable decision.

The first two decisions use the normal universal_template or source_evidence scope. The restricted force action alone uses mapping_scope=force_exact. A source-specific approval requires source_laboratory. A universal approval retains the originating laboratory as provenance but does not make it a runtime prerequisite.

If a legacy source_evidence row already exists for the same alias and target, the universal approval adds a universal template and preserves the source row for audit/history. Candidate merging gives the replay-certified clinician universal entry priority over the legacy source evidence for that same code. If the existing legacy row has the same scope, publication upgrades it in the candidate snapshot instead of adding a duplicate sibling.

When a universal and source-specific entry point to different codes, both stay in the candidate union. Explicit row facts and hard validation always run first. An exact local source_test_id may select a documented local assay; otherwise, a verified clinician universal template outranks a lab-name-only source alias. A source-laboratory name by itself is retrieval evidence, never a global veto over universal mapping.

Laboratory names

A source-scoped entry is admitted for a row when the two laboratory names name the same laboratory, not only when they are spelled alike. config/laboratory_vendors.json lists the practice's own vendors: each Elation vendor id with the surfaces that vendor prints (911547826422: Quest, Quest Diagnostics, Quest Diagnostics Incorporated; 911547957494: Labcorp, Laboratory Corporation of America, ...). A vendor id or any listed surface resolves to the vendor's canonical name (normalization.laboratory_key); every other name keeps its normalized text. Two keys match when they are equal, or when the shorter one has at least five letters and its words are a prefix of the longer one's (WellSpan matches Wellspan Health), unless that prefix would also open a second laboratory already in the registry: the registry computes its distinct laboratory keys once, and Health alone matches nothing. A name that prints longer than the registry's (Bellin Health against Emplify Health Bellin Health) still misses; a clinician approval from such a fax publishes the printed name and closes the gap. The same rule scopes the registry's exact, contained, fuzzy and force lookups and the publish conflict rule.

Publication canonicalises the name: a case whose fax says Quest Diagnostics publishes source_laboratory: Quest, so the guard is not defeated by the first approval, and supporting_source_laboratory stays provenance. The runtime manifest lists the file (configs.laboratory_vendors) with its version; editing it bumps the version.

Registry constraints are discriminators

A specimen or method constraint on an approval (required_specimen, required_method) fails only on a printed contradiction; absence never contradicts (2026-09-24). The row's analytical specimen is compared, or, when the row prints none, the confirmed panel's reviewed default; a collection source contradicts only when it names other material (a urine collection against Ser/Plas); a report-level specimen is left to the validator, which reviews it. A method constraint fails only on a printed incompatible method. So Estimated Average Glucose (Bld), CO2 and hs-CRP approvals now apply to rows that print no specimen or method.

A constraint that restates the approved term's own System or Method adds nothing the validator does not already check. New approvals therefore publish a specimen or method constraint only when the decision names it as a discriminator (discriminators: ["specimen"], payload key and CSV column discriminators; allowed values specimen and method): a printed fact that tells this approval apart from another approval of the same printed name. tools/registry_constraint_report.py lists the constraints of a snapshot and classifies each as a restatement, a discriminator (another approval of the same folded name targets another code) or a scope limit, for the clinician to decide which to keep; removing any is a new snapshot through this contract, never an edit. The 2026-09-22 snapshot holds 59 method and 136 specimen restatements and 6 specimen constraints on names with a second approved target (the LDL cholesterol calculated-versus-unspecified pair and the urine-versus-blood WBC / RBC approvals).

Approvals that meet

A printed name reaches an approval through its normalized text (which already corrects common misspellings), the closed fold, the joined surface and the bounded fuzzy lookup (which since 2026-10-09 never crosses a number printed as a word of its own: the universal OMEGA-6 TOTAL approval had reached OMEGA-3 TOTAL as a close spelling). Two approvals that meet there with different targets make the outcome depend on how the report spelled the name. tools/registry_lint.py lists them. On the 2026-09-22 snapshot it finds 16 conflicts, all the legacy 2026-08-27 pair LDL cholestrol / Low dessity lipoprotein cholesterol -> 13457-7 beside LDL cholesterol / Low Density Lipoprotein Cholesterol -> 2089-1 (the normalizer corrects both misspellings, so they are one surface at runtime). It also finds 16 scoped pairs: a laboratory's own approval beside the universal one (Cortisol, A.M., Vitamin D, 25-Hydroxy, Immature Granulocyte %, WBC / RBC for one laboratory), and SHBG 13967-5 / 2942-1 kept apart by their required units. The clinician decides; a change is a new snapshot through this contract. Approvals that record the question they answered (round K) are classified by the contract's closest-question rule; pass --catalog and --panel-index so the lint reads what a printed panel implies, as publication does.

Correcting an approval: supersedes

An approval may name the code it replaces for its printed surface (supersedes, payload key and CSV column; universal or source-specific approvals only, never the selected code itself). Validation then treats the superseded entry as the correction instead of a conflict, and fails when no active approval of that surface targets the named code. Compile removes every same-surface entry with that code (fold, joined and spelling-corrected forms included) and records each in the manifest (superseded_entries: alias, target, replaced_by, decision_id); a decision that superseded something counts as a change even when its own code already existed. Round 20260924 used it for CO2 / Total CO2 / TCO2 / CO2 Total Plasma (1963-8 bicarbonate -> 2028-9 total carbon dioxide) and for the two misspelled LDL-cholesterol entries (13457-7 -> 2089-1).

Round 20260925 (evaluation/clinical_review/publish_round_20260925.json, Dr Tro's decisions delegated to the mapper session by Dhruv on 2026-09-25, compiled from de-identified facts with --facts on the recertified 20260924 registry into registry-20260925T181028-ed9e7ab4b19a, 17 added, strict full and fast-path replays 16/16): the coagulation surfaces PT / Prothrombin Time -> 5902-2, INR -> 6301-6, APTT / aPTT / PTT -> 14979-9; Quest's immunoassay total testosterone -> 83116-4 and bioavailable testosterone -> 2990-0; the LC/MS/MS insulin and C-peptide surfaces -> the practice's 20448-7 and 1986-9; INSULIN RESISTANCE SCORE -> 92845-7; ALBUMIN/CREATININE RATIO, RANDOM URINE -> 9318-7 in mg/g creat; RDW Coeff of Var -> 788-0; PSA Screen -> 2857-1 (judgement call, listed for Dr Tro); Vitamin B6 -> 30552-4; Folate -> 2284-8; HIV 1&2 Antibody -> 7918-6 (flagged for Dr Tro: the antigen/antibody combination 56888-1 is another analyte); Bilirubin, Direct / Direct Bilirubin -> 1968-7 (the registry had no direct bilirubin approval). Panel index 2.82.7 adds the practice record Cardio IQ Insulin Resistance Panel (members 20448-7, 1986-9, 92845-7; no Elation order carries the heading) and every member's release names; 2.82.8 gives the Cardio IQ record Quest's printed heading LIPOPROTEIN FRACTIONATION, ION MOBILITY, so it no longer falls to the NMR record (whose members it shares) through an enum guess.

Round 20261008 (evaluation/clinical_review/publish_round_20261008.json, Dr Tro's review of 2026-10-08 relayed by Dhruv, compiled from de-identified facts with --facts on the 20260925 ACTIVE registry into registry-20261008T172328-65b01b4d64f1, 19 added, 15 unchanged, strict full and fast-path replays 17/17; Release A with panel index 2.82.10 and package 0.3.5): the NMR LipoProfile's HDL particle surfaces (HDL-P (total), HDL-P, HDL Particle Number, Total HDL-P, HDL-P, Total, HDL particles) -> 49748-7 in umol/L as normal approvals, never force entries; LDL-P / LDL-P (LDL Particle Number) -> 54434-6; LDL-C (calculated) -> 13457-7; and LDL-C with the printed method calculated -> 13457-7 as the second question of the name beside ldlc -> 2089-1, which stays for rows that print no method (the method is a question fact, so no rule was needed). Vitamin D: the practice's default for a total 25-hydroxyvitamin D print is the D2+D3 total 62292-8 — the Labcorp Vitamin D, 25-Hydroxy and Quest Vitamin D, 25-OH, Total source-specific approvals and the universal 25-hydroxyvitamin D, total approval of the D3 term 1989-3 were superseded (supersedes 1989-3, recorded in the manifest), every other total print of the order sets and the week's reports -> 62292-8, the term's own name, the immunoassay print (VITAMIN D,25-OH,TOTAL,IA -> 83070-3) and the force entry kept, and the fractions added (Vitamin D3, 25-Hydroxy / 25-OH Vitamin D3 -> 1989-3, Vitamin D2, 25-Hydroxy / 25-OH Vitamin D2 -> 49054-0). registry_lint lists no vitamin D surface as divided any more. The round needed one validator change (0.3.5): HDL-P (total) had been refused as "a total against an unstated sub-fraction" because LOINC writes HDL as Lipoprotein.alpha; a printed fraction acronym the term's release names carry now names that sub-fraction (Safety). Publication order: the registry before or with index 2.82.10 — with the NMR heading resolved and no HDL-P approval, HDL-P (total) would file HDL cholesterol in moles (14646-4).

Round 20261009 (evaluation/clinical_review/publish_round_20261009.json, Dhruv's decisions of 2026-10-09 under Dr Tro's delegation after the final sandbox run, compiled with --facts on the 20261008 ACTIVE registry; the sandbox round, package 0.3.6 with index 2.82.11): the eight OmegaCheck fraction surfaces (EPA, DHA, DPA, EPA+DPA+DHA, ARACHIDONIC ACID, LINOLEIC ACID, OMEGA-6 TOTAL, ARACHIDONIC ACID/EPA RATIO) -> the Blood terms Dr Tro approved on 2026-09-18 for one laboratory (90912-7, 90914-3, 90913-5, 90911-9, 90916-8, 90917-6, 90915-0, 90909-3), now universal — the source-specific approvals had left Quest's whole-blood OmegaCheck rows to the serum series under the heading; LEAD / LEAD (VENOUS) -> 77307-7 universal (ug/dL, mcg/dL); NUCLEATED RBCS printed per 100 WBC'S -> the ratio 58413-6.

Fast-path certification of older approvals

Only the approvals a review round adds are certified for the registry fast path by that round's replay. The 2026-08-27 order-set universal approvals never were, so each of their rows paid for the full semantic pipeline. tools/recertify_fast_path.py gives them the same strict replay: each entry's own surface with its first required unit and its constraints, with no laboratory and at an unseen laboratory, through the full pipeline and with the entry's fast path on. An entry is certified only when all four results map its target. Panel-scoped (required_panel) and force entries are never candidates. Both runs switch the force table off: the force table answers before any entry is read, so with it on an entry whose surface is also a force entry was never tested (its row came back force_mapped); the snapshot keeps the force table unchanged. The output is a new active snapshot (fast_path_certification: recertify_existing_strict, parent and report hash in the manifest) that the operator publishes like any other.

Every snapshot manifest this contract writes (compile and recertification) records its full ancestry, lineage_versions, nearest first, as the main project's publish page does: the parent, then the parent's own lineage read from the manifest beside the parent file (review_contract.manifest_lineage). A service pinned to an older snapshot trusts a descendant only through it.

Panel-scoped approvals

Some labels are complete only under their heading: % SATURATION is iron saturation on an iron panel and oxygen saturation on a blood-gas page. A universal approval may therefore carry required_panel, a record key of the panel index (panel_context_enum.values[].value, compared by equality, no folding). The published entry keeps the field; at runtime the candidate passes the registry constraint only when the row's resolved panel is that record through a trusted heading, an explicit panel LOINC, or a sibling inference of it (a heading-less iron table still names the iron panel through its rows). A hint or an unrecognized heading never satisfies it, and a blocked entry takes the blocked-approved-code outcome (missing_required_clinical_fact naming the approved code) instead of letting the blood-gas survivor file. Outside its panel the entry is inapplicable, not a veto of its code (2026-09-25): a row that merely contains the scoped alias (Iron Saturation contains saturation) keeps 2502-3 as an ordinary candidate with its retrieval score, so the code is judged on the row's own evidence instead of vanishing behind the entry (the row had filed the molar twin). The registry fast path cannot see the heading, so an entry with a panel scope always runs the full pipeline.

required_panel is allowed with approve_universal only (the CSV has a required_panel column). When the validator is given the panel index, the value must be a record key and the review case itself must resolve to that record through its heading or panel LOINC hint; without an index the key is accepted and the replay proves it. One label may be approved once per practice panel that prints it (% SATURATION under Iron, TIBC, and Ferritin Panel and under Iron and TIBC are two entries), while a bare approval of the same label to another code still conflicts with a panel-scoped one: the reviewer resolves it, never a silent coexistence of two codes for one printed label.

A decision can only name the key the case resolved to when it was mapped (stages.panel_context.panel_key), and published entries carry that key. Record keys are therefore stable identifiers: if an index rebuild renames a record, older review cases fail the check with must be the trusted panel this review case resolves to (re-map the case; that is the index, not the review form), and the builder refuses an index whose keys no longer cover every required_panel in the registry it was given, so a renamed record cannot silently orphan a published approval. Rename by adding the new spelling as an alias of the same key instead.

One question, one answer (round K, 2026-09-29)

The review redesign of the main project's BUG-88 changes what a review case is: one question per test, answered once. The mapper owns the question, the scope of an approval, the preview a reviewer sees before approving, the registry view that lets an approval apply before it is published, and a publish that returns a failing approval alone. The sections below are added chunk by chunk; the plan is the last addendum of PLAN.md.

Fields the contract gained

Every change is additive. A registry, a case or a decision written before round K loads and behaves as before.

Where Field Meaning
registry entry question the facts the reviewed rows printed, fact name to code; recorded on every new approval, and what the approval is held to: a row that prints a contradicting fact does not take it
registry entry decision_id, reviewer_name the decision behind the entry and the reviewer's display name (reviewer stays the id)
registry entry review_status: approved_pending approved and applying to its own question, not yet certified by a publish replay; verified_active stays the published status
candidate evidence approval_state, approved_at, decision_id, reviewer_name what a filing note says about the decision behind a code: pending or published
observation printed_label the label the row prints once a verified cleaning removed layout, captions and the row's own facts; never an expansion
observation name_grounding text, image_only or name_unverified
review case schema_version: 2 the observation is the row's own (raw_name, printed_label, mapping_name, input_quality, repair_operations, specimen_source, result_kind, name_grounding); version 1 still loads
review decision supersedes_decision_id the one approval this decision replaces (payload key and CSV column); supersedes still names a code

The new entry fields are written to a snapshot only when they are set, so the payload and the content hash of an older snapshot are what they were. Two helpers replace the literal status comparisons: RegistryEntry.clinician_decision (published or pending: authority for the row's own question) and RegistryEntry.clinician_published (published only: the practice's evidence, such as its approved targets and materials, and everything a publish certifies). A pending entry is never returned by the fast-path lookup.

The question a row asks

loinc_mapper.question owns the question key; the main project never computes it. A question is the facts the report prints about a test:

Fact Read from Normal form Absent when
name printed_label when the application sends it, else raw_name; a row whose name_grounding is name_unverified keeps raw_name the registry's closed folds (plural, %/percent, word order, word breaks) and the closed OCR fold (1/I/l, 0/O) never
value_kind the value's shape quantitative (number, comparator, range, or a word beside a unit that states a dimension), titer, ordinal (answer word, grade), narrative (sentence) no value, or a mixed shape
unit_family the unit grammar the dimension (mass per volume), never the spelling; dimensionless prints split into percent, ratio by basis, titer, index, plain number missing, unreadable, note only, magnitude only
panel a printed heading or order code (official_panel_heading, confidence 0.95 or more) the panel index record key, with the member codes the record names for the row and the material it implies a hint, a sibling inference, an unrecognized heading
specimen the row's or its section's specimen the material family (serum and plasma are one) not printed, or specimen_source: report
method the method column the reviewed method family, else the normalized text not printed
collection_time time_context the LOINC time code not printed

Never part of it: the laboratory, the unit's spelling, the mapper's outcome, how the row filed, the run. The name is the print, not a model's expansion: Basos keys as Basos, and its verified expansion becomes an alias at approval. A cell that glues a caption to its label (Sodium Normal Range: 135 - 145 mmol/L) keys as the label the application cleaned (printed_label: Sodium).

A row whose value is a pointer or a status, whose cell status is not result, or whose structure is invalid is not a result and asks no question.

Every mapping result carries the question of its row:

"question": {
  "id": "b77ef284-d218-57bf-90d4-4c6e83e74de0",
  "name_id": "8a9589d68e34a1d1e87d3964",
  "key_version": "1",
  "facts": {
    "name": {"code": "glucose", "words": "GLUCOSE"},
    "value_kind": {"code": "quantitative", "words": "a number"},
    "unit_family": {"code": "mass^1/volume^-1", "words": "mass per volume", "unit": "mg/dL"},
    "panel": {"code": "Comprehensive Metabolic Panel (CMP)", "words": "Comprehensive Metabolic Panel (CMP)", "members": ["2345-7"], "material": "ser/plas"}
  },
  "refused_approvals": []
}

id is the question the row's own facts open (a UUID, usable as a case id); name_id is shared by every question of the name and is the main project's index. An absent fact is omitted, never null. key_version (question.KEY_VERSION, also in the runtime manifest as question_key) changes whenever a key computed before would differ; a case stores it.

refused_approvals lists the approvals of the row's own name that did not apply, each {code, axis, reason_code, kind, state, decision_id | alias}: axis is the printed fact the refusal turns on (panel, specimen, method, collection_time, value_kind, unit_family, name, other); reason_code is one of the closed set panel_conflict, specimen_conflict, method_conflict, collection_time_conflict, value_kind_conflict, unit_family_conflict, unit_review, name_conflict, other; kind is rejected or unit_review; state is published or pending. What the main project read before round K is unchanged: the primary outcome's reason codes, blocked_approved_candidates, inapplicable_approved_candidates, registry_evidence[].applicable and the unit diagnostics' details.expected.

Which question a row joins

  • question.compatible(question, row): one name, and no fact both print that differs. An absent fact matches anything.
  • Two different printed panels split a question only when they change what the name means (practice_policy.panel_splits_question: meaning): the records name different member codes for the row, or, when a record names none, imply different materials. GLUCOSE under CMP and under BMP is one question (both name 2345-7); under the urinalysis heading it is another. With always every printed panel is its own question.
  • question.merge_facts(question, row): a question's facts are the first printed value of each fact across its rows. Its id never changes.
  • question.join(row_question, questions) takes the row's whole stages.question and a list of {id, facts, decision_ids, state: open | decided} and returns {index, facts, reason}. A row joins the one compatible question (reason: joined, the merged facts). It never joins a decided question whose approval is among its refused_approvals (matched by decision id): it joins or opens the narrower question the printed fact keys (refused). With no compatible question it is its own (no_compatible_question).
  • The closest question wins. Among the compatible questions a row joins the one that shares the most printed facts with it, then the one with the fewest facts the row does not print (question.closeness, question.closest). Albumin asked under the chemistry panel and under an electrophoresis heading: a row under either heading joins its own; a bare row ties, so it is its own question, asked once (ambiguous); the next bare row is closest to that bare question and joins it. Only a true tie has no winner.

The key never decides a code. It decides which card a row belongs to and which approvals it may use; every filing still passes the six-axis and unit check on the row's own facts, so a grouping mistake can ask one question too many or too few and cannot file a wrong code.

Conflict means the same question

Two codes for one printed name conflict only when they answer the same question. Serum glucose under the chemistry panel and urine glucose under the urinalysis heading coexist, as do a titer and a presence code of one antibody, and a mass and a molar code kept apart by the printed unit. One rule, the closest question, decides it everywhere (approval_questions.closest_entries): at validation and preview, in compile_registry_snapshot, in tools/registry_lint.py, in tools/compile_clinician_approvals.py and at runtime.

  • The candidates are the approvals of the name (the alias folds and the scope, as before) and the open sibling questions the application passes (siblings).
  • An approval counts only when its code could be filed for the reviewed row: the six-axis and unit check does not hard-reject it, and its System agrees with the material the row prints or its printed panel implies. A serum code approved for the bare label is no answer for a row under the urinalysis heading, whatever the order the approvals were made in. A Keep or Replace choice is therefore always between two codes the row could file.
  • The closest question answers. When the best is shared by different questions that all carry one code, that code answers; when they carry different codes, or one of them is an open sibling, the question is its own and nothing answers it yet.
  • An answer with the same code confirms; an answer with another code is a conflict that names the earlier decision; the approval a decision replaces (supersedes_decision_id, or supersedes for an approval published before round K, which has no decision id) is not counted.

An approval published before round K recorded no question: its facts are read from what it enforces (its unit, panel, specimen and method constraints) and, for the kind of value, from its target's Property and Scale.

What an approval is held to

An approval records the question it answered and applies to a row unless the row prints a fact that contradicts it. An absent fact contradicts nothing, so a bare row takes the approval its fuller sibling earned; a row printing another panel meaning, specimen, method, collection time, kind of value or unit family asks another question. No required_panel, required_specimen or required_method is computed from the question (a computed panel requirement would have refused the bare row); the ones a decision sends explicitly (required_panel, discriminators) still win and behave as before.

  • The alias is the printed name (printed_label, else the print). A verified expansion (Basos mapped as Basophils) becomes a verified alias, because the runtime looks the registry up by the mapping input name; a row whose name_grounding is name_unverified adds none.
  • The unit limit and the six-axis check use the reviewed row completed with the question's accumulated facts (question=): a case keeps its first fax's row, which may have printed no unit. ReviewValidation.corrected_observation is that completed row.
  • A contradicted approval is inapplicable, never a veto of its code, and is listed in stages.question.refused_approvals with the axis (registry question <fact> differs from the printed <fact>).

Which approval files a row

Among the approvals of the row's own name that passed validation and carry different codes, the closest question leads (Mapper._closest_approval_first): its candidate gets the evidence tier above the others and closest_question: true, and it leads its meaning group as well (an approval made for a row that prints another method never takes the bare row from the approval made for it, although the group would prefer the method-less code between the two). One code approved twice (under the panel and as a bare label) holds one candidate; each of its approvals is weighed and the one that answers the row becomes the candidate's own evidence (decision_id, approval_state, reviewer_name, approved_at, question), the others travel as registry_alternatives. A tie between different codes stays a review. A name with a clinician decision for another code never takes the fast path, and the fast path holds its entry to the recorded question as well.

Preview and the structured outcome

preview_review_selection(case, selected_loinc, catalog, registry, panel_index=None, *, question=None, siblings=()) says what approving a code would mean before anyone approves; validate_review_decision takes the same keywords and returns the same dictionary as ReviewValidation.outcome (also in to_dict()), so a second approval of one question comes back as a conflict naming the earlier decision. The sentence in errors stays for existing callers.

Key Content
state new, confirms, conflict or refused (PREVIEW_STATES)
message card-ready words; never the printed value, only its shape
selected {code, display}
earlier on confirms and conflict: {code, display, decision_id, alias, reviewer, reviewer_name, reviewed_at, state: published \| pending, registry_version}; decision_id is null for an approval published before round K, registry_version null for a pending one
refusal on refused: {axis, reason_code}; reason_code is one of PREVIEW_REASON_CODES
method_hint when the code names a method the report does not print: {method, laboratory, message}
scope the facts the approval will record
options ["keep", "replace"] on a conflict, else empty
replaces the decision id the decision's supersedes_decision_id named, else null

PREVIEW_REASON_CODES is closed: the refusal codes of refused_approvals and the contract's own (code_not_active, not_a_lab_result, source_laboratory_missing, supersedes_not_found, panel_not_resolved, unit_not_readable, force_target_not_a_result). The structured fields are the contract; message is the fallback for a reason the application does not know.

Replace sends supersedes_decision_id = earlier.decision_id, or supersedes = earlier.code when the earlier approval has no decision id. A replacement of an approval that never reached the registry is no error.

Preview and validation are pure reads over the loaded terms, the registry passed in and the panel index: no SQL, no mapper. The application calls them without its mapper lock while a fax is mapping.

Approve applies now: the pending view

registry_with_pending(active_registry, validations, *, catalog=None, panel_index=None) returns a PendingView: the published registry (the base, what ACTIVE names) plus the approved, unpublished decisions that hold on it. It is pure, it never changes the base, and nothing stands behind it on disk. The main project rebuilds it whenever the set of approved decisions changes and installs it with Mapper.replace_registry(view) between mapping batches.

Attribute Content
registry a MappingRegistry with registry_status: pending_view, base_version, pending_fingerprint, pending_count
applied the decision ids in force, in the order taken
skipped [{decision_id, reason_code, message, other?}]; reason_code is one of VIEW_REASON_CODES: not_eligible, conflict (with other, the earlier decision), force_exact_needs_publish
superseded [{decision_id, alias, target, state: published \| pending, replaced_by}]: base entries hidden and pending ones withdrawn by a replacement
superseded_unpublished [{decision_id, supersedes_decision_id}]: a replacement of an approval that never reached the registry; no error
base_version, fingerprint the ACTIVE version and pending_fingerprint(base_version, decision_ids) over every supplied decision, applied or not
to_dict() everything but the registry
  • The decisions are taken in the order of reviewed_at (the order given, where equal) and each is judged against the base and the earlier ones with the rule of validation, so the earlier decision stands. The view reads from a validation its eligibility, its entry, its completed row and its decision; nothing else depends on the registry the validation was made against. The main project validates each pending decision against the view of the ones before it, which is the registry its reviewer saw.
  • A pending approval answers the row's own name only: the normalizer's reading and the closed folds (word breaks, plural, percent, word order). Never a retrieval variant, a contained alias or a fuzzy one, and never the fast path. Its registry_match is exact or exact_verified_variant, the kinds a filing gate treats as authoritative. A force-exact exception applies once it is published.
  • A pending approval is no practice evidence: the approved targets, materials, properties, methods and scales, the default material and the panel member confirmation read published entries only.
  • The winning candidate's evidence says approval_state: pending | published, decision_id, reviewer, reviewer_name, approved_at and registry_version.
  • The view's version is <active version>+pending.<applied count>.<first eight of the fingerprint>. runtime_manifest.registry reports status (the registry's own, no longer a fixed active), base_version and pending: {count, fingerprint}. BatchWorker accepts a batch that names the view or its base.
  • A view is never a base of another view, its payload is refused by MappingRegistry.from_json, and compile_registry_snapshot refuses it as a parent.

Publish certifies, per item

compile_registry_snapshot(active_registry, approvals, catalog, *, collect=False, panel_index=None, ...). With collect=True a decision that cannot be published never stops the others: the decisions are taken in the order of reviewed_at, the earlier one stands, and each that fails comes back alone in RegistrySnapshot.returned_decisions (also manifest.returned_decisions):

{
  "decision_id": "...", "case_id": "...", "reviewed_at": "...",
  "reason_code": "conflict",
  "message": "<reviewer> approved <code> (<name>) for this test on <date> (not yet published). Keep it, or replace it with <code> (<name>) for all future reports?",
  "error": "review decision ... conflicts with active entry ...",
  "other": {"decision_id": "...", "alias": "...", "target": "...", "code": "...", "display": "...", "state": "pending", "reviewer_name": "...", "reviewed_at": "...", "registry_version": null}
}
  • reason_code is one of COMPILE_REASON_CODES: not_eligible, publication_policy, force_exact_conflict, conflict. other.state is published for an entry of the parent and pending for an earlier decision of the same list; other.decision_id is null for an approval older than round K, which is named by alias and target.
  • manifest.decision_ids lists every decision the snapshot publishes (added and unchanged); returned_decisions lists the rest; they never overlap.
  • A confirmation is never returned. It is unchanged when an earlier approval already says it (the same code for a question at least as general), and an entry of its own when it records a question the earlier one does not cover (the bare label confirming the code approved under the chemistry panel): that entry is what lets the next bare row file beside another panel's approval.
  • Two questions of one name publish side by side.
  • A replacement removes what it replaces only once it is known to publish (manifest.superseded_entries, with replaced_decision_id). An earlier decision of the list that a later one replaces is published by neither. A replacement of an approval that never reached the registry compiles and is recorded in manifest.superseded_unpublished; it excuses nothing else.
  • With nothing publishable the snapshot is no_change. Without collect the first failure raises ReviewContractError with the sentence it always had, and the manifest has no returned_decisions key.
  • panel_index lets the conflict rule compare panels by meaning, as validation does.

tools/compile_clinician_approvals.py compiles with collect=True and prints each returned decision; python -m loinc_mapper compile-review-snapshot --collect does the same for a CSV batch.

Practice policy

The decisions of the main project's BUG-88 are practice policy, not analyte knowledge. They live in config/axis_safety_policy.json (practice_policy, version 2026.10.08.1) and a different answer is a change there, with a version bump:

Key Value Meaning
question_facts the six facts beside the name which printed facts make up a question
panel_splits_question meaning two printed panels split a question only when they change the meaning; always splits on every panel
value_kind_binds_approvals true an approval made on one kind of value does not file another kind
method_default practice_then_method_less with no method printed: the practice's approved code for the analyte, else the method-less term; method_less; ask
material_default min_codes: 5, min_share: 0.70 the default material of a test family comes from the practice's own approvals of that family (and, since 2026-10-08, of its microscopy terms: method:microscopy)
property_default MCnc with no unit printed at all, the Property siblings of an analyte the practice never approved join under this Property family (Dr Tro, 2026-10-08); the row files flagged unit_missing
scale_default urine_microscopy: {number: count_per_field, word: presence} a urine microscopy result with no unit reads as the count per field on a bare number and as the presence term on a degree word (Dr Tro, 2026-10-08)

Publisher Job Status

The package emits PublicationProgressEvent values through an optional callback. The main project stores and displays these job states without holding an HTTP request open: queued, validating, batch_queued, batch_completed, finalization_ready, published, no_change, and failed_safety. The package emits the compile/replay/finalization states; the main project emits published only after it atomically advances its active registry pointer.

Only new or superseding decision revisions should be supplied to the publisher. If all submitted approvals already exist in the active snapshot, candidate.requires_publication is false and the main project records those decisions as audit-only without replaying or writing another registry object.

Specimen Contract

The UI should use select controls for canonical specimen values and preserve the original OCR wording separately:

AnalyticalSpecimen:
unknown, blood_unspecified, whole_blood, serum, plasma, serum_or_plasma,
urine, cerebrospinal_fluid, saliva, stool, sputum, amniotic_fluid,
synovial_fluid, pleural_fluid, peritoneal_fluid, body_fluid_unspecified,
tissue, other

CollectionSpecimen:
unknown, blood_venous, blood_arterial, blood_capillary,
urine_clean_catch, urine_catheter, urine_24_hour, other

Send the enum value in specimen_code or collection_specimen_code and keep the raw wording in specimen or collection_specimen. For example, collection_specimen_code=blood_venous with collection_specimen="Blood, Venous". The mapper rejects contradictory code and text instead of silently choosing one. Bld, Ser/Plas, and similar values are LOINC System-axis representations, not clinician-facing UI values.

Units and Result Styles

unit is intentionally not an enum. UCUM units are composable expressions, so a finite dropdown would reject valid forms and would confuse a result style with a physical unit. The main project should preserve the report's unit text and pass it to the mapper; the mapper canonicalizes and validates it with the UCUM-compatible parser. Examples include mg/dL, mcg/dL, K/mcL, mL/min/1.73mE2, and %.

Some LOINC results do not have a physical UCUM unit. For presence, interpretation, and titer results, pass unit=null and use the report/result style in scale when it is known. In particular, never send unit="Titr": Titr is a LOINC Property, not a unit, and stays invalid. {titer} is the release's annotation token: a printed titer parses as the non-dimensional {titer} and is compatible only with a titer-style term, so a report that prints it can be passed through unchanged. For an immunofluorescence titer such as LOINC 5307-4, the correct review correction is still unit=null, with the selected code providing the titer semantics.

The same rule applies to a LOINC result whose official example is {ratio}. For example, CHOL/HDLC RATIO may be approved as 9830-1 with unit=null; the printed number is a unitless ratio. This is target-specific, not a broad exception for every Property containing Ratio: a target whose official example is mg/g still needs a compatible reported unit. Since policy 2026.09.17.1 a ratio-of-like-quantities property (Rto, MRto, SRto, CRto, DRto, NRto) whose term has no example unit at all, such as the GGT/AST ratio 2325-9, is also approvable with unit=null. The review API enforces this distinction before publication.

LOINC method values are compact axis codes. The validator recognizes standard report surfaces such as ImmunoBlot -> IB and IFA -> IF while retaining the target method as the authoritative axis. This is method-axis normalization, not a manual analyte mapping.

An explicit antibody isotype is a safety discriminator. A row named IgM cannot be approved as an IgG code. For example, 6320-6 is the LOINC code for B. burgdorferi IgG by immunoblot; the corresponding serum IgM immunoblot code in the pinned release is 6321-4. A TBRF IgM row is not a Lyme IgG row.

Infectious-report example

The IGeneX infectious report bootstrap is append-only at config/mapping_registry.igenex.20260828.json. It contains 14 structured single-result mappings plus the report-only B. hermsii mapping. The remaining 12 report rows are retained in bootstrap_audit.unresolved_observations and are intentionally not given unrelated LOINC codes. TBRF immunoblots cover multiple Borrelia organisms, while the LOINC antibody terms are generally organism and/or isotype specific; a single-organism code would change the clinical meaning.

For this report, use the snapshot as the runtime registry and pass the exact row to the review contract. A 6320-6 approval is valid for a Lyme IgG immunoblot row, not for TBRF Borrelia ImmunoBlot IgM. A 5307-4 approval for R. rickettsii IFA - IgG must use unit=null; Titr is a LOINC property and {titer} is a result-style example, not a UCUM unit. ImmunoBlot and IFA are normalized to the LOINC method surfaces IB and IF during validation.

This distinction is deliberate: clinician authority can clear confidence and margin thresholds only after the selected active LOINC, specimen, method, scale, property, and unit rules pass. It cannot convert a vendor-specific multi-organism assay into a different standardized observation.

CSV Review Workflow

The module can export an uploadable CSV with original facts as read-only context. The clinician-facing required columns are:

case_id,case_revision,decision,selected_loinc,mapping_scope,
reviewer_id,reviewer_name,rationale

Conditional columns are:

corrected_raw_name,corrected_unit,corrected_specimen,
corrected_specimen_code,corrected_collection_specimen,
corrected_collection_specimen_code,corrected_method,
corrected_time_context,corrected_scale,clear_fields,
source_test_id,source_document_reference,decision_id,reviewed_at

The exported file also includes original name, value, unit, specimen, source laboratory, report/page reference, mapper status, and candidate codes. Those are review context, not fields that a CSV import is allowed to rewrite. The full CSV also includes original_specimen_code and original_collection_specimen_code.

python -m loinc_mapper export-review-csv `
  --cases review_cases.jsonl `
  --output clinician_review_template.csv

python -m loinc_mapper validate-review-csv `
  --cases review_cases.jsonl `
  --input clinician_reviewed.csv `
  --registry config/mapping_registry.json `
  --catalog assets/loinc/2.82/catalog_v5.sqlite3 `
  --output validation_report.json

python -m loinc_mapper compile-review-snapshot `
  --cases review_cases.jsonl `
  --input clinician_reviewed.csv `
  --registry config/mapping_registry.json `
  --catalog assets/loinc/2.82/catalog_v5.sqlite3 `
  --registry-output candidate_registry.json `
  --manifest-output candidate_manifest.json

The last command produces a candidate, not a publishable runtime registry. The main project must run replay through the matching artifact bundle, call finalize_registry_snapshot, upload the finalized registry/manifest, and only then move its active pointer.

CSV uploads are revision-safe: the case revision must match the stored case, one decision is allowed per case revision in one upload, and repeated identical rows receive a deterministic upload decision ID. The main project should also enforce uniqueness of decision_id in its durable ledger.

Batch approvals from stored review cases

When a clinician decides a whole abstain set at once (a publish round), the operator can compile it on their own machine with tools/compile_clinician_approvals.py instead of clicking through the UI. The tool reads the main project's stored review-case JSON files, keeps only the mapping facts and the stored mapping result, gives every case a synthetic id and an empty report reference, and pairs each case with a decision from a small JSON list (analyte name, optional unit, target code, action, rationale, optional corrected_fields, clear_fields and, for a universal approval, required_panel). With --facts instead of --cases the round starts from a list of de-identified mapping fact rows (name, unit, specimen, laboratory, panel heading with source and confidence, panel LOINC hint, report section), each becoming a synthetic case with an abstain result, so an abstain seen in production is approved from its printed facts without a stored case. Every decision becomes a ReviewDecision under the reviewer id and name given on the command line, so the registry entries carry that clinician's authority, and the mapper's panel index checks every required_panel.

The rest is the contract above, unchanged: validate_review_decision against the active snapshot, compile_registry_snapshot, a full-pipeline replay of every approval (universal templates also against an unseen laboratory), a registry fast-path replay that must agree with it, and finalize_registry_snapshot. The finalized registry and manifest are written locally; nothing is uploaded. Two rules make the batch safe:

  • An alias whose surfaces collide with an active entry, or with an earlier decision in the same batch, in the same scope (any panel) and with a different target is skipped and reported, never overwritten. The identity is the folded surface, not the compact form: PSA, FREE and PSA, % FREE share only psafree (also a verified alias of each) and may both be universal, because % makes them different analytes and the runtime keeps them apart by unit and by exactness for the row. The same label approved to the same code under a second required_panel is a second entry.
  • A unit cell that is not a unit ((calc)) is cleared with clear_fields rather than published as a required_units constraint.
  • When the stored result records a verified input repair, the case's name is the repaired mapping_name, because the runtime looks the registry up by the mapping input name: an approval of the printed MERCURY, BLOOD row publishes the alias MERCURY, and a bare repaired name is approved source-scoped rather than as a universal alias. A decision may name either surface. The contract's ReviewCase observation has no mapping_name field, so a review case built from the raw surface alone would publish an alias the runtime never looks up.
  • A second reviewed entry for a code that already has one (a source-scoped HS CRP approval beside a universal HS CRP template that requires a method column) is not a conflict: the runtime keeps one registry candidate per code, carries the other entries' constraints as registry_alternatives, and accepts the code when any matching entry is satisfied.

The decisions file for a round lives under evaluation/clinical_review/ with the finalized snapshot under config/, so the round is reproducible from the stored cases.

These are recommended routes for the larger FastAPI project. They are not implemented in this module:

GET  /api/loinc/review-cases?status=open
GET  /api/loinc/review-cases/{case_id}
POST /api/loinc/review-cases/{case_id}/decisions
GET  /api/loinc/terms?query=...
POST /api/loinc/review-imports
POST /internal/loinc/registry-publications

The term-search route should call search_active_loinc(catalog, query) and return the official name, active status, six axes, example units, class, and ORDER_OBS. It must not expose inactive or order-only terms as result-row choices. The decision route should call validate_review_decision before storing an approval event. A publisher process should call compile, replay, and finalize; it should never trust a browser-submitted registry object.

Main-Project Integration Checklist

  1. Install the pinned package and artifact bundle in the mapper worker image; do not initialize SapBERT, ScispaCy, or UMLS inside the FastAPI request process.
  2. After row extraction, build LabObservation objects with every available fact and submit them in a MappingBatch. Include registry_snapshot with the finalized snapshot version, private URI, storage generation, and SHA-256.
  3. Persist mapper results and create review cases for abstain, invalid_candidate, and no_candidate. Keep report identity and document links in the main project, not in this module.
  4. Build a review screen and an optional CSV upload that only submit the versioned ReviewDecision fields described above. Staff may draft; an authorized clinician must approve reusable mappings.
  5. Use search_active_loinc for the clinician search box and validate_review_decision before writing an approval event. Return field errors to the UI instead of publishing an unsafe decision.
  6. Run a publisher worker that compiles validated events, replays candidate snapshots with the next production artifact bundle, finalizes passing snapshots, and atomically advances the active registry pointer.
  7. Launch new Cloud Run Jobs with that exact snapshot reference. A running job must finish with its own explicit reference, never reload a mutable file.

The current version does not query patient-chart pending orders. In a later version, after demographic matching, the main project may pass an advisory OrderContext containing an Elation order-set/test identifier, local lab test identifier, CPT, diagnosis, and ordering context. It must never become a hard global LOINC filter; absent, stale, or ambiguous order context falls back to normal laboratory-wide retrieval.

Recommended durable records in the main project are:

review_cases/{case_id}: immutable case, current revision, report/page reference, mapper result
review_decisions/{decision_id}: case ID/revision, reviewer identity, action, corrections, rationale
registry_publications/{version}: parent version, event IDs, manifests, replay report, active state

Keep report files, OCR artifacts, large batch payloads, results, manifests, and registry snapshot JSON in private object storage; save only their stable object references in the review ledger. The main project must enforce authorization, audit logging, retention, deletion, and incident-response policies.

Immediate Learning Versus Model Learning

There are two safe learning loops:

  1. Immediate registry learning: a finalized clinician decision appears in the next immutable registry snapshot. No neural-model training occurs.
  2. Offline model learning: retained reviewed outcomes become training and evaluation examples only after there are enough labels. A challenger ranker may become active only after a champion/challenger evaluation preserves accepted precision, unseen-laboratory performance, and zero dangerous safety errors.

The main project should create a regression case for every approved mapping, track repeat abstentions by normalized phrase and laboratory, and surface those clusters to clinicians. It must never silently turn a repeated unknown phrase into a production mapping.

Cloud Run and Audit Boundary

The deployment sequence is detailed in Deployment and release operations are in Maintenance. In short:

  1. Main API persists review cases, decisions, status history, mapper provenance, report references, and registry version references in Firestore or an equivalent durable ledger.
  2. Main API stores PDFs, OCR artifacts, batch inputs/outputs, candidate and finalized registry snapshots in private Cloud Storage.
  3. The registry publisher writes immutable object names such as mapping_registry.<version>.<sha>.json, then changes the active pointer with a Cloud Storage generation precondition.
  4. A Cloud Run Job receives a MappingBatch with an explicit registry_snapshot reference, downloads that small finalized snapshot to ephemeral storage, constructs the mapper, and verifies its loaded version.
  5. The job records snapshot URI, generation, checksum, and registry version in results. Jobs already running finish with their previous explicit snapshot.

Firestore and Cloud Storage are deployment choices, not package dependencies. Use them only under the organization's executed BAA, least-privilege IAM, audit logging, retention policy, and security review.